How tokens are counted
Models read text as tokens: pieces of words from a fixed vocabulary. This counter uses o200k_base, the encoding OpenAI introduced with GPT-4o, so counts are exact for models that use it. For anything else, including newer models that may ship a new encoding, treat it as an estimate. Anthropic, Google and others use their own tokenizers; for those the count is a close estimate, and the exact number is in each API's usage field.
English prose runs at roughly three quarters of a word per token. Code, numbers and non-Latin scripts use more tokens per word. Convert quickly with tokens to words, or price a month of usage with the cost calculator.