Search DevTools

Jump to any tool or page

Token Counter

Calculate the number of tokens in your text using our Token Counter tool.

Input

0 tokens

Analysis

gpt-3.5-turbo

Token Count

0

Word Count

0

Estimated Cost (USD)

$0.000000

Based on input tokens only. Output tokens may vary.

How It Works

Large Language Models (LLMs) process text using tokens, which are chunks of text such as words, characters, or punctuation. The number of tokens determines both the cost and the limits of what you can send to a model.

  • Tokens vs Words: Tokens are usually smaller than words. For example, "ChatGPT" may count as two tokens: "Chat" and "GPT".
  • Cost Calculation: Each model has a price per 1M tokens. We calculate your cost on the backend.
  • Input vs Output Tokens: You are billed for both the text you send (input) and the text the model generates (output). This tool estimates input costs only.
  • Model Differences: Advanced models like GPT-5 cost more per token but are more powerful. Smaller models are cheaper and faster.

Use this tool to estimate costs before running large prompts. Always check the official provider’s documentation for the latest pricing.

Developer Utilities

About Token Counter

Count how many tokens a piece of text consumes for a given model, which is what actually determines context limits and API cost. Tokens are not words: modern models use byte-pair encoding, splitting text into subword fragments learned from a training corpus, so "tokenisation" may cost four tokens while "the" costs one, and the same string differs across tokenisers.

Frequently asked questions

Why do different models give different counts for the same text?
Each model family ships its own vocabulary learned from its own corpus. GPT-3.5 and GPT-4 use cl100k_base with about 100k merges; GPT-4o moved to o200k_base with roughly 200k, which encodes non-English text noticeably more compactly. Claude and Llama use different tokenisers again, with different vocabulary sizes and different treatment of whitespace. A prompt that fits one model's window can overflow another's even when the window size in tokens is identical.
Why does non-English text cost so much more?
BPE vocabularies are dominated by the languages in their training data, so English words often map to a single token while other scripts fall back to sub-character byte sequences. A CJK character can take two or three tokens; Devanagari, Thai and Amharic frequently cost more per character still. Emoji are especially expensive because many are multi-codepoint sequences joined by zero-width joiners, each part encoded separately. Ratios of two to four times English are common for the same meaning.
How does code tokenise compared with prose?
Worse than prose, for structural reasons. English averages roughly four characters per token; code often lands nearer three. Identifiers in camelCase or snake_case are split at the boundaries, indentation consumes tokens (though cl100k_base and later merge runs of spaces, which older encodings did not), and punctuation-dense lines fragment heavily. Minified code is the pathological case: long unique identifiers have no learned merges, so it can tokenise close to one token per two characters.
Does the token count I see here equal what I will be billed for?
It is the count for the raw text, which is the body of the cost but not the whole bill. Chat APIs wrap each message in role and delimiter tokens — a handful per message plus a few for the reply priming — and system prompts, tool and function schemas, and injected retrieval context all count as input. Output tokens are billed separately and usually at a higher rate. For budgeting, add the per-message overhead and treat this figure as a floor.
Why does adding a leading space change the count?
Because in GPT-family encodings the space is part of the token. " hello" is a single token while "hello" at the start of a string is another, distinct one, and "hello world" is two tokens rather than three. This is also why concatenating strings can produce a different total than tokenising them separately: merges apply across the join. Never assume token counts are additive over fragments — tokenise the final assembled string.