AI Token & API Cost Calculator
Estimate token usage, input/output costs, and API pricing for LLMs including OpenAI GPT-4o, Claude 3.5, and Gemini.
Estimated Token Cost
What is an AI Token?
Large Language Models (LLMs) like GPT-4o, Claude, and Gemini don't process text character by character or word by word — they process tokens. A token is a chunk of text that can be as short as a single character or as long as a common word. OpenAI's tokenizer (used in GPT models) operates on the principle that approximately 1 token ≈ 4 characters or ~0.75 words in English. Thus, 1,000 words of English text is roughly 1,333–1,500 tokens.
Token boundaries are not simply at word spaces — the tokenizer (typically based on Byte Pair Encoding, or BPE) splits text in ways that may seem unintuitive. "Tokenization" may split into ["Token", "ization"]. Numbers often take more tokens than you'd expect: "1234567890" might be split into ["123", "456", "78", "90"]. This is why token estimators use approximate ratios rather than exact word counts.
Understanding AI API Pricing Models
All major AI APIs price usage based on token counts, typically split into input tokens (your prompt + conversation history) and output tokens (the model's response). Output tokens are consistently more expensive than input tokens because generation is computationally more intensive than processing.
| Model | Input Cost (per 1M tokens) | Output Cost (per 1M tokens) | Context Window |
|---|---|---|---|
| GPT-4o | $5.00 | $15.00 | 128K tokens |
| GPT-4o mini | $0.15 | $0.60 | 128K tokens |
| Claude 3.5 Sonnet | $3.00 | $15.00 | 200K tokens |
| Gemini 1.5 Pro | $3.50 | $10.50 | 1M tokens |
| Llama 3.1 70B | ~$0.50–$0.90 | ~$0.50–$1.20 | 128K tokens |
Prices are approximate as of mid-2026 and subject to change. Always verify current pricing at each provider's official pricing page.
Practical Tips to Reduce Your AI API Costs
- Use smaller models for simple tasks: GPT-4o mini is 33x cheaper than GPT-4o for input tokens and 25x cheaper for output. For classification, summarisation of short texts, or structured data extraction, smaller models often perform just as well at a fraction of the cost.
- Optimise your system prompt: System prompts are sent with every API call. A 500-token system prompt on 10,000 daily calls costs 5 million input tokens per day. Trim every unnecessary word.
- Use caching: OpenAI and Anthropic offer prompt caching (discounting repeated prefixes at 50–90% off). If your system prompt is constant, caching can dramatically reduce costs at high volumes.
- Limit max output tokens: Set a
max_tokensparameter to prevent runaway responses. If you need a short answer, tell the model explicitly: "Reply in 2 sentences or fewer." - Batch requests where possible: OpenAI's Batch API offers 50% discounts on jobs that can tolerate up to 24-hour completion windows — ideal for bulk document processing.
How many tokens does a typical ChatGPT conversation use?
A typical back-and-forth conversation exchange (one user message + one assistant response) might use 200–800 tokens in total. A longer analytical task — like summarising a long document or writing a detailed essay — might consume 2,000–8,000 tokens per call. At GPT-4o pricing, a 5,000-token call (3,000 input + 2,000 output) costs approximately $0.015 + $0.03 = $0.045 per call.