TokenHot
Home
Models
ModelsGPT-5.6Claude Opus 5Claude Fable 5Gemini 3.5 FlashClaude Sonnet 5DeepSeek V4 ProKimi K3Seedance 2.5

Providers

OpenAIAnthropicGoogleDeepSeekQwenByteDanceDoubaoMiniMaxZ.ai (GLM)
ConsoleDocumentationBlog
✓ English简体中文繁體中文日本語FrançaisРусскийTiếng Việt
TokenHot

One API. A model catalog. Usage-based billing.

Product

  • Models
  • Pricing
  • About
  • Support

Popular Models

  • GPT-5.6
  • Claude Opus 5
  • Claude Fable 5
  • Gemini 3.5 Flash
  • Claude Sonnet 5
  • DeepSeek V4 Pro
  • Kimi K3
  • Seedance 2.5

Model Providers

  • OpenAI
  • Anthropic
  • Google
  • DeepSeek
  • Qwen
  • ByteDance
  • Doubao
  • MiniMax
  • Z.ai (GLM)

Resources

  • Docs
  • Blog
  • hi@tokenhot.ai
  • Terms
  • Privacy
  • Refund Policy
© 2026 TokenHot Inc. — Built for builders.
HomeBlogPricingLLM API Pricing Comparison 2026: Cost Formula and Rates
Pricing

LLM API Pricing Comparison 2026: Cost Formula and Rates

TTokenhot Team·August 14, 2026Updated September 14, 2026·8 min read
LLM API Pricing Comparison 2026: Cost Formula and Rates

An LLM API pricing comparison is useful only when every number uses the same billing unit and the calculation reflects the workload you actually run. Input, cached input, cache writes, output, hidden reasoning, long-context premiums, tools, and retries can all land on different lines of the bill.

This guide provides a limited, reproducible snapshot of first-party text-model list prices checked on September 14, 2026. It then shows how to turn request telemetry into a monthly estimate. Prices are in US dollars before tax, enterprise discounts, cloud-marketplace adjustments, or regional premiums. Always recheck the linked source before committing a budget.

2026 text-model API pricing snapshot

The table is a practical sample, not a catalog of every model. Rates are per one million tokens on each vendor's first-party standard service unless a condition is stated. A lower token price does not establish equal quality, latency, reliability, or tokenization across models.

Provider and model Standard input Cached input or cache hit Output Conditions that change the bill
OpenAI gpt-6-astra $10.00 $1.00 $50.00 Cache writes are $12.50. Above 272K input tokens, the full request is billed at 2x input/cache rates and 1.5x output. Batch and Flex are 50% of Standard; Fast is 2x.
OpenAI gpt-5.6-luna $0.20 $0.02 $1.20 Cache writes are 1.25x uncached input, or $0.25. Above 272K input tokens, the full request costs 2x input and 1.5x output.
Anthropic Claude Sonnet 5 $2.00 $0.20 $10.00 A 5-minute cache write is $2.50 and a 1-hour write is $4.00. Batch is $1.00 input and $5.00 output.
Anthropic Claude Haiku 4.5 $1.00 $0.10 $5.00 A 5-minute cache write is $1.25 and a 1-hour write is $2.00. Batch is $0.50 input and $2.50 output.
Google gemini-3.7-flash $0.75 $0.075 $3.75 These rates run through December 31, 2026. Cached context also costs $0.50 per million tokens per hour to store. Batch input/output is $0.375/$1.875.
DeepSeek deepseek-flash $0.15-$0.30 $0.003-$0.006 $0.60-$1.20 Lower figures are off-peak; higher figures apply 01:00-04:00 and 06:00-10:00 UTC, Monday-Friday.
DeepSeek deepseek-v4-pro $0.66-$1.32 $0.022-$0.044 $1.98-$3.96 The same time bands apply. DeepSeek says V4 Pro service continues after September 14, 2026 with this billing method.
Mistral Mistral Large 3 $0.50 $0.05 $1.50 Mistral's table labels these as standard USD rates; Batch, Priority, and regional inference are separate selections.

The DeepSeek rows replace the older deepseek-chat and deepseek-reasoner prices that still appear in stale comparisons. For model-specific context and migration considerations, see our DeepSeek V4 Pro pricing guide and guide to using the DeepSeek API outside China.

First-party prices are not Tokenhot prices

The table above is a vendor list-price baseline. It is not a Tokenhot quote. Tokenhot publishes gateway routes and current rates in its model catalog, and those rates can differ from first-party pricing by model, channel, or promotion. Confirm the exact route and displayed billing unit before purchase. The examples below use only the first-party rates in the table.

This separation also matters when comparing gateways. A provider can bundle routing, payment, fallback, or model access into a different price. Our OpenRouter alternatives guide covers the operational questions to compare alongside unit cost.

The LLM API cost formula

Start from metered usage, not word counts. Let each token quantity represent the total across successful calls and any billed failed or retried calls:

token_cost = (
    standard_input_tokens * standard_input_rate
  + cache_write_tokens     * cache_write_rate
  + cached_input_tokens    * cached_input_rate
  + billed_output_tokens   * output_rate
) / 1,000,000

total_cost = token_cost
           + cache_storage_cost
           + tool_call_cost
           + media_or_other_unit_cost

Use the provider's usage object or invoice to populate the buckets. Do not estimate tokens from characters across vendors: tokenizers differ. Do not add reasoning tokens twice. If the API reports reasoning or thinking tokens inside billed output, billed_output_tokens already includes them. Google explicitly prices output including thinking tokens; OpenAI usage records expose reasoning tokens as a breakdown of output tokens.

Worked example: caching plus retry overhead

Assume a monthly gpt-5.6-luna workload, with every request below the long-context pricing threshold, records:

  • 10 million ordinary input tokens
  • 5 million cache-write tokens
  • 80 million cached-input tokens
  • 10 million billed output tokens, including any reasoning tokens

At the September 14 rates, the token bill is:

(10 * $0.20) + (5 * $0.25) + (80 * $0.02) + (10 * $1.20)
= $2.00 + $1.25 + $1.60 + $12.00
= $16.85

If the same 95 million input tokens were all billed as ordinary input, the bill would be $19.00 + $12.00 = $31.00. Under these assumptions, caching saves $14.15 before any extra tool charges.

Now add a budgeting scenario in which retries replay the same mix of tokens and raise token usage by 5%. The estimate becomes $16.85 x 1.05 = $17.6925, or $17.69 after rounding at the end. In production, calculate retry cost from actual duplicate usage because a timeout before output, a cache hit, and a full replay do not cost the same.

Six billing details that change the answer

1. Output and reasoning can dominate cost

Output often costs several times more than ordinary input. Agent loops can also generate reasoning tokens that are not visible in the final answer. Set output and reasoning limits where the API supports them, and log the complete usage record. A request that returns 300 visible tokens can still have a materially larger billed output count.

2. Caching has writes, reads, and sometimes storage

A cache hit is cheaper only after content has been written and reused. OpenAI and Anthropic charge a premium for cache writes in the examples above. Google lists both a cached-token rate and token-hour storage. Forecast stable prompt prefixes separately from changing user content, then measure the hit rate rather than assuming every eligible prompt is cached.

3. Batch discounts exchange speed for price

OpenAI GPT-6 Astra, Anthropic's listed models, and Google Gemini 3.7 Flash show 50% batch token rates in their official documentation. Batch jobs are asynchronous and subject to each provider's supported endpoints and completion window. Keep interactive traffic at standard rates and move only delay-tolerant work such as offline classification, evaluation, or backfills.

4. Long context may trigger a whole-request multiplier

Do not multiply only the tokens above a threshold unless the vendor says to. Above 272K input tokens, GPT-6 Astra explicitly applies 2x input/cache rates and 1.5x output to the full request. GPT-5.6 Luna's page specifies 2x input and 1.5x output without explicitly defining a cached-input multiplier; confirm that cache treatment before estimating a large cached request. Anthropic says Claude 4.6 and later models include their full one-million-token context at standard pricing. Rules vary by model, so record prompt size per request instead of applying one global assumption.

5. Tools and retries sit outside the headline table

Server-side search, code execution, grounding, storage, and other tools may add per-call or usage-based fees. Retries can duplicate token and tool charges even when the application shows one logical user action. Use capped exponential backoff, retry only errors documented as transient, and attach an application request ID so duplicated work can be traced.

6. Image and video prices use different units

Text-model token rates cannot price every multimodal workload. Image APIs may charge per generated image, per call, or by image tokens and quality. Video APIs commonly use seconds, resolution, or generation tiers. Keep separate formulas and subtotal the results.

Seedream is for image generation; Seedance is the ByteDance video family. Keep their prices and billing units separate. If you are changing video providers, use a video-specific comparison and the task-lifecycle checks in our Sora API migration guide.

How to choose on real workload cost

Export at least one representative week of usage by model and endpoint. Separate standard input, cache writes, cache hits, billed output, retries, long-context requests, batch traffic, and tools. Replay a fixed evaluation set against candidate models, then compare task success, latency distribution, rate limits, and total billed units. The cheapest row is useful only if the model meets your quality and operating requirements.

Recalculate after any model alias, prompt, reasoning setting, output cap, routing rule, or tokenizer change. Keep the pricing page URL and the date of every rate snapshot in your spreadsheet or cost service so later invoice differences can be explained.

Frequently asked questions

What is the cheapest LLM API in 2026?

There is no universal cheapest API. In this dated sample, DeepSeek Flash has the lowest listed off-peak cache-hit rate, while GPT-5.6 Luna has a simple low standard rate. Your cheapest valid option depends on output volume, cache reuse, time bands, tools, retries, quality, and whether batch processing is acceptable.

Are reasoning tokens billed?

Often, yes. Google states that its output price includes thinking tokens, and OpenAI reports reasoning tokens within output-token details. Use the provider's billed output total and documentation for the exact model; do not infer cost from visible answer length.

Does prompt caching always reduce cost?

No. A workload must reuse an eligible stable prefix enough times to recover cache-write and any storage charges. Track cache-write and cache-read fields, then calculate the break-even point using that model's rates.

Should I use input price alone to compare models?

No. Use the complete workload formula. Output-heavy generation, agent reasoning, tools, and retries can outweigh input price, while batch and cache hits can reduce it.

Are the prices in this article Tokenhot prices?

No. The comparison table and worked example use first-party vendor list prices checked on September 14, 2026. Check the Tokenhot model catalog separately for the current gateway route, price, and billing unit.

Summary

A dated LLM API pricing comparison with a reproducible cost formula, a worked example, and the billing details that headline token rates leave out.

Back to Blog

Related Articles

Nano Banana Pro and 2 API: Choose a Tokenhot Route and Save Images in Python

Nano Banana Pro and 2 API: Choose a Tokenhot Route and Save Images in Python

September 29, 2026
OpenAI API Timeout: Diagnose Slow or Interrupted Streams

OpenAI API Timeout: Diagnose Slow or Interrupted Streams

September 29, 2026
GPT Image 2 API: Generate, Edit, and Save Images in Python

GPT Image 2 API: Generate, Edit, and Save Images in Python

September 24, 2026