TokenHot
Home
Models
ModelsGPT-5.6Claude Opus 5Claude Fable 5Gemini 3.5 FlashClaude Sonnet 5DeepSeek V4 ProKimi K3Seedance 2.5

Providers

OpenAIAnthropicGoogleDeepSeekQwenByteDanceDoubaoMiniMaxZ.ai (GLM)
ConsoleDocumentationBlog
✓ English简体中文繁體中文日本語FrançaisРусскийTiếng Việt
TokenHot

One API. A model catalog. Usage-based billing.

Product

  • Models
  • Pricing
  • About
  • Support

Popular Models

  • GPT-5.6
  • Claude Opus 5
  • Claude Fable 5
  • Gemini 3.5 Flash
  • Claude Sonnet 5
  • DeepSeek V4 Pro
  • Kimi K3
  • Seedance 2.5

Model Providers

  • OpenAI
  • Anthropic
  • Google
  • DeepSeek
  • Qwen
  • ByteDance
  • Doubao
  • MiniMax
  • Z.ai (GLM)

Resources

  • Docs
  • Blog
  • hi@tokenhot.ai
  • Terms
  • Privacy
  • Refund Policy
© 2026 TokenHot Inc. — Built for builders.
HomeBlogAPI GuidesSonnet 5.5 vs Opus 5.5: Compare Accepted Tasks and Cost
API Guides

Sonnet 5.5 vs Opus 5.5: Compare Accepted Tasks and Cost

TTokenhot Team·October 11, 2026·7 min read
Sonnet 5.5 vs Opus 5.5: Compare Accepted Tasks and Cost

Choose between Sonnet 5.5 and Opus 5.5 by asking which configuration produces acceptable work under your constraints—not which answer looks more elaborate. The useful comparison includes failed attempts, human corrections and the cost of completing the task, not just the first response.

This article supplies an evaluation method and reusable starter kit, not a benchmark result. No model was called to generate comparison data. The included task fixtures are authored examples; the result template is empty, and the summarizer deliberately reports no winner.

Fix the decision before choosing the model

Write one sentence describing the deployment choice. For example: “Which candidate should handle well-specified support-ticket extraction while meeting our validation rules?” That is a hypothetical decision, not a finding about either model.

Next, define a successful result before seeing outputs. For coding, that may mean specified tests pass and a reviewer accepts the change without unrequested dependencies. For extraction, it may mean exact fields and types with no invented values. For a tool workflow, it may mean using the required source rather than answering from memory.

Keep a small set of genuinely representative tasks separate from the starter kit. Three toy tasks can help verify your evaluation plumbing, but they cannot establish how either model performs on your entire workload.

Make configuration differences visible

Anthropic lists the native IDs claude-sonnet-5-5 and claude-opus-5-5. Its current documentation also lists different defaults: Sonnet 5.5 uses high effort, while Opus 5.5 uses medium. Opus 5.5's adaptive thinking cannot be turned off. Check the Opus model reference and effort guide rather than treating an omitted setting as a controlled experiment.

The kit starts both candidates at an explicit medium for one clearly named condition. That does not mean equal compute, identical latency or equal reasoning depth. A second condition can tune each model for your acceptance threshold, provided you report the tuning and do not mix its results into the first condition.

Keep fixed or record explicitly Why it matters to this comparison
Provider, route and endpoint A model-name comparison can otherwise also become an infrastructure comparison.
Input fixture and acceptance rubric Both candidates need the same task and standard.
System prompt, tools and local tool outputs A different tool result can change the problem itself.
Effort, output cap and streaming mode Configuration is part of the candidate being evaluated.
Cache state, retries and stop policy Record their effects rather than hiding them in a headline cost.
Transport/SDK version, run time and environment Keep enough context to reproduce a discrepant result.

Run each candidate from its own fresh history. Current Sonnet 5.5 thinking documentation describes asymmetric cross-model readability and account/history conditions. Copying signed blocks between candidates is not a neutral shortcut. Cross-model handoff belongs in a separate test with its own acceptance criteria.

Use the supplied starter cases

The companion evaluation_cases.json contains three synthetic tasks with explicit checks:

Ticket extraction. A ticket specifies a Team plan and 12 seats but no renewal month. The expected object includes renewal_month: null. A response that invents a date fails even when it is valid JSON.

Quantity-aware basket total. The candidate writes total_cents(items) using integer unit prices and quantities, with an empty-basket case and a zero-quantity case. Tests also require the inputs to remain unchanged and prohibit unrequested dependencies or external behavior. The kit provides cases but does not execute returned code. Run generated code only in an isolated, authorized environment.

A required plan lookup. A local mock tool returns the plan's seat cap. Acceptance requires the correct tool and arguments, a properly paired result, and a final answer based on that returned value. Text that happens to contain the right number without the required lookup does not satisfy the tool-use criterion.

These cases distinguish output appearance from task success. For the extraction case, either use a supported schema on both routes or run a separately labeled plain-prompt condition. Anthropic's structured-output reference distinguishes JSON output constraints from strict tool arguments; neither substitutes for checking whether the answer is substantively correct.

Record attempts, not just final successes

Copy evaluation_template.json to a new results file. It contains one unrun row for each candidate/task pair. Fill provider, endpoint and other configuration fields before conducting separately authorized calls.

After each actual attempt, record its request identity, observed status, output artifact, acceptance decision and elapsed time. Add another row for a retry or model-directed revision. Do not overwrite the failure row with the later success.

The template distinguishes these outcomes:

Observation status Meaning
not_run No attempt was made. Acceptance and charge remain null.
completed A result was obtained; it can still fail your acceptance rubric.
failed The attempt failed. Any observed charge still belongs in the cost record.
unknown The outcome cannot yet be resolved, such as an interrupted request.

A human grader should apply the declared rubric consistently. Hide candidate labels during subjective review where practical. Record reviewer identity or procedure and time spent correcting the output. The local summarizer does not grade quality or verify that a supplied billing reference is genuine.

Calculate the cost of accepted work

For a defined cohort, use this accounting measure:

cost per accepted task =
  reconciled charges for every attempted task in the cohort
  ------------------------------------------------------
             number of accepted tasks

Count failed tasks and retries in the numerator. Count an accepted task once even if it took several attempts. If there are no accepted tasks, the ratio is undefined—not zero. If any charge is unresolved, keep the cost result unknown rather than presenting an incomplete subtotal as the full cost.

Record human review time separately. You may later apply your own declared labor rate, but do not silently mix that estimate with provider charges. This separation makes it possible to distinguish “less API spend” from “less work to get an acceptable result.”

A public rate card is useful for planning, but it is not your account's invoice. Keep cached usage, discounts, retry charges and credits tied to actual records. The kit uses charge_reconciled and billing_evidence to distinguish a recorded number from a reconciled charge; it cannot prove those entries are correct.

Run the offline summarizer:

python summarize_eval.py your-results.json

The summarizer reports observed attempts, attempted and accepted tasks, first-attempt acceptances, reconciled charges and recorded review time. Missing charges stay null. Duplicate candidate/task/attempt identifiers are rejected. Running it on the untouched template reports zero observed attempts and unknown cost, not a performance score.

Before comparing its candidate summaries, confirm that both evaluated the same planned cases, with comparable repetition counts and conditions. A cheap result on one easy case cannot be compared with a complete cohort on the other candidate.

Make a conditional deployment choice

Prefer a candidate only when it meets your required acceptance and operational thresholds on the relevant cohort. A reasonable decision could be to use one configuration for routine cases and reserve another for a separately identified harder class. That is a possible routing design, not a recommendation based on results produced here.

Escalation also has a cost. A failed first attempt plus a successful second model must be counted as the two-step workflow it actually is. Do not compare that workflow's final answer against a single model while charging it for only the final step.

The starter kit makes the decision reproducible, but it leaves the outcome open. A publishable claim that one model is stronger or cheaper for a workload needs actual outputs, declared grading, a sufficiently relevant sample and complete cost records. Until those exist, the honest comparison is a method and a pending decision—not a winner.

Download the evaluation instructions.

Summary

Use reusable evaluation cases and attempt-level records to compare Sonnet 5.5 and Opus 5.5 on accepted work, retries, review time and reconciled charges.

Back to Blog

Related Articles

Sonnet 5.5 API Migration: Update More Than the Model ID

Sonnet 5.5 API Migration: Update More Than the Model ID

October 10, 2026
Claude tool_result Ordering Errors: Fix the Messages Sequence

Claude tool_result Ordering Errors: Fix the Messages Sequence

October 5, 2026
MiMo TTS API: Turn Text into a Playable WAV File

MiMo TTS API: Turn Text into a Playable WAV File

October 11, 2026

Related Models

Claude Sonnet 5Claude Opus 5