TokenHot
Home
Models
ModelsGPT-5.6Claude Opus 5Claude Fable 5Gemini 3.5 FlashClaude Sonnet 5DeepSeek V4 ProKimi K3Seedance 2.5

Providers

OpenAIAnthropicGoogleDeepSeekQwenByteDanceDoubaoMiniMaxZ.ai (GLM)
ConsoleDocumentationBlog
✓ English简体中文繁體中文日本語FrançaisРусскийTiếng Việt
TokenHot

One API. A model catalog. Usage-based billing.

Product

  • Models
  • Pricing
  • About
  • Support

Popular Models

  • GPT-5.6
  • Claude Opus 5
  • Claude Fable 5
  • Gemini 3.5 Flash
  • Claude Sonnet 5
  • DeepSeek V4 Pro
  • Kimi K3
  • Seedance 2.5

Model Providers

  • OpenAI
  • Anthropic
  • Google
  • DeepSeek
  • Qwen
  • ByteDance
  • Doubao
  • MiniMax
  • Z.ai (GLM)

Resources

  • Docs
  • Blog
  • hi@tokenhot.ai
  • Terms
  • Privacy
  • Refund Policy
© 2026 TokenHot Inc. — Built for builders.
HomeBlogGuidesOpenAI-Compatible APIs: What to Verify Before You Migrate
Guides

OpenAI-Compatible APIs: What to Verify Before You Migrate

TTokenhot Team·October 4, 2026·13 min read
OpenAI-Compatible APIs: What to Verify Before You Migrate

Cover: a conceptual illustration of checking interface fit; it does not represent a completed API compatibility test.

Changing an API base URL can get a basic request through, but it does not show that a production workload will behave the same way. The client can reach a compatible route while a particular model, endpoint, stream event, tool round trip, structured response, or usage field still differs.

Use a small compatibility suite against the exact provider, API surface, model ID, and features your application needs. This guide includes an offline-tested Python harness with a live mode you can opt into yourself. The offline tests check the harness's parsers and pass/fail rules; they do not call a model or establish compatibility for Tokenhot or any other provider.

Compatibility is specific to a route and a workload

“OpenAI-compatible” describes an interface relationship, not one universal conformance guarantee. A provider may expose the familiar Chat Completions shape while offering Responses on a different route, or a model may accept a feature only with particular parameters. Treat the following as one test identity:

provider/account + base URL + endpoint + exact model ID + SDK version + feature set + test date

Write down what the application actually uses. If the service handles ordinary text through Chat Completions, but your production agent relies on Responses, function calls, and JSON Schema output, a passing “hello” request only verifies the first item.

Also distinguish the evidence types. A provider page that lists an endpoint is documented support. A request that returned an HTTP 200 is request accepted. A completed stream with the expected terminal event is protocol checked. A validated tool round trip or schema-checked output is feature tested. None alone proves output quality, production reliability, privacy terms, or the final charge.

Separate the base URL, endpoint, and model ID

The client base URL is usually the API root expected by that SDK. An endpoint is the operation path the client appends or invokes. The hostname identifies the service, while the key belongs to whichever service issued it. Never move a credential between providers just because both accept a bearer token.

For example, Tokenhot's public GPT Chat Completions reference and GPT Responses reference list separate routes: POST https://api.tokenhot.ai/v1/chat/completions and POST https://api.tokenhot.ai/v1/responses. Their request examples use gpt-5.6-sol. These are documentation examples, not a live test result or a promise that every model/feature combination is available. Recheck the current model and endpoint documentation before running a test.

Common setup errors include putting /chat/completions in a setting that expects a base URL, omitting /v1, appending /v1 twice, and using a key from a different account. Verify the assembled request URL without printing the authorization header. Avoid diagnosing solely from an HTTP status: a 400 can mean an unsupported parameter or malformed payload; a 401/403 can involve the key, account, permissions, or endpoint; a 404 can be a wrong path or unavailable model; and a 429 can represent a rate or quota condition. Read the provider's error body in your own secure environment and retain its request ID, but redact tokens and sensitive prompt text from reports.

Identify Chat Completions and Responses separately

Chat Completions commonly uses messages and returns choices[].message; Responses uses input and returns a response object with output items and a completion status. Their event streams and tool-result formats also differ. Being able to send a request to one does not establish that the other surface works.

Tokenhot documents both routes above, and its Chat reference shows a choices response with a usage object. Its visible Responses reference lists the endpoint and fields such as input, stream, and tools; it does not establish successful responses, usage fields, or behavior for every feature. Treat the two route listings as separate entry points, not evidence of identical feature support. The harness records each surface independently, and documentation stays separate from test results.

If you use only one surface, test that one. If you are migrating an application that calls both endpoints, exercise both with the same small, non-sensitive input. Do not infer that a model alias works across surfaces merely because the provider's model catalog displays it; verify the exact model ID against the exact operation.

Test streams for content and a real terminal event

A stream can deliver text and then disconnect before it completes. A non-empty first chunk proves only that some data arrived. For Chat Completions, the harness requires non-empty assistant text and the normal terminal finish_reason="stop", matching the completion state described in OpenAI’s Chat Completions API reference. For Responses, it requires a non-empty text delta and response.completed; a failed or incomplete response fails the check.

The check is intentionally conservative for its ordinary text fixture: Chat must end with exactly one unambiguous finish_reason="stop"; any disallowed or conflicting terminal reason fails even if a later chunk says stop. Responses must emit a text delta and a response.completed event in a fully blank-line-terminated SSE frame, without error, response.failed, response.incomplete, or refusal events. The parser discards a pending final SSE event when the body ends before its blank-line terminator, consistent with the SSE event-stream processing rule. For tool-only or refusal workflows, use their own selected test and outcome criteria; a refusal is an explicit result, not a pass for the ordinary text fixture. Also check cancellation and partial-output handling in your application separately; this harness does not test abort propagation or long-running disconnect recovery. See OpenAI's streaming guide and Responses streaming events reference for the protocol context.

Verify tools through the entire round trip

A request containing a tools field may be accepted even if the application cannot safely finish the call. A useful test needs to confirm that the model returns the expected tool name and arguments, that your dispatcher validates those arguments, that the tool result is sent back in the correct format, and that the follow-up response completes.

The included live fixture uses an inert local function that returns a fixed value. It has no outside side effects. The harness performs a full round trip separately for Chat Completions and Responses. For Responses, it reads raw HTTP JSON from output message and function-call items; it does not depend on the SDK-only output_text convenience property. A text message accompanied by an unhandled function_call is not a final text result, so the basic/JSON checks and the tool check’s final follow-up reject unresolved calls. OpenAI’s Responses text-generation guide shows the wire output structure and notes that the top-level output_text convenience field is provided by SDKs. Passing means that this example's protocol and payload completed for the tested route; it does not validate your production tool authorization, argument policy, retries, or business operation. OpenAI's function-calling guide shows the call/result cycle; each provider can differ in support and details.

Distinguish parseable JSON from schema-valid structured output

Plain JSON mode can return text that parses, while still omitting a required key or giving a field the wrong type. The harness parses the result and checks a tiny local schema (ok, count, and label). That verifies application-side validation for the fixture, not strict server-side JSON Schema enforcement on Chat Completions. OpenAI distinguishes JSON mode from schema adherence in its Structured Outputs guide.

For Responses, it asks for a strict JSON Schema format and then validates the returned object locally too. It reads text or refusal content from the raw Responses wire shape: message items in output contain typed content items. A refusal is not a schema pass, nor is an incomplete response. The script requires status="completed" and a parseable, schema-valid object. Before using structured output in production, add required/optional-field cases, boundary values, invalid inputs, and refusal handling relevant to your product. Check the target model's current API reference for its supported schema subset.

Run the harness offline, then opt into live calls

The harness uses Python's standard library. Save compat_check-v3.py and its companion compat-fixtures-v3.json in the same directory. Start with offline tests:

python3 compat_check-v3.py --self-test

Expected result:

offline self-tests: PASS (43 fixture assertions; no network)

These local fixtures cover fully framed event parsing, malformed or conflicting terminal states, unresolved tool calls, body-read failure metadata, empty output, and basic schema validation. They make no network request and incur no model usage.

Live mode is opt-in and has not been run for this guide. It requires explicit surface and feature selection, so you can run just what you need without repeating checks that already passed. One check may contain more than one request: the tool check needs a call and its follow-up. Selected checks run independently and the script makes no automatic retries. Each request's HTTP status and available request ID, plus each pass/fail/unknown result, is saved to a new JSON report. If the body read fails after headers arrive, the report still retains the observed status and request ID without recording the raw body or exception message. A failed check does not erase earlier results. Run only a check that the current provider documentation says to try; requests may incur charges. Use a harmless, non-sensitive prompt and a key with appropriate limits where available. Set credentials in your shell or secret manager; do not put them in a shared file or paste them into a ticket.

export LLM_BASE_URL="https://api.tokenhot.ai/v1"  # replace after checking provider docs
export LLM_MODEL="gpt-5.6-sol"                   # replace with the exact current route ID
export LLM_API_KEY="<set-this-outside-the-script>"
python3 compat_check-v3.py --live --surface responses --check basic

To test just Chat streaming, select --surface chat --check stream. To test both surfaces' basic requests, repeat both surface flags: --surface chat --surface responses --check basic. Check options are basic, stream, tools, and json. An unselected cell is recorded as unknown with reason not selected; it is not a pass or fail. For a Chat tool round trip, use --surface chat --check tools. Reuse an earlier report instead of rerunning checks that already passed; each report path must be new, and the script refuses to overwrite one. The script never prints the key or upstream error body. It saves each selected result as it finishes, continues with other explicitly selected checks after a failure, and does not automatically repeat requests. Usage objects are copied from responses when present and labeled provider-reported usage. Billing remains unknown because the harness cannot query an account ledger or invoice.

The gpt-5.6-sol example reflects Tokenhot's published request examples checked on September 28, 2026. Inspect the current Chat or Responses route documentation and model directory for the operation and model you plan to test. Public documentation can change, and the offline run does not certify that either endpoint, model, stream, tool feature, schema feature, or usage object passed on a live service.

Record the result by surface and feature

Keep one row per endpoint and model combination. “Unknown” is a useful state: do not convert a missing usage field, undocumented endpoint, skipped test, timeout, or unavailable result into a pass or a zero charge.

Surface / feature Documentation checked Offline fixture Live route status Pass condition Result here
Chat Completions basic request Tokenhot GPT Chat reference Synthetic raw fixture checked Not run HTTP success; finish_reason="stop"; non-empty text; no refusal/tool handoff Unknown
Chat streaming Provider's Chat stream spec Synthetic positive and negative events checked Not run Non-empty text, no refusal/error, exactly one unambiguous normal finish_reason="stop" Unknown
Chat tool round trip Provider's function-call spec Synthetic request/result fixtures checked Not run tool_calls handoff; expected name/arguments; local result returned; final text completes normally Unknown
Chat JSON + local schema Provider's JSON-mode spec Synthetic JSON and invalid-schema fixtures checked Not run Normal completion; parses and meets local field/type checks; not proof of server schema enforcement Unknown
Responses basic request Tokenhot GPT Responses reference Synthetic raw wire response checked Not run status="completed"; non-empty text in output message content; no refusal or unresolved function call Unknown
Responses streaming Provider's Responses event spec Synthetic positive and negative events checked Not run Non-empty text delta and fully framed response.completed; no explicit error, failed, incomplete, or refusal event Unknown
Responses tool round trip Provider's function-call spec Synthetic raw call/result fixtures checked Not run Completed function call with valid arguments; output submitted; raw follow-up message completes normally with no unresolved subsequent call Unknown
Responses JSON Schema Provider's structured-output spec Synthetic raw output and validator fixtures checked Not run Completed raw response; valid JSON meets local schema; refusal/incomplete is not counted as pass Unknown
Usage and billing Usage reference + account billing terms No billing simulation Not run Request-level usage fields reconciled to account usage/charge record Unknown

After a live run, record provider/service, region/account context if relevant, base URL, endpoint, exact model identifier, SDK version, date/time, test script version/hash, status per row, request IDs, returned usage object, and any account billing record. The JSON report retains selected results; unselected features remain unknown. Never include keys, full sensitive prompts, or personal data. Compare token counts or other units using that provider's current pricing unit. Reconcile billed totals after the provider's reporting delay; request usage and account invoice data can be different views and may include caching, tools, retries, taxes, or credits differently.

Diagnose from the smallest failing case

Test in layers so a failure points to a smaller cause:

  1. Check hostname, TLS, and the assembled endpoint path without sending a secret.
  2. Send one tiny non-streaming request with the service's own key and current model ID.
  3. Add only the operation your app uses (Chat or Responses), then one feature at a time.
  4. Read the service's documented error code and request ID; vary one field per retry and keep automatic retries disabled during the first diagnosis.
  5. Run the feature fixture with the exact model and capture its terminal state, not just an HTTP success.
  6. Compare returned usage with the provider's account-level usage or billing record before projecting production costs.

A 2xx response is a starting point, not an end-to-end compatibility result. A completed text request says nothing by itself about tool execution, strict schema behavior, data handling, cancellation, quota behavior, or the bill. Use the feature and operational checks your application depends on, then release traffic in a controlled way consistent with your own change process.

For Tokenhot's documented setup fields, start with the Quick Start, then open the relevant Chat or Responses reference. The OpenRouter alternatives comparison covers service selection; use this checklist for the narrower task of checking your own API workload. Browse the current model directory to confirm a route before making a request.

Summary

Check base URLs, Chat Completions, Responses, streaming, tools, structured output, and usage before moving an OpenAI client to another API provider. Includes an offline-tested, opt-in live test harness.

Back to Blog

Related Articles

Best OpenRouter Alternatives in 2026: A Practical Comparison

Best OpenRouter Alternatives in 2026: A Practical Comparison

August 14, 2026
OpenAI API Timeout: Diagnose Slow or Interrupted Streams

OpenAI API Timeout: Diagnose Slow or Interrupted Streams

September 29, 2026
API 429 Errors: Diagnose Limits and Retry Safely

API 429 Errors: Diagnose Limits and Retry Safely

October 4, 2026