TokenHot
Home
Models
ModelsGPT-5.6Claude Opus 5Claude Fable 5Gemini 3.5 FlashClaude Sonnet 5DeepSeek V4 ProKimi K3Seedance 2.5

Providers

OpenAIAnthropicGoogleDeepSeekQwenByteDanceDoubaoMiniMaxZ.ai (GLM)
ConsoleDocumentationBlog
✓ English简体中文繁體中文日本語FrançaisРусскийTiếng Việt
TokenHot

One API. A model catalog. Usage-based billing.

Product

  • Models
  • Pricing
  • About
  • Support

Popular Models

  • GPT-5.6
  • Claude Opus 5
  • Claude Fable 5
  • Gemini 3.5 Flash
  • Claude Sonnet 5
  • DeepSeek V4 Pro
  • Kimi K3
  • Seedance 2.5

Model Providers

  • OpenAI
  • Anthropic
  • Google
  • DeepSeek
  • Qwen
  • ByteDance
  • Doubao
  • MiniMax
  • Z.ai (GLM)

Resources

  • Docs
  • Blog
  • hi@tokenhot.ai
  • Terms
  • Privacy
  • Refund Policy
© 2026 TokenHot Inc. — Built for builders.
HomeBlogGuidesOpenAI API Timeout: Diagnose Slow or Interrupted Streams
Guides

OpenAI API Timeout: Diagnose Slow or Interrupted Streams

TTokenhot Team·September 29, 2026·10 min read

An openai api timeout says your client stopped waiting; it does not prove whether OpenAI received, accepted, or completed the POST. Diagnose it by saving the status and request ID if headers arrived, recording the last Responses API event seen, and separating HTTP success from application-level completion. If no documented completion event arrived, preserve the result as incomplete or unknown and do not automatically replay the POST.

This guide uses a direct HTTP request to OpenAI’s Responses API, so it can show the actual response and transport boundaries without relying on methods that vary between SDK releases. The same diagnosis applies conceptually to other clients, but their timeout controls, error objects, request IDs, and stream events can differ. OpenAI-compatible gateways also need their own endpoint and event documentation; the script below is for api.openai.com only.

What a timeout tells you

Timeouts describe a wait limit. A configured limit is not a measured phase duration, and a timeout after the request may have been written does not prove the service did no work. Keep three kinds of evidence separate:

Boundary Example setting or observation What it establishes
Connection-pool wait and connection setup aiohttp connect limit; optional transport tracing A connection could not be acquired or established before the configured wait expired. connect covers both pool waiting and connection establishment; timestamps alone do not separate those subphases.
Socket connection sock_connect limit Time allowed to connect to a peer for a new connection. It is a limit, not an observed DNS/TCP/TLS breakdown.
Request write Transport tracing, if available This example has no distinct write-timeout field. The total application deadline still bounds the whole operation. A write failure can leave the server outcome uncertain.
Response headers headers_received timestamp, HTTP status, x-request-id if supplied An HTTP response arrived. It does not prove the stream or model response completed.
Incoming bytes sock_read limit aiohttp waited too long for the next data portion. Bytes may arrive without completing an SSE event, so this is not an “event idle” timer.
Total operation Separate asyncio.timeout deadline Your application ended the operation after a wall-clock budget. The deadline is not an HTTP status.
Responses API terminal event Last event type and time response.completed documents a completed response; response.incomplete, response.failed, or error records a different application outcome.

The aiohttp timeout documentation defines connect as covering connection establishment or waiting for an available pool connection, sock_connect as the new peer connection wait, and sock_read as the wait between data portions. Use TraceConfig or your transport’s supported hooks if you need observed pool, DNS, TLS, or write timings. Do not infer them from timeout settings.

Read HTTP status and stream completion separately

A response can have HTTP status 200 and still fail to produce a complete answer. The OpenAI streaming guide describes the Responses API as server-sent events (SSE). Its event reference documents response.completed, response.incomplete, and response.failed; the stream can also emit error.

Use these distinctions when triaging:

  • No headers: there is no observed status or response request ID. If the POST might have been written, the server-side outcome is unknown.
  • Non-2xx headers: keep the HTTP status and request ID. The example does not read or log an error body, because that body can include private details. Follow the endpoint’s error documentation separately.
  • 2xx headers, then response.completed: the API reported completion. A later local cleanup or connection error does not erase that observed event; record both facts.
  • response.incomplete, response.failed, or error: record the explicit event type. Do not treat it as a completed answer.
  • EOF or transport error without a terminal event: retain any partial output in your application’s protected state, but treat the result as incomplete or unknown. A broken connection cannot establish what happened after the last event you received.

An SSE event is dispatched only when its blank-line separator arrives. If EOF occurs after a data: line but before that separator, discard the pending frame; it is not an observed response.completed event. The HTML Standard’s SSE processing rules specify that pending data is discarded at end of file. The diagnostic records an unterminated_sse_frame flag without promoting that frame to the last event.

A direct connection close after text deltas is not the same thing as a documented terminal event. Avoid parsing partial JSON, invoking tools from an unfinished response, or automatically replaying a POST with potentially duplicated side effects. Check the endpoint’s documented reconciliation or idempotency mechanism before deciding whether a retry is safe.

Run a single controlled reproduction

The runnable example below targets the OpenAI Responses API directly. It makes one POST, disables redirects, does not implement retries, and logs event names rather than response text. Running it contacts OpenAI and can incur usage charges. Use an authorized key and an exact model ID from your account; the code itself has not been run against a live API.

Download timeout_probe.py.

Use Python 3.11 or later and the exact aiohttp version validated for this script. Save the following code as timeout_probe.py:

python3.11 -m venv .venv
. .venv/bin/activate
python -m pip install 'aiohttp==3.13.5'

Set OPENAI_API_KEY through your shell’s secret manager or secure environment injection, and set OPENAI_MODEL to an available model ID. Avoid putting the key literal in shell history, source control, logs, or screenshots. Then run python timeout_probe.py.

"""One-request OpenAI Responses stream diagnostic. See article setup and limits."""
import asyncio
import json
import os
import time
from datetime import datetime, timezone

import aiohttp

OPENAI_RESPONSES_URL = "https://api.openai.com/v1/responses"
CONNECT_LIMIT_SECONDS = 5.0
SOCKET_CONNECT_LIMIT_SECONDS = 5.0
SOCKET_READ_IDLE_SECONDS = 30.0
TOTAL_DEADLINE_SECONDS = 90.0


def utc_now():
    return datetime.now(timezone.utc).isoformat(timespec="milliseconds")


def safe_exception_chain(exc):
    """Return class names only; never serialize exception text or headers."""
    names = []
    seen = set()
    current = exc
    while current is not None and id(current) not in seen and len(names) < 5:
        seen.add(id(current))
        names.append(type(current).__name__)
        current = current.__cause__
    return names


def parse_sse_data(data_lines):
    if not data_lines:
        return None
    data = "\n".join(data_lines)
    if data == "[DONE]":
        return {"type": "sse_done_marker"}
    try:
        event = json.loads(data)
    except json.JSONDecodeError:
        return {"type": "unparsed_sse_data"}
    if isinstance(event, dict):
        return event
    return {"type": "non_object_sse_data"}


async def run_probe(url, token, model, *, total_seconds=TOTAL_DEADLINE_SECONDS,
                    sock_read_seconds=SOCKET_READ_IDLE_SECONDS):
    started = time.monotonic()
    state = {
        "status": None,
        "request_id": None,
        "phase": "request_start",
        "first_event_at": None,
        "first_event_seconds": None,
        "last_event_type": None,
        "last_event_at": None,
        "terminal_event": None,
        "application_outcome": None,
        "exception_classes": [],
        "deadline_expired": False,
        "sse_frame_pending": False,
    }

    def log(record):
        print(json.dumps(record, separators=(",", ":")))

    log({"stage": "request_start", "at": utc_now(), "model": model,
         "retry_count": 0, "timeouts_s": {"connect_including_pool": CONNECT_LIMIT_SECONDS,
         "socket_connect": SOCKET_CONNECT_LIMIT_SECONDS,
         "read_between_chunks": sock_read_seconds,
         "total_application_deadline": total_seconds}})

    deadline = None
    cancellation = None
    try:
        timeout = aiohttp.ClientTimeout(
            total=None,
            connect=CONNECT_LIMIT_SECONDS,
            sock_connect=SOCKET_CONNECT_LIMIT_SECONDS,
            sock_read=sock_read_seconds,
        )
        headers = {"Authorization": "Bearer " + token,
                   "Content-Type": "application/json",
                   "Accept": "text/event-stream"}
        payload = {"model": model, "input": "Reply with the single word OK.",
                   "stream": True}
        async with aiohttp.ClientSession(timeout=timeout, trust_env=False) as session:
            deadline = asyncio.timeout(total_seconds)
            try:
                async with deadline:
                    state["phase"] = "waiting_for_response_headers"
                    async with session.post(url, headers=headers, json=payload,
                                            allow_redirects=False) as response:
                        state["status"] = response.status
                        state["request_id"] = response.headers.get("x-request-id")
                        state["phase"] = "response_body"
                        log({"stage": "headers_received", "at": utc_now(),
                             "seconds_from_start": round(time.monotonic() - started, 3),
                             "status": state["status"],
                             "request_id": state["request_id"]})

                        if not 200 <= response.status < 300:
                            state["application_outcome"] = "http_rejection"
                            state["phase"] = "http_rejection_body_not_read"
                        else:
                            data_lines = []
                            async for line_bytes in response.content:
                                state["phase"] = "response_body"
                                line = line_bytes.decode("utf-8", errors="replace").rstrip("\r\n")
                                if line == "":
                                    event = parse_sse_data(data_lines)
                                    data_lines = []
                                    state["sse_frame_pending"] = False
                                    if event is None:
                                        continue
                                    now = utc_now()
                                    event_type = event.get("type", "unknown")
                                    state["last_event_type"] = event_type
                                    state["last_event_at"] = now
                                    if state["first_event_at"] is None:
                                        state["first_event_at"] = now
                                        state["first_event_seconds"] = round(time.monotonic() - started, 3)
                                        log({"stage": "first_event", "at": now,
                                             "seconds_from_start": state["first_event_seconds"],
                                             "event_type": event_type})
                                    if event_type in {"response.completed", "response.incomplete",
                                                      "response.failed", "error"}:
                                        state["terminal_event"] = event_type
                                        if event_type == "response.completed":
                                            state["application_outcome"] = "complete"
                                        elif event_type == "response.incomplete":
                                            state["application_outcome"] = "incomplete"
                                        else:
                                            state["application_outcome"] = "application_error"
                                    log({"stage": "event_observed", "at": now,
                                         "event_type": event_type,
                                         "request_id": state["request_id"]})
                                elif line.startswith("data:"):
                                    value = line[5:]
                                    data_lines.append(value[1:] if value.startswith(" ") else value)
                                    state["sse_frame_pending"] = True
            except TimeoutError as exc:
                state["deadline_expired"] = bool(deadline and deadline.expired())
                state["exception_classes"] = safe_exception_chain(exc)
                if state["deadline_expired"]:
                    state["phase"] = "total_application_deadline"
                else:
                    state["phase"] = "transport_timeout"
                if state["application_outcome"] is None:
                    state["application_outcome"] = "unknown_after_request_start"
            except Exception as exc:
                state["exception_classes"] = safe_exception_chain(exc)
                state["phase"] = "transport_or_stream_error"
                if state["application_outcome"] is None:
                    state["application_outcome"] = "unknown_after_request_start"
    except asyncio.CancelledError as exc:
        cancellation = exc
        state["exception_classes"] = safe_exception_chain(exc)
        state["phase"] = "caller_cancelled"
        if state["application_outcome"] is None:
            state["application_outcome"] = "unknown_after_request_start"
    except Exception as exc:
        state["exception_classes"] = safe_exception_chain(exc)
        state["phase"] = "client_setup_or_request_error"
        if state["application_outcome"] is None:
            state["application_outcome"] = "unknown_after_request_start"

    if state["application_outcome"] is None:
        state["application_outcome"] = (
            "stream_ended_with_unterminated_sse_frame" if state["sse_frame_pending"]
            else "stream_ended_without_terminal_event")
    log({"stage": "request_end", "at": utc_now(),
         "phase": state["phase"], "application_outcome": state["application_outcome"],
         "status": state["status"], "request_id": state["request_id"],
         "seconds_total": round(time.monotonic() - started, 3),
         "seconds_to_first_event": state["first_event_seconds"],
         "last_event_type": state["last_event_type"],
         "last_event_at": state["last_event_at"],
         "terminal_event": state["terminal_event"],
         "exception_classes": state["exception_classes"],
         "total_deadline_expired": state["deadline_expired"],
         "unterminated_sse_frame": state["sse_frame_pending"]})
    if cancellation is not None:
        raise cancellation
    return state


async def main():
    token = os.environ["OPENAI_API_KEY"]
    model = os.environ["OPENAI_MODEL"]
    await run_probe(OPENAI_RESPONSES_URL, token, model)


if __name__ == "__main__":
    asyncio.run(main())

The script emits newline-delimited JSON. Its end record separates the last fully dispatched event from the unterminated_sse_frame flag. An unfinished data: line is not parsed into an event at EOF; if an earlier response.completed was fully framed, that outcome remains recorded.

Keep a safe incident record

Incident ID / UTC date:
Client and exact version: Python ___ / aiohttp ___
Endpoint owner: direct OpenAI / named compatible gateway (verify its contract)
Endpoint path (no query string):
Model or route ID:
Streaming: yes     Client retries: none in this script     Other retries: ___
Configured limits: connect ___ / sock_connect ___ / sock_read ___ / total deadline ___
Request start UTC:
Response headers UTC / status:
Response request ID (if supplied):
First parsed event UTC / type:
Last parsed event UTC / type:
Terminal event observed: yes/no; event type:
Unterminated SSE frame at EOF: yes/no:
End condition: EOF / read timeout / total deadline / cancellation / other
Exception class chain (types only):
Business outcome: complete / explicit incomplete / application error / unknown
Prompt fingerprint (not prompt text):
Redaction checked: authorization / prompt / response text / error body / query parameters

Send a provider support request only through its documented support route. Include the request ID if available, timestamp and time zone, endpoint/model, status, sanitized exception types, and last event type. Never attach an authorization header, full prompt, generated text, or unredacted error body to a general log.

FAQ

Does an OpenAI API timeout mean the request failed?

No. It means the client stopped waiting at a transport limit or the application’s deadline. If the POST might have reached the service, the final server-side outcome is unknown unless you received a status or documented terminal event that resolves it.

Is HTTP 200 enough to accept a streamed response?

No. It establishes that an HTTP response began successfully. For the Responses API stream, wait for the documented response.completed event before marking the answer complete; preserve response.incomplete, response.failed, error, and EOF without a terminal event as different outcomes.

Should I retry after a stream disconnect?

Not automatically. First check whether the operation is safe to repeat and whether the endpoint documents idempotency or request reconciliation. If the POST may have been accepted and there is no supported reconciliation path, record an unknown outcome and resolve it before resubmission.

References

  • OpenAI streaming responses
  • OpenAI Responses API streaming events
  • aiohttp client quickstart: streaming and timeouts
  • aiohttp client reference
  • Python asyncio.timeout
  • HTML Standard: interpreting an SSE stream
  • Tokenhot’s existing DeepSeek API setup and checks
Summary

Trace an OpenAI API timeout by stage, record status and stream events, and separate a completed response from an unknown POST outcome.

Back to Blog

Related Articles

How to Use the DeepSeek API Outside China: Setup and Checks

How to Use the DeepSeek API Outside China: Setup and Checks

August 17, 2026
N

Nano Banana Pro and 2 API: Choose a Tokenhot Route and Save Images in Python

September 29, 2026
GPT Image 2 API: Generate, Edit, and Save Images in Python

GPT Image 2 API: Generate, Edit, and Save Images in Python

September 24, 2026