TokenHot
Home
Models
ModelsGPT-5.6Claude Opus 5Claude Fable 5Gemini 3.5 FlashClaude Sonnet 5DeepSeek V4 ProKimi K3Seedance 2.5

Providers

OpenAIAnthropicGoogleDeepSeekQwenByteDanceDoubaoMiniMaxZ.ai (GLM)
ConsoleDocumentationBlog
✓ English简体中文繁體中文日本語FrançaisРусскийTiếng Việt
TokenHot

One API. A model catalog. Usage-based billing.

Product

  • Models
  • Pricing
  • About
  • Support

Popular Models

  • GPT-5.6
  • Claude Opus 5
  • Claude Fable 5
  • Gemini 3.5 Flash
  • Claude Sonnet 5
  • DeepSeek V4 Pro
  • Kimi K3
  • Seedance 2.5

Model Providers

  • OpenAI
  • Anthropic
  • Google
  • DeepSeek
  • Qwen
  • ByteDance
  • Doubao
  • MiniMax
  • Z.ai (GLM)

Resources

  • Docs
  • Blog
  • hi@tokenhot.ai
  • Terms
  • Privacy
  • Refund Policy
© 2026 TokenHot Inc. — Built for builders.
HomeBlogGuidesAPI 429 Errors: Diagnose Limits and Retry Safely
Guides

API 429 Errors: Diagnose Limits and Retry Safely

TTokenhot Team·October 4, 2026·10 min read
API 429 Errors: Diagnose Limits and Retry Safely

Cover: a conceptual illustration of paced requests and waiting; it does not represent measured traffic or provider-specific retry behavior.

An HTTP 429 means the endpoint refused a request under a limit or policy, but it does not identify one universal cause. Read the status, provider-documented error code or type, and any retry header before deciding what to do. A temporary request-rate limit may be retriable; an exhausted balance or spend/usage cap needs an account action. The response contract belongs to the service that received the request.

For direct OpenAI API calls, current documentation distinguishes rate-limit and rapid-traffic-increase errors from credit-balance and spend/usage limits. It also describes Retry-After and OpenAI rate-limit headers. Those body codes and x-ratelimit-* headers are OpenAI-specific examples. An OpenAI-compatible gateway may return a different body, different headers, different limits, and different retry behavior; check the gateway's own current error reference instead of assuming the OpenAI fields are present.

Diagnose the 429 before retrying

Start with the request ID, timestamp and time zone, HTTP status, endpoint owner, model/route identifier, and a redacted error code. Inspect the response body only to the extent the provider's documentation says the fields are meaningful. Never include an API key, full prompt, customer data, or raw response body in routine logs.

OpenAI's current API error guide documents several distinct 429 cases. This table describes that guide, not a universal status-code mapping:

Direct OpenAI documented case What it means Next action
Request rate limit Requests or tokens arrived faster than the assigned limit allows Reduce or pace traffic. Follow a valid Retry-After header if present; otherwise use bounded exponential backoff with jitter. Check concurrent jobs sharing the applicable organization/project/model limits.
slow_down Traffic ramped up too quickly; it can happen even when RPM and TPM limits are not exceeded Reduce the request rate and ramp up gradually. Follow Retry-After when present. A blind immediate retry adds pressure.
credit_balance_exhausted The organization's prepaid credit balance is depleted Stop retrying and check the account's billing state through the authorized owner path. Waiting or backoff does not restore credits.
Organization or project spend limit An enforced configured spending limit has been reached Stop retrying. Ask the appropriate organization/project owner to review the limit under the provider's normal administrative process.
Organization usage limit The provider-assigned usage limit has been reached Stop retrying and use the provider's documented support or approved-limit process.

The current OpenAI error-code guide names these cases separately. A gateway may use another status, body format, or distinction. If its docs do not tell you whether a response is temporary or requires action, treat the retryability as unknown and stop automated retries while you investigate. Do not switch accounts, create keys, or reroute traffic to get around a limit or provider policy.

Read headers only within their documented scope

For direct OpenAI responses, the rate-limits guide says responses can include Retry-After and request/token limit, remaining, and reset metadata. These fields help you see the state represented by that OpenAI response. They are not a promised gateway interface. A gateway can omit or transform upstream headers, so confirm what its own docs and response actually support.

When the endpoint documents a valid Retry-After value for a temporary error, treat it as a minimum: wait at least that long, then add a small random delay to avoid synchronized retries. If it is missing or invalid, use exponential backoff with jitter only for an error documented as transient. If the server's wait exceeds your job's retry window, defer or fail the job; do not retry earlier to satisfy an application timeout. A Retry-After value does not make a quota or billing error transient.

Bound retries and avoid multiplying attempts

Retries exist at more than one layer. The current OpenAI Python SDK documentation says it retries eligible failures, including 429 responses, twice by default and exposes a max_retries setting. Check the version and configuration actually installed. If your application also retries, the total number of HTTP attempts can multiply across the SDK, job queue, service mesh, and caller. Either let one layer own retries or account for every nested layer in one attempt and elapsed-time budget.

For each logical operation, define: the specific provider error codes that are safe to retry; the maximum number of total HTTP attempts; an overall deadline; the largest allowed server-directed delay; and what the job does when the deadline is reached. Apply a retry only when the 429 is documented as a temporary rejection before output or side effects began. Do not replay after consuming stream output, invoking a downstream tool, or performing another side effect unless the endpoint provides an idempotency/reconciliation contract that makes the repeat safe.

Download the retry policy helper.

The following policy helper has no API client dependency. Your adapter must extract the documented error code and headers for the actual endpoint and call send_once only after verifying it is a transient 429 that can be retried safely. Unknown codes and quota/billing cases stop. max_attempts includes the first request.

"""Bounded retry policy helper for a caller-verified transient 429."""
from datetime import datetime, timezone
import math
import random
import re
import time


_DIGITS = re.compile(r"[0-9]+", re.ASCII)
_IMF_FIXDATE = re.compile(
    r"(?P<weekday>Mon|Tue|Wed|Thu|Fri|Sat|Sun), "
    r"(?P<day>[0-9]{2}) (?P<month>Jan|Feb|Mar|Apr|May|Jun|Jul|Aug|Sep|Oct|Nov|Dec) "
    r"(?P<year>[0-9]{4}) (?P<hour>[0-9]{2}):(?P<minute>[0-9]{2}):(?P<second>[0-9]{2}) GMT",
    re.ASCII,
)
_RFC850_DATE = re.compile(
    r"(?P<weekday>Monday|Tuesday|Wednesday|Thursday|Friday|Saturday|Sunday), "
    r"(?P<day>[0-9]{2})-(?P<month>Jan|Feb|Mar|Apr|May|Jun|Jul|Aug|Sep|Oct|Nov|Dec)-"
    r"(?P<year>[0-9]{2}) (?P<hour>[0-9]{2}):(?P<minute>[0-9]{2}):(?P<second>[0-9]{2}) GMT",
    re.ASCII,
)
_ASCTIME_DATE = re.compile(
    r"(?P<weekday>Mon|Tue|Wed|Thu|Fri|Sat|Sun) "
    r"(?P<month>Jan|Feb|Mar|Apr|May|Jun|Jul|Aug|Sep|Oct|Nov|Dec) "
    r"(?P<day> [1-9]|[12][0-9]|3[01]) "
    r"(?P<hour>[0-9]{2}):(?P<minute>[0-9]{2}):(?P<second>[0-9]{2}) "
    r"(?P<year>[0-9]{4})",
    re.ASCII,
)
_MONTHS = {name: index for index, name in enumerate(
    ("Jan", "Feb", "Mar", "Apr", "May", "Jun", "Jul", "Aug", "Sep", "Oct", "Nov", "Dec"), 1)}
def _parse_http_date(value, now):
    match = _IMF_FIXDATE.fullmatch(value)
    kind = "imf"
    if match is None:
        match = _RFC850_DATE.fullmatch(value)
        kind = "rfc850"
    if match is None:
        match = _ASCTIME_DATE.fullmatch(value)
        kind = "asctime"
    if match is None:
        return None

    fields = match.groupdict()
    year = int(fields["year"])
    if kind == "rfc850":
        current_century = now.year - (now.year % 100)
        year = current_century + year
        if year - now.year > 50:
            year -= 100

    try:
        parsed = datetime(year, _MONTHS[fields["month"]], int(fields["day"]),
                          int(fields["hour"]), int(fields["minute"]),
                          int(fields["second"]), tzinfo=timezone.utc)
    except ValueError:
        return None

    expected_weekday = fields["weekday"]
    actual_weekday = (parsed.strftime("%a") if kind != "rfc850"
                      else parsed.strftime("%A"))
    if actual_weekday != expected_weekday:
        return None
    return max(0.0, (parsed - now).total_seconds())


def retry_after_seconds(value, now=None):
    """Parse RFC 9110 delay-seconds or HTTP-date; return None if invalid."""
    if not isinstance(value, str):
        return None
    raw = value.strip(" \t")
    if not raw:
        return None
    if _DIGITS.fullmatch(raw):
        digits = raw.lstrip("0") or "0"
        # A valid but enormous integer is represented as infinity so policy
        # defers it as over-budget instead of falling back to a shorter wait.
        if len(digits) > 308:
            return math.inf
        try:
            return float(int(digits))
        except (ValueError, OverflowError):
            return math.inf

    current = now or datetime.now(timezone.utc)
    if current.tzinfo is None:
        current = current.replace(tzinfo=timezone.utc)
    else:
        current = current.astimezone(timezone.utc)
    return _parse_http_date(raw, current)


def bounded_retry(send_once, *, retryable_codes, repeat_is_safe,
                  max_attempts=4, max_elapsed=45.0, max_server_wait=30.0,
                  base_delay=0.5, cap_delay=8.0, jitter_seconds=0.5,
                  sleep=time.sleep, clock=time.monotonic, rng=random.uniform):
    """Return (result, attempts); max_attempts includes the first request.

    send_once(timeout_s=...) makes exactly one request and returns status, code,
    headers and body. It enforces the supplied remaining-time budget.
    """
    if not repeat_is_safe:
        raise ValueError("retry requires a documented safe-to-repeat operation")
    if max_attempts < 1:
        raise ValueError("max_attempts includes the first attempt and must be >= 1")
    started = clock()
    deadline = started + max_elapsed
    last = None
    for attempt in range(1, max_attempts + 1):
        remaining = deadline - clock()
        if remaining <= 0:
            return last, attempt - 1
        status, code, headers, body = send_once(timeout_s=remaining)
        last = (status, code, headers, body)
        if status != 429 or code not in retryable_codes or attempt == max_attempts:
            return last, attempt

        remaining = deadline - clock()
        header_value = next((v for k, v in headers.items()
                             if k.lower() == "retry-after"), None)
        server_wait = retry_after_seconds(header_value)
        if server_wait is not None:
            if server_wait > max_server_wait:
                return last, attempt
            delay = server_wait + rng(0.0, jitter_seconds)
        else:
            ceiling = min(cap_delay, base_delay * (2 ** (attempt - 1)))
            delay = rng(0.0, ceiling)

        if delay >= remaining:
            return last, attempt
        sleep(delay)

    raise AssertionError("unreachable")

Retry-After numeric values are accepted only when they are non-negative integer seconds, as defined by RFC 9110. Values such as nan, inf, -1, or 1.5 are invalid and take the fallback path. A syntactically valid wait above max_server_wait defers the operation; the helper does not shorten it.

The code intentionally does not guess which error codes are transient. For direct OpenAI, use its current documented error guide; for a compatible gateway, use that route's own codes and header contract. Put only the provider-documented temporary rate-limit code(s) in retryable_codes; never add a broad “all 429” rule. Set repeat_is_safe=True only when the response represents a retryable rejection before any output or side effect, or when a documented idempotency mechanism covers the repeat. If either condition is uncertain, do not retry automatically. The adapter must apply the supplied timeout_s to each HTTP request so the full retry sequence respects the wall-clock deadline.

Reduce the work that caused the limit

Before raising concurrency, inspect the workload that shares the limit:

  • Count overlapping requests across workers, scheduled jobs, and other services using the same organization, project, or route.
  • Separate request rate from token volume. Shortening input/output, smoothing bursts, and putting a queue in front of a worker can reduce pressure; do not assume only requests-per-minute matters.
  • Increase traffic gradually after a sustained period without errors. A sudden ramp can trigger a different policy than a steady workload.
  • Recheck the provider's current account limit page and model-specific documentation. Limits can vary by model, organization, project, or endpoint and can change over time.
  • Treat balance, spending, usage, and administrative restrictions as stop conditions. Correct the authorized configuration or ask the account owner/provider through its normal support path.

Repeatedly resending a rejected request can consume more rate-limit capacity rather than resolve the cause. OpenAI's guide explicitly warns that unsuccessful attempts count toward per-minute limits. Backoff works only for the transient category it is meant to handle; it does not repair an exhausted quota or invalid setup.

Keep a useful, redacted incident record

Incident ID / timestamp (UTC + time zone):
Endpoint owner: direct provider / named gateway:
Endpoint path (host + path only; no key/query string):
Client and exact version; retry layers/settings:
Model/route ID:
HTTP status and documented error type/code:
Request ID (if supplied):
Retry-After value (if supplied; original format):
Documented rate-limit headers present (record names/values only if appropriate):
Attempt count / overall retry deadline / stop reason:
Requests in flight / approximate workload concurrency:
Stream/output or downstream side effect already started: yes/no/unknown
Outcome: retried and completed / deferred / quota-action stop / unknown:
Prompt or payload fingerprint (not contents):
Redaction checked: authorization / key / prompt / response / customer data

Share the request ID and sanitized metadata through the provider's approved support channel. Do not paste keys, full prompts, raw response bodies, or private account details into public tickets. The current Tokenhot DeepSeek setup guide also notes that rate-limit errors and gateway codes vary by service; use its generic check as an entry point, then verify the actual route's error contract.

FAQ

Should every 429 be retried?

No. Retry only a response that the receiving service documents as a temporary rate-limit rejection and only within a bounded attempt/deadline budget. Stop for quota, balance, spend, or other errors that need an account action, and leave unknown error codes for investigation.

What if there is no Retry-After header?

If the endpoint documents the specific error as transient, use capped exponential backoff with jitter and a limit on both attempts and total elapsed time. If the error is unknown or requires billing/administrator action, stop. Do not assume a compatible gateway provides OpenAI's retry or rate-limit headers.

Does an OpenAI-compatible endpoint return OpenAI 429 fields?

Not necessarily. Compatibility does not prove identical error bodies, quota semantics, headers, retry defaults, or upstream behavior. Check the endpoint owner's current documentation and actual sanitized response.

References

  • OpenAI API error codes
  • OpenAI API rate limits
  • OpenAI Python SDK
  • RFC 9110, Retry-After
  • Tokenhot's current DeepSeek API setup and checks
Summary

Tell temporary API rate limits from quota or billing stops, honor documented Retry-After guidance, and use bounded jittered retries without multiplying requests.

Back to Blog

Related Articles

OpenAI API Timeout: Diagnose Slow or Interrupted Streams

OpenAI API Timeout: Diagnose Slow or Interrupted Streams

September 29, 2026
How to Use the DeepSeek API Outside China: Setup and Checks

How to Use the DeepSeek API Outside China: Setup and Checks

August 17, 2026
OpenAI-Compatible APIs: What to Verify Before You Migrate

OpenAI-Compatible APIs: What to Verify Before You Migrate

October 4, 2026