API 429 Errors: Diagnose Limits and Retry Safely

Cover: a conceptual illustration of paced requests and waiting; it does not represent measured traffic or provider-specific retry behavior.
An HTTP 429 means the endpoint refused a request under a limit or policy, but it does not identify one universal cause. Read the status, provider-documented error code or type, and any retry header before deciding what to do. A temporary request-rate limit may be retriable; an exhausted balance or spend/usage cap needs an account action. The response contract belongs to the service that received the request.
For direct OpenAI API calls, current documentation distinguishes rate-limit and rapid-traffic-increase errors from credit-balance and spend/usage limits. It also describes Retry-After and OpenAI rate-limit headers. Those body codes and x-ratelimit-* headers are OpenAI-specific examples. An OpenAI-compatible gateway may return a different body, different headers, different limits, and different retry behavior; check the gateway's own current error reference instead of assuming the OpenAI fields are present.
Diagnose the 429 before retrying
Start with the request ID, timestamp and time zone, HTTP status, endpoint owner, model/route identifier, and a redacted error code. Inspect the response body only to the extent the provider's documentation says the fields are meaningful. Never include an API key, full prompt, customer data, or raw response body in routine logs.
OpenAI's current API error guide documents several distinct 429 cases. This table describes that guide, not a universal status-code mapping:
| Direct OpenAI documented case | What it means | Next action |
|---|---|---|
| Request rate limit | Requests or tokens arrived faster than the assigned limit allows | Reduce or pace traffic. Follow a valid Retry-After header if present; otherwise use bounded exponential backoff with jitter. Check concurrent jobs sharing the applicable organization/project/model limits. |
slow_down |
Traffic ramped up too quickly; it can happen even when RPM and TPM limits are not exceeded | Reduce the request rate and ramp up gradually. Follow Retry-After when present. A blind immediate retry adds pressure. |
credit_balance_exhausted |
The organization's prepaid credit balance is depleted | Stop retrying and check the account's billing state through the authorized owner path. Waiting or backoff does not restore credits. |
| Organization or project spend limit | An enforced configured spending limit has been reached | Stop retrying. Ask the appropriate organization/project owner to review the limit under the provider's normal administrative process. |
| Organization usage limit | The provider-assigned usage limit has been reached | Stop retrying and use the provider's documented support or approved-limit process. |
The current OpenAI error-code guide names these cases separately. A gateway may use another status, body format, or distinction. If its docs do not tell you whether a response is temporary or requires action, treat the retryability as unknown and stop automated retries while you investigate. Do not switch accounts, create keys, or reroute traffic to get around a limit or provider policy.
Read headers only within their documented scope
For direct OpenAI responses, the rate-limits guide says responses can include Retry-After and request/token limit, remaining, and reset metadata. These fields help you see the state represented by that OpenAI response. They are not a promised gateway interface. A gateway can omit or transform upstream headers, so confirm what its own docs and response actually support.
When the endpoint documents a valid Retry-After value for a temporary error, treat it as a minimum: wait at least that long, then add a small random delay to avoid synchronized retries. If it is missing or invalid, use exponential backoff with jitter only for an error documented as transient. If the server's wait exceeds your job's retry window, defer or fail the job; do not retry earlier to satisfy an application timeout. A Retry-After value does not make a quota or billing error transient.
Bound retries and avoid multiplying attempts
Retries exist at more than one layer. The current OpenAI Python SDK documentation says it retries eligible failures, including 429 responses, twice by default and exposes a max_retries setting. Check the version and configuration actually installed. If your application also retries, the total number of HTTP attempts can multiply across the SDK, job queue, service mesh, and caller. Either let one layer own retries or account for every nested layer in one attempt and elapsed-time budget.
For each logical operation, define: the specific provider error codes that are safe to retry; the maximum number of total HTTP attempts; an overall deadline; the largest allowed server-directed delay; and what the job does when the deadline is reached. Apply a retry only when the 429 is documented as a temporary rejection before output or side effects began. Do not replay after consuming stream output, invoking a downstream tool, or performing another side effect unless the endpoint provides an idempotency/reconciliation contract that makes the repeat safe.
Download the retry policy helper.
The following policy helper has no API client dependency. Your adapter must extract the documented error code and headers for the actual endpoint and call send_once only after verifying it is a transient 429 that can be retried safely. Unknown codes and quota/billing cases stop. max_attempts includes the first request.
"""Bounded retry policy helper for a caller-verified transient 429."""
from datetime import datetime, timezone
import math
import random
import re
import time
_DIGITS = re.compile(r"[0-9]+", re.ASCII)
_IMF_FIXDATE = re.compile(
r"(?P<weekday>Mon|Tue|Wed|Thu|Fri|Sat|Sun), "
r"(?P<day>[0-9]{2}) (?P<month>Jan|Feb|Mar|Apr|May|Jun|Jul|Aug|Sep|Oct|Nov|Dec) "
r"(?P<year>[0-9]{4}) (?P<hour>[0-9]{2}):(?P<minute>[0-9]{2}):(?P<second>[0-9]{2}) GMT",
re.ASCII,
)
_RFC850_DATE = re.compile(
r"(?P<weekday>Monday|Tuesday|Wednesday|Thursday|Friday|Saturday|Sunday), "
r"(?P<day>[0-9]{2})-(?P<month>Jan|Feb|Mar|Apr|May|Jun|Jul|Aug|Sep|Oct|Nov|Dec)-"
r"(?P<year>[0-9]{2}) (?P<hour>[0-9]{2}):(?P<minute>[0-9]{2}):(?P<second>[0-9]{2}) GMT",
re.ASCII,
)
_ASCTIME_DATE = re.compile(
r"(?P<weekday>Mon|Tue|Wed|Thu|Fri|Sat|Sun) "
r"(?P<month>Jan|Feb|Mar|Apr|May|Jun|Jul|Aug|Sep|Oct|Nov|Dec) "
r"(?P<day> [1-9]|[12][0-9]|3[01]) "
r"(?P<hour>[0-9]{2}):(?P<minute>[0-9]{2}):(?P<second>[0-9]{2}) "
r"(?P<year>[0-9]{4})",
re.ASCII,
)
_MONTHS = {name: index for index, name in enumerate(
("Jan", "Feb", "Mar", "Apr", "May", "Jun", "Jul", "Aug", "Sep", "Oct", "Nov", "Dec"), 1)}
def _parse_http_date(value, now):
match = _IMF_FIXDATE.fullmatch(value)
kind = "imf"
if match is None:
match = _RFC850_DATE.fullmatch(value)
kind = "rfc850"
if match is None:
match = _ASCTIME_DATE.fullmatch(value)
kind = "asctime"
if match is None:
return None
fields = match.groupdict()
year = int(fields["year"])
if kind == "rfc850":
current_century = now.year - (now.year % 100)
year = current_century + year
if year - now.year > 50:
year -= 100
try:
parsed = datetime(year, _MONTHS[fields["month"]], int(fields["day"]),
int(fields["hour"]), int(fields["minute"]),
int(fields["second"]), tzinfo=timezone.utc)
except ValueError:
return None
expected_weekday = fields["weekday"]
actual_weekday = (parsed.strftime("%a") if kind != "rfc850"
else parsed.strftime("%A"))
if actual_weekday != expected_weekday:
return None
return max(0.0, (parsed - now).total_seconds())
def retry_after_seconds(value, now=None):
"""Parse RFC 9110 delay-seconds or HTTP-date; return None if invalid."""
if not isinstance(value, str):
return None
raw = value.strip(" \t")
if not raw:
return None
if _DIGITS.fullmatch(raw):
digits = raw.lstrip("0") or "0"
# A valid but enormous integer is represented as infinity so policy
# defers it as over-budget instead of falling back to a shorter wait.
if len(digits) > 308:
return math.inf
try:
return float(int(digits))
except (ValueError, OverflowError):
return math.inf
current = now or datetime.now(timezone.utc)
if current.tzinfo is None:
current = current.replace(tzinfo=timezone.utc)
else:
current = current.astimezone(timezone.utc)
return _parse_http_date(raw, current)
def bounded_retry(send_once, *, retryable_codes, repeat_is_safe,
max_attempts=4, max_elapsed=45.0, max_server_wait=30.0,
base_delay=0.5, cap_delay=8.0, jitter_seconds=0.5,
sleep=time.sleep, clock=time.monotonic, rng=random.uniform):
"""Return (result, attempts); max_attempts includes the first request.
send_once(timeout_s=...) makes exactly one request and returns status, code,
headers and body. It enforces the supplied remaining-time budget.
"""
if not repeat_is_safe:
raise ValueError("retry requires a documented safe-to-repeat operation")
if max_attempts < 1:
raise ValueError("max_attempts includes the first attempt and must be >= 1")
started = clock()
deadline = started + max_elapsed
last = None
for attempt in range(1, max_attempts + 1):
remaining = deadline - clock()
if remaining <= 0:
return last, attempt - 1
status, code, headers, body = send_once(timeout_s=remaining)
last = (status, code, headers, body)
if status != 429 or code not in retryable_codes or attempt == max_attempts:
return last, attempt
remaining = deadline - clock()
header_value = next((v for k, v in headers.items()
if k.lower() == "retry-after"), None)
server_wait = retry_after_seconds(header_value)
if server_wait is not None:
if server_wait > max_server_wait:
return last, attempt
delay = server_wait + rng(0.0, jitter_seconds)
else:
ceiling = min(cap_delay, base_delay * (2 ** (attempt - 1)))
delay = rng(0.0, ceiling)
if delay >= remaining:
return last, attempt
sleep(delay)
raise AssertionError("unreachable")
Retry-After numeric values are accepted only when they are non-negative integer seconds, as defined by RFC 9110. Values such as nan, inf, -1, or 1.5 are invalid and take the fallback path. A syntactically valid wait above max_server_wait defers the operation; the helper does not shorten it.
The code intentionally does not guess which error codes are transient. For direct OpenAI, use its current documented error guide; for a compatible gateway, use that route's own codes and header contract. Put only the provider-documented temporary rate-limit code(s) in retryable_codes; never add a broad “all 429” rule. Set repeat_is_safe=True only when the response represents a retryable rejection before any output or side effect, or when a documented idempotency mechanism covers the repeat. If either condition is uncertain, do not retry automatically. The adapter must apply the supplied timeout_s to each HTTP request so the full retry sequence respects the wall-clock deadline.
Reduce the work that caused the limit
Before raising concurrency, inspect the workload that shares the limit:
- Count overlapping requests across workers, scheduled jobs, and other services using the same organization, project, or route.
- Separate request rate from token volume. Shortening input/output, smoothing bursts, and putting a queue in front of a worker can reduce pressure; do not assume only requests-per-minute matters.
- Increase traffic gradually after a sustained period without errors. A sudden ramp can trigger a different policy than a steady workload.
- Recheck the provider's current account limit page and model-specific documentation. Limits can vary by model, organization, project, or endpoint and can change over time.
- Treat balance, spending, usage, and administrative restrictions as stop conditions. Correct the authorized configuration or ask the account owner/provider through its normal support path.
Repeatedly resending a rejected request can consume more rate-limit capacity rather than resolve the cause. OpenAI's guide explicitly warns that unsuccessful attempts count toward per-minute limits. Backoff works only for the transient category it is meant to handle; it does not repair an exhausted quota or invalid setup.
Keep a useful, redacted incident record
Incident ID / timestamp (UTC + time zone):
Endpoint owner: direct provider / named gateway:
Endpoint path (host + path only; no key/query string):
Client and exact version; retry layers/settings:
Model/route ID:
HTTP status and documented error type/code:
Request ID (if supplied):
Retry-After value (if supplied; original format):
Documented rate-limit headers present (record names/values only if appropriate):
Attempt count / overall retry deadline / stop reason:
Requests in flight / approximate workload concurrency:
Stream/output or downstream side effect already started: yes/no/unknown
Outcome: retried and completed / deferred / quota-action stop / unknown:
Prompt or payload fingerprint (not contents):
Redaction checked: authorization / key / prompt / response / customer data
Share the request ID and sanitized metadata through the provider's approved support channel. Do not paste keys, full prompts, raw response bodies, or private account details into public tickets. The current Tokenhot DeepSeek setup guide also notes that rate-limit errors and gateway codes vary by service; use its generic check as an entry point, then verify the actual route's error contract.
FAQ
Should every 429 be retried?
No. Retry only a response that the receiving service documents as a temporary rate-limit rejection and only within a bounded attempt/deadline budget. Stop for quota, balance, spend, or other errors that need an account action, and leave unknown error codes for investigation.
What if there is no Retry-After header?
If the endpoint documents the specific error as transient, use capped exponential backoff with jitter and a limit on both attempts and total elapsed time. If the error is unknown or requires billing/administrator action, stop. Do not assume a compatible gateway provides OpenAI's retry or rate-limit headers.
Does an OpenAI-compatible endpoint return OpenAI 429 fields?
Not necessarily. Compatibility does not prove identical error bodies, quota semantics, headers, retry defaults, or upstream behavior. Check the endpoint owner's current documentation and actual sanitized response.
References
Tell temporary API rate limits from quota or billing stops, honor documented Retry-After guidance, and use bounded jittered retries without multiplying requests.


