By Abdullah Zulfiqar · Vibe Coder Journal · Checked against Anthropic’s API, SDK and Claude Code documentation on 26 September 2026
“API Error: Overloaded” (HTTP 529) is the Claude API’s way of saying it is temporarily at capacity. Anthropic’s docs say it happens when the API “experiences high traffic across all users”. The trigger is load on Anthropic’s side, not your prompt or your code.
If you saw API Error: 529 Overloaded or overloaded_error, here’s what to do:
- In Claude Code: it has already retried for you. Run /model to switch models, or start it with –fallback-model sonnet so it switches automatically whenever your main model is overloaded (each switch lasts for that turn).
- In your own app: retry with growing, randomised waits, and fall back to a second model if the first stays overloaded. The code is below.
- Either way: check status.claude.com for a posted incident.
Switching models works because Anthropic tracks capacity per model. Claude Code’s own documentation also says that a 529 “is not your usage limit and doesn’t count against your quota”.

Pick the branch that matches where you hit the error
What does “API Error: Overloaded” mean?
Anthropic’s API reference lists status 529 as overloaded_error, described in one line: “The API is temporarily overloaded.” The error body looks like this:

The three parts of a 529 response
Two things follow from Anthropic’s wording, and together they decide the fix:
- It’s temporary. Claude Code’s message says it plainly: “this is usually temporary”.
- It’s caused by overall load. The trigger is high traffic across all users, so rewording your prompt or tweaking settings won’t fix it. Waiting, or sending the request to a different model, will.
You’ll meet the same error in a few different forms depending on where you hit it:
| Where you see it | What it looks like |
|---|---|
| Raw API response | HTTP 529, “type”: “overloaded_error”, “message”: “Overloaded” |
| Claude Code | API Error: Repeated 529 Overloaded errors. The API is at capacity — this is usually temporary. Try again in a moment. If it persists, check https://status.claude.com. |
| Python SDK | An anthropic.OverloadedError (status_code 529) |
| A streaming response | An event: error carrying overloaded_error, after the stream has already started |
529 vs 429 vs 500: which Claude error do you have?
These three get mixed up all the time, and each has a different fix. The quickest way to tell them apart is to ask who caused it.

A 529 is caused by overall load. A 429 is caused by your own usage.
- 529 overloaded_error: the API is busy across all users. Back off and retry, or change model.
- 429 rate_limit_error: your organisation hit a limit. That can be a rate limit, your usage tier’s monthly spend cap, or a spend limit on the Claude Code workspace. According to Anthropic, a 429 from a monthly spend cap has no retry-after header and “keeps failing until access resumes”, so retrying won’t help there.
- 500 api_error: something failed inside Anthropic’s systems. Retry with backoff, and contact support with the request ID if it keeps happening.
One edge case is worth knowing if you’re launching something new. Anthropic says that, in rare cases, a sharp jump in your organisation’s usage can produce 429 errors from “acceleration limits”, even while you’re under your normal limits. The advice is to ramp traffic up gradually. That’s a 429, not a 529, but it tends to appear on launch day, which is exactly when people assume the API is overloaded.
Why do API users hit 529 before Claude Code users do?
Two people can hit the same busy period and have very different experiences. The difference is how many retries happen before the error reaches them.

Automatic retries before a 529 reaches you
| Where you call Claude from | Automatic retries before you see the error |
|---|---|
| Claude Code | Up to 10 by default, with exponential backoff |
| Anthropic SDKs (Python, TypeScript and others) | 2 by default, with a short exponential backoff |
| Raw HTTP (curl, fetch) | 0, unless you write them |
When Claude Code shows Repeated 529 Overloaded errors, it has already failed several times in a row. Its automatic retries apply to overloads that arrive before any of Claude’s response has streamed.
An SDK client left on default settings stops after two quick retries. During a busy period that often isn’t enough, so the error reaches your code.
If you build on the API, the retry setup is the first thing to change.
How do I fix Claude error 529 in my own code?
The approach has three parts:
1. Retry with growing waits.
2. Randomise each wait (“jitter”), so thousands of clients don’t all retry at the same moment.
3. Move to a second model if the first one stays overloaded.

What the code below does, step by step
Here’s a Python version built on the official SDK:
import random
import time
import anthropic
# Turn off the SDK’s own retries so this function is the only retry layer.
# Otherwise each attempt below would quietly become three requests.
client = anthropic.Anthropic(max_retries=0)
# Capacity is tracked per model, so a second model is often free
# when the first is overloaded.
MODELS = [“claude-opus-5-5”, “claude-sonnet-5”]
def retry_after_seconds(error):
“””Wait time the API asked for, if any (same headers the SDK checks).”””
headers = error.response.headers
for name, scale in ((“retry-after-ms”, 1000), (“retry-after”, 1)):
try:
return float(headers.get(name, “”)) / scale
except ValueError:
continue
return None
def create_with_529_handling(attempts_per_model=3, max_wait=30,
deadline_seconds=90, **kwargs):
start = time.monotonic()
last_error = None
for model in MODELS:
for attempt in range(attempts_per_model):
try:
return client.messages.create(model=model, **kwargs)
except anthropic.OverloadedError as e: # raised for HTTP 529
last_error = e # keeps the request_id
if attempt == attempts_per_model – 1:
break # last try on this model: switch now
wait = retry_after_seconds(e) or random.uniform(
0, min(max_wait, 2 ** (attempt + 1)))
wait = min(wait, max_wait)
if time.monotonic() – start + wait > deadline_seconds:
raise # out of time: surface the 529
time.sleep(wait)
raise RuntimeError(“Every model in MODELS stayed overloaded”) from last_error
message = create_with_529_handling(
max_tokens=1024,
messages=[{“role”: “user”, “content”: “Hello, Claude”}],
)
We tested this. We ran the function against the real Python SDK (v1.8.0) with a simulated API that returns 529 for Opus. It made three Opus attempts with two jittered waits (random, up to 2 s and then up to 4 s), skipped the wait after the last attempt, and Sonnet answered on the fourth call.
Why it’s written this way:
- max_retries=0 avoids stacked retries. Anthropic’s SDKs retry 5xx errors twice by default. Without this line, three attempts across two models would become up to 18 requests. Using a separate client for this function keeps the SDK’s default handling of 429s and connection errors everywhere else in your app.
- Catch OverloadedError, not InternalServerError. The current Python SDK (v1.8.0) raises its own anthropic.OverloadedError for a 529, and that isn’t a subclass of InternalServerError. Anthropic’s SDK docs page still lists every status of 500 or above under InternalServerError, but the SDK source maps 529 separately. Code that catches only InternalServerError lets every 529 straight through, which is an easy bug to ship.
- Respect the API’s hint. If a response includes retry-after-ms or retry-after, the helper uses it, checking the same two headers in the same order as the SDK.
- Keep the last error. The final exception is chained to the last 529, so its request ID still shows up in your logs and traceback.
- Cap everything. Each wait is capped, and so is the total time. A user waiting 90 seconds for an answer is already a bad experience.
- Pick a fallback your product can live with. A smaller model that answers usually beats a bigger one that doesn’t.
Does the API’s fallbacks parameter handle 529 errors?
No, although the name makes it easy to assume it does.

What server-side fallback does and doesn’t catch
Anthropic has a beta fallbacks parameter for server-side fallback. It needs the server-side-fallback-2026-07-01 beta header, and it applies to Claude Fable 5 and 5.1 and Claude Opus 5 and 5.5. It’s available on the Claude API only, not on Amazon Bedrock, Google Cloud, Microsoft Foundry or the Batches API. It exists to retry safety-classifier refusals on another model.
The documentation is explicit about everything else: “Only a safety classifier decline triggers the fallback. A rate limit, overload, or server error on the requested model is returned to you as-is.” It also says that if the fallback model is itself overloaded, the fallback attempt isn’t made.
So fallbacks=”default” won’t protect you from 529s. The SDK’s refusal-fallback middleware doesn’t either, since it is also refusal-only. Overload fallback has to live in your own retry code.
Why do I get a 529 after a 200 OK?
When you stream a response, the HTTP status goes out before Claude writes anything. If the API becomes overloaded partway through, the error arrives as an event inside a stream that has already returned 200.

The error arrives mid-stream, after partial text
Here’s how it looks, from Anthropic’s streaming docs:
event: error
data: {“type”: “error”, “error”: {“type”: “overloaded_error”, “message”: “Overloaded”}}
Anthropic notes that errors inside a stream don’t follow the standard error-handling mechanisms. In practice, the SDK’s automatic retries don’t cover them. You’re left holding a partial answer, and your code has to decide what happens next.
Anthropic describes a capture-and-resume approach. Keep the text you’ve received, then send a new request asking the model to continue. For Claude 4.6 and later models, the docs say to put the partial response inside a user message rather than in an assistant message. The sketch below uses the max_retries=0 client from the first example:
RESUME = (“Your previous response was interrupted and ended with:\n\n”
“{partial}\n\nContinue from where you left off.”)
def is_overloaded(e):
if isinstance(e, anthropic.OverloadedError): # 529 before streaming began
return True
body = getattr(e, “body”, None) # error event mid-stream
return isinstance(body, dict) and body.get(“error”, {}).get(“type”) == “overloaded_error”
def stream_with_resume(model, prompt, max_resumes=2, max_tokens=4096):
text = “”
for i in range(max_resumes + 1):
content = prompt if not text else (
prompt + “\n\n” + RESUME.format(partial=text[-1000:]))
try:
with client.messages.stream(model=model, max_tokens=max_tokens,
messages=[{“role”: “user”, “content”: content}]) as stream:
for chunk in stream.text_stream:
text += chunk
return text
except anthropic.APIStatusError as e:
if not is_overloaded(e) or i == max_resumes:
raise
time.sleep(random.uniform(0, 2 ** (i + 1))) # growing, jittered wait
return text
We tested this one too. With a simulated stream that sent part of an answer and then an overloaded_error event, the function caught the error, resent the request with the continuation prompt and returned both halves joined together.
This sketch resumes on the same model. If that model stays overloaded, pass MODELS[1] from the first example on the retry instead, the same fallback idea applied to streaming.
Two limitations are worth knowing. The joined output can repeat a few words at the seam, so trim any overlap before showing it. And according to the docs, tool-use and thinking blocks can’t be partially recovered, so resume from the most recent text block.
Why check the error body and not just the class? In the current SDK, an error that arrives mid-stream is raised as a generic APIStatusError carrying the stream’s original 200 status, with overloaded_error in its body. A 529 that arrives before streaming starts is raised as OverloadedError. The helper handles both.
How do I fix “529 Overloaded” in Claude Code?
Claude Code has done most of the work before you see the error. Here’s what’s left:
1. Check status.claude.com. If there’s a posted incident, waiting is the fix. Recent versions print this link just below the retry countdown.
2. Switch model with /model. Capacity is tracked per model, so another model often works right away. Claude Code sometimes suggests this itself, with a message like “Opus is experiencing high load, please use /model to switch to Sonnet.”
3. Set up automatic fallback so you don’t have to switch by hand. Start Claude Code with claude –fallback-model sonnet,haiku, or make it permanent in your settings:
{
“fallbackModel”: [“claude-sonnet-5”, “claude-haiku-4-5”]
}
The documentation says this kicks in “when the primary model is overloaded, unavailable, or returns another non-retryable server error.” There are a few quirks:
- The switch lasts for the current turn only, so your next message tries your main model first.
- Chains are capped at three models.
- /status doesn’t show the chain. The first sign it’s working is the notice Claude Code shows when it switches.
- Rate-limit and billing errors never trigger it.
4. Wait a few minutes and retry. Your message is still in the conversation, so typing try again works.
5. For unattended runs, set CLAUDE_CODE_RETRY_WATCHDOG=1. In CI jobs and other sessions nobody is watching, this makes Claude Code keep retrying 429 and 529 capacity errors instead of stopping. It still fails straight away on a 429 that reports a spend limit or exhausted credits.
If the overload arrives while Claude is already writing, Claude Code v2.1.199 and later keep what was produced and show API Error: Server error mid-response. The response above may be incomplete. Earlier versions discarded the partial output.
To change how many times Claude Code retries, set CLAUDE_CODE_MAX_RETRIES. The default is 10, and the value is capped at 15 unless the watchdog is on. Lowering it is useful in scripts that should fail fast.
Setting up for the first time? Our guides to installing Claude Code and checking your Claude Code usage cover the parts that are in your control. Usage causes 429s, not 529s.
Is it Claude, or is it me? A quick check
| You see | Most likely cause | What to do |
|---|---|---|
| 529, overloaded_error, “Overloaded” | The API is at capacity | Wait, back off, or switch model |
| 429, rate_limit_error | Your organisation’s limits | Slow down, spread requests out, or raise your tier |
| 429 that keeps failing, with no retry-after | Possibly your monthly spend cap | Read the error message; if it’s the cap, retrying won’t help |
| 500, api_error | A server-side fault | Retry with backoff, then contact support |
| 401, authentication_error | Your API key | Check, rotate or re-create the key |
When should I contact Anthropic support?
A 529 during a busy period is normal and doesn’t need a ticket. Get in touch if 529s continue for a long time with nothing posted on the status page. When you do, include the request ID. Every API response carries one in the request-id header, and error bodies include it as request_id. Anthropic says it helps them resolve issues quickly.
Frequently asked questions
Does a 529 error count against my Claude usage?
For Claude Code, Anthropic’s error reference says a 529 “is not your usage limit and doesn’t count against your quota.” The API documentation doesn’t say how overloaded requests are billed. Keep in mind that a 529 in the middle of a stream happens after some output has already been generated.
How long does “Claude API overloaded” last?
Anthropic only describes it as temporary and doesn’t give a typical duration. If it keeps happening, check status.claude.com for a posted incident, and switch models while you wait.
Why do I get 529 errors but my coworker doesn’t?
They may be on a different model, since capacity is tracked per model. Or they may be using Claude Code, which retries up to 10 times before showing an error, while a default SDK client retries twice.
Will upgrading my plan or usage tier stop 529 errors?
Not by itself. Raising your usage tier lifts your rate limits, which fixes 429s, not 529s. Anthropic’s Priority Tier was designed to “minimize ‘server overloaded’ errors”, but Anthropic says Priority Tier commitments are no longer available to buy. The docs also list Opus 5.5, Opus 5 and Sonnet 5 among the models it doesn’t support. For guaranteed capacity, Anthropic points you to its sales team.
Is “overloaded_error” the same as “rate_limit_error”?
No. overloaded_error (529) means the API is busy across all users. rate_limit_error (429) means your organisation went over its own limits. They need different fixes.
Can I avoid 529 errors completely?
Not completely, because they come from shared capacity. Backoff with jitter, a fallback model and stream recovery mean most of your users will never notice one. For work that doesn’t need an instant answer, the Message Batches API lets you poll for results instead of holding a live connection open. Anthropic doesn’t say batches avoid overloads, but your users aren’t left waiting on a live request.
Sources
- Claude API errors — Anthropic
- Streaming messages: error events and recovery — Anthropic
- Refusals and fallback — Anthropic
- Service tiers — Anthropic
- Python SDK: errors and retries — Anthropic
- anthropic-sdk-python source (_client.py, _streaming.py), v1.8.0 — how 529 and mid-stream errors are actually raised
- Claude Code error reference — Anthropic
- Claude Code model configuration: fallback model chains — Anthropic
- Claude status page
Related on Vibe Coder Journal: [Claude pricing plans explained](https://vibecoderjournal.com/claude-pricing-plans/) · [The best Claude model for coding](https://vibecoderjournal.com/best-claude-model-for-coding/)
Leave a Reply