Claude Code Retry Watchdog: Keep Unattended Runs Alive

TL;DR
CLAUDE_CODE_RETRY_WATCHDOG=1 makes Claude Code wait out 429 and 529 errors instead of failing after 10 retries. When to use it, how to cap it, what we measured.
For an unattended Claude Code run (CI, cron, an overnight agent), set CLAUDE_CODE_RETRY_WATCHDOG=1 so 429 and 529 errors are waited out instead of failing after 10 retries, and on v2.1.295 or later add CLAUDE_CODE_RETRY_WATCHDOG_MAX_WAIT_MS so a long outage still ends the job. Leave both off at your desk.
Last updated: October 9, 2026. Behaviour below was tested on Claude Code v2.1.295 against a local mock API that answers every request with a 529.
Official Sources#
| Resource | What it confirms |
|---|---|
| Claude Code error reference: automatic retries | Default of 10 retries, what is and is not retried, the retry tuning table |
| Claude Code environment variables | CLAUDE_CODE_RETRY_WATCHDOG, CLAUDE_CODE_RETRY_WATCHDOG_MAX_WAIT_MS, CLAUDE_CODE_MAX_RETRIES, CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS and their minimum versions |
| Claude Code model configuration: fallback model chains | --fallback-model and the fallbackModel setting |
| Claude Code v2.1.295 release notes | CLAUDE_CODE_RETRY_WATCHDOG_MAX_WAIT_MS added |
| Claude API errors | What 429 rate_limit_error and 529 overloaded_error mean |
Why this matters now#
A Claude Code session at your desk can afford to fail fast. You see API Error: Repeated 529 Overloaded errors, wait a minute, and type the prompt again. A claude -p job in CI, a Ralph-style loop that runs for hours or a scheduled agent has nobody to type it again. When the API is at capacity for ten minutes, the default retry budget runs out and the job exits non-zero with half its work done.
Anthropic has been filling in this gap release by release. The watchdog itself needs v2.1.186, v2.1.199 widened it to other transient errors, v2.1.239 stopped it waiting forever on a spend limit, v2.1.292 added a tunable base delay for 529 backoff, and v2.1.295 (October 8) added a cap on how long the watchdog waits. If you read our outage playbook, this is the Claude Code half of "wait, retry with backoff, check status".
The four retry settings, side by side#
Every cell below comes from the retry tuning table and the environment variable reference.
| Variable | Default | What it changes | Needs |
|---|---|---|---|
CLAUDE_CODE_MAX_RETRIES | 10 | Number of retry attempts for a failed request. Capped at 15 unless the watchdog is on | Cap added in v2.1.186 |
CLAUDE_CODE_RETRY_WATCHDOG | unset | 1 retries 429 and 529 capacity errors indefinitely, backing off up to 5 minutes between attempts or until a stated reset time. Raises other transient-error retries to 300, roughly three hours | v2.1.186 (wider scope v2.1.199) |
CLAUDE_CODE_RETRY_WATCHDOG_MAX_WAIT_MS | unset (no limit) | Maximum time each request spends waiting out 429 and 529 errors while the watchdog is on; the next such error after that ends the request | v2.1.295 |
CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS | 500 | Starting delay of the 529 backoff, 500 to 32000. Ignored when the watchdog is on or in fast mode | v2.1.292 |
Two details in the docs are easy to miss. First, the watchdog is not a blank cheque: Claude Code "fails at once when a standard-speed request gets a 429 that reports a spend limit or exhausted usage credits", including a gateway spend cap that resets on a schedule. Before v2.1.239 it retried those indefinitely too. Second, when a 429 carries a reset time, the watchdog waits until the reset, "so a session that hits a usage limit waits out the remaining window". On a subscription that can mean a job that sits quietly for hours rather than failing, which is what you want in a nightly batch and not what you want in a PR check. Our usage limits playbook covers how those windows work.
What we measured#
We pointed Claude Code v2.1.295 at a local HTTP server that answered every POST with a 529 overloaded_error (or a 429 rate_limit_error for one row), ran claude -p "say hi" --model claude-haiku-5-5, and logged each request the server saw. This tests Claude Code's retry logic only; no real API capacity was involved.
| Setup | Requests sent | First to last request | Result |
|---|---|---|---|
| Default (no variables) | 11 (1 + 10 retries) | 176.2s | Failed with 529 |
CLAUDE_CODE_MAX_RETRIES=2 | 3 | 1.9s | Failed with 529 |
CLAUDE_CODE_MAX_RETRIES=3 + CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS=4000 | 4 | 32.8s | Failed with 529 |
CLAUDE_CODE_RETRY_WATCHDOG=1 | 8, then still waiting | 70.3s (8th request) | Still retrying when timeout killed it at 90s |
CLAUDE_CODE_RETRY_WATCHDOG=1 + ..._MAX_WAIT_MS=30000 | 7 | 30.1s | Failed with 529 |
CLAUDE_CODE_MAX_RETRIES=2 + --fallback-model claude-sonnet-5-5 | 3 on Haiku, then 3 on Sonnet | 3.5s | Failed with 529 |
429 throttle, CLAUDE_CODE_MAX_RETRIES=2 | 3 | 1.9s | Failed: "Server is temporarily limiting requests (not your usage limit)" |
What the numbers show:
- The default budget is short. Gaps between attempts roughly doubled from about half a second (0.55s, 1.26s, 2.05s, 4.04s, 9.5s, 17.7s) and then levelled off at 34 to 38 seconds, so the whole default budget was about three minutes of waiting. Any overload that outlasts that window fails the job.
- The base delay works as documented. Starting at 4000ms stretched three retries over 32.8 seconds instead of under two.
- The cap ends the wait on time. With
CLAUDE_CODE_RETRY_WATCHDOG_MAX_WAIT_MS=30000, the run gave up 30.1 seconds after its first request, on the first 529 past the limit. - A fallback chain is a second budget, not a shortcut. With
--fallback-model, Claude Code spent its retries on Haiku, then ran a fresh set on Sonnet. Against a real outage that only hits one model, that is the difference between finishing and failing, since the docs say capacity "is tracked per model".
We also confirmed that the v2.1.295 binary contains all four variable names, so a typo is the likeliest reason one of them appears to do nothing.
Pick by run type#
| Run | Settings | Why |
|---|---|---|
| Interactive session at your desk | Defaults | You are there to retry or /model switch; long silent waits hide problems |
Script you watch (claude -p in a terminal) | CLAUDE_CODE_MAX_RETRIES=3 | The docs suggest lowering it "to surface failures faster in scripts" |
| PR check in CI with a job timeout | Watchdog on, MAX_WAIT_MS well under the job timeout | Survives a short incident, still fails with a clear 529 instead of a killed job |
| Nightly or weekend batch | Watchdog on, MAX_WAIT_MS of an hour or more, a fallback chain | Waiting is cheap; a failed batch costs a day |
| Eval harness or remote worker | Watchdog on, no cap | The docs name these as the watchdog's intended use |
| Any run at 529s without the watchdog | CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS=8000 or higher | Spreads the same retries over a longer window |
Steps#
-
Check your version. Run
claude --version. The watchdog needs v2.1.186, the base delay v2.1.292 and the wait cap v2.1.295.claude updatemoves a native install forward. -
Turn on the watchdog for the unattended job only. Set it in the job's environment, not in your shell profile:
Terminalexport CLAUDE_CODE_RETRY_WATCHDOG=1 export CLAUDE_CODE_RETRY_WATCHDOG_MAX_WAIT_MS=1800000 # 30 minutes per request claude -p "Run the test suite and fix any failures" --fallback-model sonnet,haiku -
Pick a cap from your job timeout. The cap is per request, and a run makes many requests, so a single bad hour can still burn several caps in a row. Leave headroom between
MAX_WAIT_MSand the CI job's own timeout so you get Claude Code's error instead of a killed runner with no message. -
Add a fallback chain.
--fallback-modeltakes a comma-separated list for one run, andfallbackModelin settings persists it as an array. Chains are capped at three models, the switch lasts one turn, and it also applies to subagents. Rate-limit, billing and auth errors never trigger a switch, so this only helps with overloads and unavailable models. -
Keep the job idempotent. Claude Code's retries, watchdog included, only re-send a request that failed before Claude completed any text or tool call. A failure after Claude has completed text or a tool call is not re-run, because "that could execute the same tool calls twice". Small, checkpointed slices, as in our headless CI comparison, still matter.
Verify It Works#
You do not need an outage to test this. Save a 15-line server that answers every request with a 529, which is how we produced the table above:
# mock529.py - answers every API request with a 529 overloaded_error
import http.server, json
class H(http.server.BaseHTTPRequestHandler):
def do_POST(self):
self.rfile.read(int(self.headers.get("content-length", 0)))
body = json.dumps({"type": "error", "error": {"type": "overloaded_error", "message": "Overloaded"}}).encode()
self.send_response(529)
self.send_header("content-type", "application/json")
self.send_header("content-length", str(len(body)))
self.end_headers()
self.wfile.write(body)
http.server.HTTPServer(("127.0.0.1", 18432), H).serve_forever()
Start it with python3 mock529.py, then in a second terminal point a throwaway claude -p at it. The server prints one log line per request it rejects:
ANTHROPIC_BASE_URL=http://127.0.0.1:18432 ANTHROPIC_API_KEY=sk-ant-test-dummy \
CLAUDE_CODE_RETRY_WATCHDOG=1 CLAUDE_CODE_RETRY_WATCHDOG_MAX_WAIT_MS=30000 \
claude -p "say hi"
With the watchdog on and the cap at 30 seconds, our run ended after 31 seconds and seven rejected requests with API Error: 529 Overloaded and exit code 1. Without the cap it should keep going until you stop it. In an interactive session, the spinner shows a Retrying in Ns · attempt x/y countdown while it waits.
What the watchdog will not wait out#
From the error reference, none of these are rescued by turning the watchdog on:
- Spend limits and exhausted usage credits. A 429 that reports either fails at once. A Claude apps gateway spend cap is marked
x-should-retry: falseand showsspend limit reachedwith its reset time. - TLS certificate failures. Reported on the first attempt so you can fix the proxy or
NODE_EXTRA_CA_CERTSsetup. - Failures after Claude has produced a block of text or a tool call. Claude Code keeps what was completed and continues from finished tool calls rather than replaying them.
- Organisation policy denials. Not retried and not sent to a fallback model.
- Fast mode limits. Fast mode handles its own rate limit by dropping to standard speed until its cooldown ends; see fast mode.
Troubleshooting#
The job sat idle for hours. The watchdog waited out a usage window because the 429 carried a reset time. Add CLAUDE_CODE_RETRY_WATCHDOG_MAX_WAIT_MS, or keep subscription-backed jobs out of time-boxed pipelines.
Request rejected (429) keeps coming back. That is the rate limit on your API key or cloud project, not an overload. The docs suggest running /status to confirm the active credential (a stray ANTHROPIC_API_KEY can route through a low-tier key) and lowering concurrency with CLAUDE_CODE_MAX_TOOL_USE_CONCURRENCY or fewer parallel subagents.
The base delay changed nothing. It is ignored when the watchdog is on, in fast mode, and on versions before v2.1.292. Values outside 500 to 32000, or with anything but plain digits, are treated as unset.
A long stream kept retrying for hours on an old version. v2.1.288 fixed unattended sessions "retrying for hours after a very long response stream failed"; Claude Code now "gives up after three timeouts". Update. If long generations are the problem rather than capacity, our note on long-running request timeouts covers the API side.
FAQ#
What does CLAUDE_CODE_RETRY_WATCHDOG do?#
Set to 1, it makes Claude Code retry 429 and 529 capacity errors indefinitely instead of failing after CLAUDE_CODE_MAX_RETRIES attempts, backing off up to 5 minutes between tries. On v2.1.199 or later it also raises retries for other transient errors to 300, roughly three hours. It still fails at once on spend limits and exhausted usage credits.
Should I just raise CLAUDE_CODE_MAX_RETRIES instead?#
Not past 15 unless the watchdog is on, because that is the cap. Raising retries also keeps the same backoff curve, which in our test covered about three minutes of waiting at the default of 10. For long outages, the watchdog with a MAX_WAIT_MS cap is the documented path.
Does a 529 count against my usage limit?#
No. Claude Code's error reference says a 529 "is not your usage limit and doesn't count against your quota". It means the API is at capacity across users, and capacity is tracked per model, which is why a fallback model such as Haiku 5.5 can keep a run moving.
Is the watchdog safe to leave on everywhere?#
It does not replay tool calls, because Claude Code only re-sends requests that failed before Claude completed any text or tool call, but it is a poor default at your desk: a session that silently waits out a usage window looks hung. Scope it to CI, cron and remote workers, and pair it with a cap.
Sources#
| Source | URL |
|---|---|
| Claude Code error reference | https://code.claude.com/docs/en/errors |
| Claude Code environment variables | https://code.claude.com/docs/en/env-vars |
| Claude Code model configuration | https://code.claude.com/docs/en/model-config |
| Claude Code fast mode | https://code.claude.com/docs/en/fast-mode |
| Claude Code v2.1.295 release notes | https://github.com/anthropics/claude-code/releases/tag/v2.1.295 |
| Claude Code releases (v2.1.288, v2.1.292) | https://github.com/anthropics/claude-code/releases |
| Claude API errors | https://platform.claude.com/docs/en/api/errors |
| Claude Code setup | https://code.claude.com/docs/en/setup |
Continue Reading#
- Claude Outages Are a Workflow Design Problem - the checkpoints and model-switch paths that make waiting worth it
- Claude API Reliability: Error Handling Best Practices - the same retry and fallback ideas for your own API calls
- Headless AI Coding Agents in CI: Claude Code, Codex CLI, Gemini CLI, and opencode Compared - how
claude -pfits into a pipeline in the first place - The Ralph Loop: Running Claude Code For Hours Autonomously - the long-running loop that needs a watchdog most
- Claude Code Usage Limits in 2026: The Practical Playbook for Pro and Max Teams - the usage windows the watchdog will wait out
Get the next deep dive like this in your inbox
One email a week on Claude Code and the rest of the AI dev stack. Free.
Read next on Claude Code
Claude Outages Are a Workflow Design Problem
Claude outages and 529 overloads expose whether your AI coding workflow has checkpoints, receipts, model-switch paths, and small enough task slices to survive provider degradation.
7 min readClaude API Reliability: Error Handling Best Practices
The defensive patterns that keep Claude integrations alive in production. Retry shapes, backoff with jitter, circuit breakers, fallback chains, and the observability you need to debug at 3am.
10 min readHeadless AI Coding Agents in CI: Claude Code, Codex CLI, Gemini CLI, and opencode Compared
A fair comparison of running Claude Code, OpenAI's Codex CLI, Gemini CLI, and opencode in non-interactive CI pipelines: invocation flags, sandboxing, auth, and output formats.
9 min readTechnical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.







