·

Parallel AI Coding Agents: Handle 429 Rate Limits

Prepared and checked using AI. Sources are linked in the article.

best-practices multi-agent cost ai-coding

Parallel AI coding agents hit 429 rate limits when their combined requests exceed a provider’s allowance for requests, tokens, usage, or spending. Several sessions can reach a shared limit even when each agent seems to be working at a reasonable pace. For API-backed agents, OpenAI applies limits at organization and project levels, with separate request and token measures. To handle a 429, identify the exhausted limit first, then slow the sessions that share it. Retry only when the error is temporary.

Find the limit the agents share

A git worktree separates files and changes; it does not create a new provider allowance. Group your running sessions by provider, account or organization, project or workspace, authentication method, and model. That gives you a useful answer to “which other agent could have consumed this capacity?” before you change the workflow.

The providers count traffic differently:

Account limits and model availability change. Use the limit shown for the account and authentication method actually serving the failing session; a published allowance for another plan is a poor concurrency target.

Diagnose the 429 before retrying

Keep the error text or code, the affected model, the account or project, and the time. If your client exposes response headers, record Retry-After and the remaining and reset values too. This is enough to distinguish a short traffic spike from an allowance that will remain exhausted.

A 429 is not always a request-per-minute error. OpenAI’s API error codes distinguish request or token limits from exhausted credits and organization or project usage and spend limits. Pacing helps with the first case; repeated retries do not replenish credits or raise a spend limit. Its API can also enforce limits over shorter intervals than the displayed minute, so starting several agents together can fail despite a modest minute-long average.

Claude API can also return a 429 for a monthly tier spend cap; that response has no retry-after header and can be identified by error.details.error_code set to enforced_spend_limit_reached. Claude API may separately rate-limit a sharp increase in traffic. Read the error details before treating every Claude 429 as an instruction to wait a few seconds.

For Gemini, compare the failing request with the project’s active limits in Google AI Studio. A remaining daily allowance does not rule out a minute-level request or token limit. Likewise, an agent’s context-window indicator describes what fits in that conversation; it does not tell you how much provider quota remains.

Pace sessions against the shared bottleneck

Start with a small active set of sessions using the same provider account. Launch the next session after the earlier ones have passed their initial burst of model calls. If 429s continue, reduce the number of active agents on that account and queue lower-priority tasks. A worktree can remain open while its agent is paused.

Use the error to choose what to reduce:

  • Request limit: Stagger task starts and avoid having every agent ask for a new response at once. Give an interactive task priority over background work when both use the same allowance.
  • Input-token limit: Narrow the files and logs each agent needs. Assign bounded tasks with clear paths instead of asking every agent to inspect the whole repository. For a practical way to define those boundaries, see splitting work for parallel AI agents.
  • Output-token limit: Ask for a focused change or diagnosis rather than a broad report plus an implementation in every session. Review the result before requesting another long response.
  • Daily, subscription, credit, or spend limit: Check its reset time or billing controls. Spacing the same total work across a few minutes will not restore an exhausted daily or monthly allowance.

Treat concurrency as a budget for each shared limiter, not as one machine-wide setting. For example, an agent using a separate provider may have a different bottleneck from two agents using the same API project. The trade-off is straightforward: fewer simultaneous sessions can mean more predictable progress, while a larger active set can spend its allowance sooner. For the spending side of that choice, see the cost of running multiple agents.

Retry temporary errors without a retry storm

When an API reports a temporary rate limit and supplies Retry-After, wait at least that long before sending the request again. If it is absent or invalid, use exponential backoff with a small random delay so agents do not all retry together. Set a maximum number of attempts and an overall time limit; then surface the failure for a person to resolve. OpenAI notes that unsuccessful API requests can count toward minute limits, so an immediate retry loop can prolong the problem.

Check what your client already does. The official OpenAI and Anthropic SDKs retry eligible transient rate-limit errors and honor Retry-After when present. If an agent CLI or wrapper already retries, putting another aggressive retry loop around the whole task can multiply requests. A wrapper should pause further work on the affected account, then resume the failed task once; it should not relaunch every agent at the same time.

If you control Gemini CLI configuration, general.maxAttempts sets the maximum attempts for requests to its main chat model; its documented default is 10. Lowering that setting can make persistent failures surface sooner, but it does not increase quota. Keep quota handling at the provider-account level as well as in each CLI, because independent agents cannot coordinate their retries by themselves.

Reduce the token pressure from each task

Repeated context can make a later turn larger than the first. In Claude Code, /clear starts a fresh conversation, /compact summarizes an ongoing one, /model changes models, and /cost shows current-session usage for API billing. Use /clear when moving to an unrelated task; use /compact when the current task still needs its history. Clearing or compacting affects the session context, not an already exhausted provider allowance.

Give each agent a specific goal, relevant file paths, and a stopping point. Avoid pasting large logs or asking several agents to rediscover the same repository context. When a task is blocked by quota, preserve its current diff and prompt so it can resume after the applicable limit resets. That is more useful than repeatedly restarting the investigation from an empty session.

Where Parallel Code fits

Parallel Code is our free, open-source desktop app for macOS and Linux. It runs several AI coding agents in parallel, each in its own git worktree, and provides a diff-first review surface. Worktree isolation helps you keep their changes separate; pace the agent CLIs you bring according to the limits of their accounts and providers.

Frequently asked questions

Does a new worktree give an agent a separate rate limit?

No. A worktree separates the files an agent changes, while provider limits are tied to the relevant account, organization, project, workspace, or subscription. Check the authentication method for each session before deciding which agents share capacity.

Should I retry every 429 automatically?

No. Retry a temporary request or token limit after Retry-After or a bounded backoff. If the error identifies exhausted credits, a spend cap, or a longer usage allowance, resolve that condition or wait for its stated reset.

Why do agents get 429s when the dashboard shows capacity left?

The dashboard may show a different measure or a longer interval than the one that failed. Check the error’s request or token measure, the model and project in use, and any response headers before increasing concurrency.

Is there one safe number of agents to run at once?

There is no universal count: tasks make different numbers and sizes of requests, and the applicable limits depend on the account and model. Begin with fewer active sessions, stagger their starts, and increase concurrency only while the shared limit has room.