Grok returned 429 Too Many Requests?

xAI enforces per-model request and token limits. Exceeding either returns 429. The safe response is paced retry with a cap, not an uncontrolled loop that repeats an expensive coding request.

Preserve first: Capture the Grok model, endpoint, HTTP 429 response, and request time. Stop synchronized retries, inspect any local agent changes, and only retry after the request budget has room.

Match the failure before retrying.

Requests per second exceeded

xAI derives a per-second cap from the request budget, so a short burst can fail even when the minute looks lightly used.

Tokens per minute exceeded

Prompt, completion, reasoning, and cached prompt tokens count toward the model limit.

The request partially changed local files

If Grok is being used through an agent, inspect the repository before retrying because tool work may have completed before the model error surfaced.

Safe recovery path.

Record the model, 429 response, request timing, and any agent tool output.
Capture git status, git diff, and the last verification result before another attempt.
Reduce request rate and use capped exponential backoff as xAI recommends.
Retry one bounded action or continue from the same packet in another supported AI.

What the next AI should receive.

xAI API endpoint and model, response status/body and request ID, request/token usage if available, retry timing, agent actions completed locally, current diff, and the retry condition supported by the response.

How ShardStitch works with Grok.

ShardStitch preserves the project state independently of the Grok request. For Grok web it produces a clipboard packet; for agent workflows it preserves the verified change set and marks any old-model claim for re-checking.

Source trail.

The linked xAI Rate Limits and Debugging Errors pages describe model-specific request/token limits and error handling. Check the current endpoint and model values in those documents; a 429 alone does not show whether requests-per-second or tokens-per-minute was exceeded.

FAQ.

Does an xAI 429 tell me whether request rate or token use was exceeded?

Not by itself. Check the response details and current model/endpoint limits. A short burst and a large token load can require different pacing changes.

Should an agent repeat the same code task after Grok returns 429?

First inspect the worktree. Local tool calls may have changed files before the model response failed. Continue from the actual diff instead of replaying the whole task.

Is a fixed retry delay safe for every xAI model?

No. Follow the current response and documentation, cap retries, and avoid synchronized bursts. This page does not prescribe a universal delay.

Scope note: This guide addresses HTTP 429 responses from the xAI API. Limits are model- and account-dependent and may change; do not apply API guidance to a different Grok surface without checking its own documentation.