Anthropic's Claude API limit story shifted in June 2026: the current public docs organize API usage into three tiers named Start, Build, and Scale. The same docs define separate ceilings for requests per minute, input tokens per minute, and output tokens per minute.

Radar takeaway: higher API limits reduce interruptions, but they do not solve the moment a coding session has already gone stale, hit a wall, or lost the useful thread of work.

What changed.

The practical change is simpler tiering and more capacity headroom for API users. Anthropic's docs now describe a three-tier API structure and make clear that rate limits apply across request count and token volume. When capacity is exceeded, the API can return a 429 response, and clients should use retry timing rather than hammering the endpoint.

Limit layer What Anthropic documents Why coders should care
Usage tier Start, Build, and Scale tiers. Capacity now maps more cleanly to workload size, but your workflow still needs a recovery path.
Request rate Requests per minute limits. Agentic coding can fan out many calls during planning, edits, tests, and retries.
Token rate Input-token-per-minute and output-token-per-minute limits. Large diffs, logs, and repo context can burn through token budgets even when request count looks fine.
Over-limit behavior 429 responses and retry timing. Your app can wait and retry; your human coding flow still needs the current task state preserved.

Higher limits are good. They are not recovery.

More Sonnet and Haiku capacity is good news for teams running parallel requests, code agents, evaluation jobs, and automated workflows. It means fewer interruptions at the infrastructure layer.

But the developer pain ShardStitch tracks is not only API throughput. It is the moment the useful working context is trapped inside a dying chat, a stale agent loop, a rate-limited session, or a tool you need to leave.

The workflow implication.

If you are building against the Claude API, use the new tier structure to size your workload. If you are coding with Claude, Claude Code, Cursor, Codex, or another AI tool, keep a separate recovery habit: preserve git state, changed files, recent commits, notes, and what the agent was trying to do.

That is why local recovery still matters even when platform limits improve. Higher ceilings delay the wall. They do not guarantee that the session will be clean when you finally hit it.

What to do in practice.

  1. Watch usage and rate-limit signals, but do not depend on them as your only safety net.
  2. Keep important planning notes on disk, not only inside the chat.
  3. Use git diff and changed files as the source of truth when a session becomes unreliable.
  4. When the session is stuck, rebuild from disk instead of asking the stuck model to explain itself.
ShardStitch angle: API limits are platform capacity. Session recovery is local continuity. They help different parts of the same workflow.

Sources.