# Stay under your limits: the playbook

Hitting "You've reached your usage limit" at 3pm is the most common Claude complaint. Most of the time it isn't your plan. It's how the conversation is built.

This playbook contains only advice backed by Anthropic's own docs or by how the product works. No "secret prompts that save 90% of tokens".

Last verified: October 2026.

---

## The one idea that explains everything

**Every message re-sends the conversation so far.** Claude doesn't "remember" a chat the way you do. Each time you hit send, the whole thread (your messages, its answers, files, tool results) is processed again as context. Anthropic's Claude Code docs say it directly: a one-line question in a session that's been open all day "still draws usage for the whole conversation."

So the cost of a message is roughly:

> **everything already in the chat + what you add + what Claude writes back (including its thinking)**

That's why message #40 in a long chat costs far more than message #1, even if both are one line. Almost every habit below is a way to keep "everything already in the chat" small.

## How the limits work (quick facts)

- **Plans** (claude.com/pricing): Free; Pro at $20/month (or $17/month billed annually); Max from $100/month with 5x or 20x the usage of Pro. Team and Enterprise for organizations.
- **Two windows on paid plans**: a session limit that resets on a rolling window (about five hours), plus a weekly limit. The message you see when you hit one tells you when it resets.
- **One shared pool**: on Pro and Max, chat on claude.ai/desktop/mobile and Claude Code count against the same limits. A heavy Claude Code morning means less chat in the afternoon.
- **See where you are**: claude.ai → Settings → Usage. In Claude Code, type `/usage` (it also shows what's eating your usage: skills, subagents, MCP servers, long context).
- **Usage credits** (optional, paid): you can turn them on in Settings → Usage to keep working past the limit at extra cost. Off unless you turn them on.

---

## Part 1: Habits for Claude.ai (chat, desktop, mobile)

### 1. One task = one chat
When the topic changes, start a new chat. A chat about your pricing page doesn't need the 30 messages about last week's newsletter. Anthropic's own tip: "Try starting a new conversation if you're approaching your usage limit in a longer chat."

**Carry-over trick**: before leaving a long chat, ask:
```
Summarize this conversation in under 200 words for a fresh chat: the goal, decisions made, the current version of [the thing], and open questions.
```
Paste the summary into a new chat. You keep the context, you drop 90% of the weight.

### 2. Put recurring material in a Project, not in the chat
If you paste the same brand guide, price list or product doc into new chats, stop. Upload it once to a **Project's knowledge**. Anthropic's help center: "Content in projects is cached and counts less against your limits when reused." On paid plans, large Project knowledge switches to retrieval (RAG) automatically, so Claude pulls only the relevant parts.

Also from Anthropic: keep Project instructions "concise and focused on essential information" and "regularly clean up files you're no longer actively using." Every file and instruction is weight.

### 3. Edit, don't follow up
If Claude misunderstood you, **edit your original message** (pencil icon) instead of sending "no, I meant…". A correction message stacks a wrong answer plus your fix on top of the conversation. Editing replaces the branch.

### 4. Batch your asks
"Batch similar requests in one message rather than sending separate ones" (Anthropic). Three questions in one message = one round-trip of context. Three separate messages = three.

### 5. Brief well the first time
Most wasted usage is back-and-forth caused by a vague first message. Use the `prompt-sharpener` skill (or the 7-part check in it) on anything important. A clear 150-word brief beats a 10-word prompt plus 6 corrections.

### 6. Send whole texts once
For writing and editing, "send entire texts at once rather than breaking them into pieces" (Anthropic). Pasting a document in 5 chunks across 5 messages means chunk 1 gets re-processed 5 times.

### 7. Turn off what you don't need for this chat
Anthropic lists these as usage drivers:
- **Extended thinking / higher effort**: great for hard problems, wasteful for "rewrite this email". Lower it or turn it off for routine tasks.
- **Web search**: "ask Claude not to search the web when you don't need current information."
- **Connectors and tools** (Gmail, Drive, etc.): "Tools and connectors are token-intensive." Turn off the ones a chat doesn't need.

### 8. Pick the model for the job
Bigger models use your allowance faster. Use a lighter model (Sonnet or Haiku, where available on your plan) for drafting, summarizing, reformatting and quick questions. Save the top model (Opus) for hard reasoning, strategy and complex code. A simple rule: **if you could give the task to a smart intern, use the lighter model.**

### 9. Use memory and chat search instead of re-explaining
Claude has memory (Settings → Memory) and, on paid plans, can search your past chats. "Use Claude's chat search and memory capabilities to reference previous conversations instead of repeating information." Say "Use what you know from our chat about the Acme proposal" instead of pasting it again.

### 10. Ask for the length you need
Output costs usage too. "Under 150 words", "just the table", "only the changed paragraph" are free savings. Put your default length in your profile instructions once.

---

## Part 2: Habits for Claude Code

All of these come from Anthropic's "Manage costs effectively" page for Claude Code.

| Habit | Command / how | Why |
|---|---|---|
| Clear between unrelated tasks | `/clear` (use `/rename` first so you can `/resume` later) | Stale context is re-sent on every message. `/clear` costs nothing. |
| Compact with instructions | `/compact Focus on the code changes and test results` | Summarizes history but keeps what matters. Note: compacting a huge context is itself a big request; if you don't need continuity, `/clear` is cheaper. |
| Check what's using context | `/context` and `/usage` | Shows the heavy parts (MCP servers, long context, cache misses). |
| Match the model to the task | `/model` | Docs: "Sonnet handles most coding tasks well and costs less than Opus. Reserve Opus for complex architectural decisions or multi-step reasoning." |
| Lower effort for simple work | `/effort` or in `/model` | Thinking tokens count as output. Low effort for renames, small fixes, formatting. |
| Keep CLAUDE.md short | Under ~200 lines | It loads into every session. Move procedures into skills (they load only when used). |
| Disable unused MCP servers | `/mcp` | Prefer CLI tools (`gh`, `aws`, etc.) when available: they add no tool listing. |
| Plan before big changes | Shift+Tab to plan mode | Avoids expensive re-work when the first direction is wrong. |
| Stop early | Esc to stop, `/rewind` or double-Esc to go back | Don't let a wrong approach run for 10 minutes. |
| Be specific | "add validation to the login function in auth.ts" | Vague asks ("improve this codebase") trigger broad file scanning. |
| Send noisy work to subagents | "Use a subagent to run the tests and report only failures" | Verbose output stays in the subagent; only the summary comes back. |
| Don't leave long sessions idle then resume | Start fresh or accept "resume from summary" | After the cache lifetime (1 hour on a subscription), the next message re-processes the full context. |
| Watch scheduled loops | `/usage` shows heavy `/loop` tasks | A recurring task re-sends your context every time it fires. |

---

## Part 3: The 60-second pre-flight (print this)

Before a big task, ask yourself:

1. **Is this a new topic?** → New chat (or `/clear`).
2. **Am I about to paste something I've pasted before?** → Put it in Project knowledge or CLAUDE.md / a skill.
3. **Is my ask complete?** → Goal, audience, format, length. Batch the related questions.
4. **Does this need the biggest model and deep thinking?** → If an intern could do it, no.
5. **Does this chat need web search or connectors?** → If not, off.

## What will NOT help (skip these)

- **"Write shorter prompts" as a blanket rule.** A clear long brief saves usage by avoiding corrections. Cut the chat history, not the brief.
- **Magic prompt phrases** claiming to "compress" Claude's thinking. There's no evidence they reduce your plan usage in a meaningful way.
- **Splitting a document across messages "so it fits".** It makes it worse (see habit 6).
- **Opening many parallel chats on the same task.** Each one rebuilds context from zero.

## When you really are at the limit

1. Check Settings → Usage for the reset time.
2. Read the message. If it's a model-specific limit (e.g., "You've hit your Opus limit"), you can keep working on a different model. If it's a session or weekly limit, it applies to all models, so switching won't help.
3. Do the thinking offline: write the brief for your next session so it starts sharp.
4. If it happens every week and Claude is making you money, compare the cost of Max with the hours you lose. That's a business decision, not a moral failing.

Sources: Anthropic Help Center ("Usage limit best practices", "How do usage and length limits work", "Using Claude Code with your Pro or Max plan", "Projects"), Claude Code docs ("Manage costs effectively"), claude.com/pricing. Checked October 2026. Limits and features change; if something here doesn't match your app, trust the app.
