Claude Code cost at team scale is a solvable problem, but the solutions are not the ones most people reach for first. This guide walks through six levers, in the order that produces the biggest returns for the least effort. Uses our reference 10-person team as the working example throughout.
The real cost drivers
For most teams, Claude Code cost is dominated by three factors:
1. Model mix. Opus costs 5x Sonnet at input, 5x Sonnet at output. Every prompt sent to Opus that could have gone to Sonnet is a 5x overspend on that prompt.
2. Input tokens (before caching). Input volume dominates. CLAUDE.md is prepended to every prompt; large CLAUDE.md means every prompt is expensive.
3. Session length. Long sessions accumulate context; each subsequent prompt has larger input than the last. Caching helps but doesn't fully eliminate this.
Other factors matter (output volume, cache miss rate, per-service MCP overhead) but these three dominate. Optimize them first.
Measurement first
Before optimizing, measure. Two data sources:
Anthropic Console usage dashboard. Shows total usage by day, model, and organization. Aggregate view; not per-developer or per-session.
Per-session tracking. Claude Code exposes session-level cost via ccusage or similar tools. Enable per-developer tracking so you can identify which developers or workflows are cost-heavy.
Baseline for a 10-person team: expect $2000-4000/month API spend if no optimization is in place. Well-optimized team: $1000-1800/month. Under $1000 typically means either the team is barely using Claude Code, or there's heavy Claude Max usage.
Lever 1: Model selection
Biggest lever, highest ROI. Default to Sonnet; use Opus specifically.
Sonnet handles 80% of coding work well. Producer work (writing tests, implementing well-specified features, refactoring within established patterns, generating documentation) is Sonnet territory. Opus earns its price on:
- Genuinely complex reasoning (architectural decisions, hard debugging).
- Multi-file refactors requiring holistic understanding.
- Reasoning across ambiguous requirements.
Configure per-subagent model. postgres-dba: Sonnet. security-reviewer: Opus (reasoning-heavy). e2e-test-writer: Sonnet. architect-reviewer: Opus. unit-test-writer: Sonnet.
Typical mix that works: 20% Opus / 70% Sonnet / 10% Haiku. If your mix is 60% Opus, you're likely overspending.
Lever 2: Prompt caching
Second biggest lever. Cache reads cost 10% of normal input pricing.
What caches well:
- CLAUDE.md (prepended to every prompt in a session).
- Codebase context (files read early in a session, referenced later).
- Documentation the developer includes as context.
Cache hit rate depends on session structure. Long sessions with stable context: high hit rate (60%+). Short sessions or sessions where context churns: low hit rate (10-30%).
To improve hit rate:
- Longer, more focused sessions vs. many short scattered ones.
- Include reference documentation upfront; don't lazy-load it mid-session.
- Structure CLAUDE.md for stability (avoid frequent edits).
Impact: raising cache hit rate from 30% to 60% cuts input token cost roughly in half.
Lever 3: CLAUDE.md discipline
CLAUDE.md is prepended to every prompt in every session. A 20KB CLAUDE.md means every prompt starts with 20KB of context. For a team of 10 sending 500 prompts/day/dev, that's 100MB of daily context that could have been smaller.
Keep CLAUDE.md focused. What belongs:
- Team conventions Claude should follow (naming, style, patterns).
- Non-obvious constraints (specific frameworks, target versions, legacy considerations).
- Discipline reminders (test discipline, security discipline).
What doesn't belong:
- Documentation Claude can find in the codebase (README, docs/ folder).
- Long tutorials or examples (Claude can read these on demand).
- Boilerplate that repeats other configuration.
Target CLAUDE.md size: 3-8KB. If yours is 30KB, aggressive editing produces immediate cost improvement.
Lever 4: Session scope
Long focused sessions cost less per outcome than many short sessions.
Short-session pattern: each prompt starts fresh; CLAUDE.md loaded fresh; no cached context from previous prompts. Every prompt has full input cost.
Long-session pattern: initial prompts load context (paid once); subsequent prompts benefit from caching (paid at 10%). Total cost per outcome is lower.
Encourage: focused work sessions of 30-60 minutes. Discourage: micro-sessions of one prompt each.
Lever 5: Hooks as budget guards
Hooks can enforce cost discipline. Two specific hooks:
Session-cost alarm. Notifies developer when session cost exceeds threshold ($1, $5, $10 — choose based on team norms). Awareness alone cuts spend; developers who see their session at $5 pause and evaluate before continuing.
Opus usage lint. Warns when a task that could be handled by Sonnet is being sent to Opus. Not blocking — sometimes Opus is needed — but requires acknowledgment. Prevents accidental Opus for routine work.
See session-cost-alarm and opus-usage-lint.
Lever 6: Workflow optimization
Some workflows are inherently expensive; some are inherently cheap. Small changes to workflow can produce large cost changes.
Expensive workflows:
- Iterative "make it work" prompting (small tweaks over many prompts).
- Exploration without specific goal ("look around and tell me what's interesting").
- Restart-heavy workflows (frequent new sessions).
Cheap workflows:
- Specific, well-defined prompts.
- Task-focused sessions with clear endpoint.
- Batch operations that share context across multiple tasks.
Cultural shift: encourage developers to write specific prompts rather than exploratory ones. "Fix the bug" is expensive; "the calculateTotal function returns wrong result when cart is empty — fix" is cheap.
The 10-person team model
Our reference team: 10 developers, moderate Claude Code use, all six levers applied.
Baseline (no optimization): ~$2800/month API spend.
After lever 1 (model mix): $2200/month. -20%.
After lever 2 (caching to 50% hit rate): $1600/month. -30%.
After lever 3 (CLAUDE.md discipline): $1400/month. -10%.
After lever 4 (session scope): $1250/month. -10%.
After levers 5+6 (hooks and workflow): $1100/month. -10%.
Total: from $2800 to $1100. About 60% reduction with 4-6 hours of team effort per lever plus ongoing discipline.
When Claude Max makes sense
Claude Max is Anthropic's flat-rate plan: $100/mo (5x usage) or $200/mo (20x usage). At team scale:
10-person team @ optimized API: $1100/mo = $110/dev/mo. Compare:
- Claude Max 5x: $100/dev/mo → $1000/mo for 10 devs. Cheaper.
- Claude Max 20x: $200/dev/mo → $2000/mo for 10 devs. Nearly 2x cost.
For a well-optimized team at this scale, Claude Max 5x is a real competitor. Considerations:
- Max has usage limits (5x or 20x the Pro tier); heavy users may hit them.
- Max is flat rate; API is usage-based. Different budgeting model.
- Max includes chat access; API doesn't (or requires separate Pro/Team).
For teams under-optimized (spending $3k+/month), API optimization is the right first step. Consider Max only after applying the six levers.
What to do next
- Measure current spend (Console + per-session).
- Set target: 40-60% reduction.
- Apply lever 1 (model mix) — biggest single lever.
- Apply lever 2 (caching) — second biggest.
- Apply lever 3 (CLAUDE.md) — often reveals surprising bloat.
- Apply levers 4-6 in parallel over the following weeks.
- Measure again after 4 weeks; iterate.
Cost optimization at scale isn't a one-time project; it's ongoing discipline. Set quarterly cost reviews and treat the levers as a checklist. The teams that keep spend flat while usage grows are the ones that revisit this deliberately.