How to Design Slash Commands That Scale — Claude Code | AI Code Toolkit
📚 Workflow · 28 min read · Ahmed R.

How to Design Slash Commands That Scale

Slash commands look easy: write a markdown file, drop it in .claude/commands/, done. That version works for one developer for a week. This guide is about the version that works for a team of ten for a year — the design decisions that make commands survive contact with real usage.

By Ahmed R.
28 min read
Updated August 2026
Guide v1.0

What this guide covers

This is a guide about command design — the discipline of building slash commands that stay useful as your team and codebase grow. It's not a tutorial on the basics (the Slash Commands library has per-command tutorials for 195 reference commands); it's the meta-guide that explains why some commands survive and others get abandoned.

The failure mode this guide is written against: teams that start with 3-4 useful slash commands, gradually accumulate 20-30 more that are half-useful, and end up in a state where nobody remembers what any command does or when to use it. The reference library exists to catch you before that happens. This guide is what to know when you start building your own.

Concretely, this guide answers: “How do I design a command that will still be useful in six months? That my teammates will actually invoke rather than working around? That won't accumulate ambiguity as the codebase changes?” If you've written a slash command and want to make it robust — or you're about to write one and want to get it right the first time — this is that guide.

The three shapes commands take

Every well-designed slash command fits one of three shapes. Recognizing which shape you're building tells you what discipline applies.

Shape 1: Invocation commands

An invocation command invokes a specific operation on a specific target. /test-unit <file> runs unit tests on a file. /rename-symbol <old> <new> renames a symbol. /deploy-prod deploys to production. These commands have clear inputs, clear outputs, and clear success/failure criteria. They're the most common shape by count.

Design discipline for invocation commands: narrow scope, explicit argument-hint, deterministic behavior for the same inputs. If invoking the command twice with the same inputs produces materially different outputs, the command is under-specified. Add rules until the behavior is predictable.

Shape 2: Workflow commands

A workflow command orchestrates multiple steps in service of a larger goal. /commit is a workflow: read diff, generate message, lint the message, run pre-commit hooks, commit. /pr-create is a workflow: create PR, fill description, apply template, assign reviewers, post to Slack. Workflow commands are fewer in count but higher in value.

Design discipline for workflow commands: document the steps explicitly in the command file. Not just “this creates a PR” but a numbered list of the six things it does in order. When something fails, you and your teammates need to be able to open the command file and see exactly which step failed. Workflow commands that don't document their steps become mysterious when they break.

Shape 3: Gate commands

A gate command provides the legitimate path through a hook-blocked capability. /deploy-prod is not just a workflow — it's also the gate command for the deploy capability that prod-deploy-block otherwise blocks. /force-push is the gate command for force-pushes that force-push-block otherwise blocks. Gate commands are rare but disproportionately important because they encode the “yes, do this thing that's normally too dangerous” safety pattern.

Design discipline for gate commands: always set the sentinel variable explicitly, always prompt for confirmation before the gated operation, always log what was gated (so the audit trail exists). The whole point of a gate command is that it turns a dangerous operation into a safe-when-intended one; the confirmation step is what carries the safety.

Which shape is your command? If you can't answer this without hesitation, the command is probably trying to be more than one shape at once — which is exactly where scalability breaks down. Split it into shape-specific commands and each becomes robust.

Anatomy of a scalable command

Every command has five parts in its frontmatter and body. Each part carries specific design weight.

allowed-tools: narrower than you think

The allowed-tools field restricts what tools the command can invoke. Default instinct is to grant broad access: Read, Write, Edit, Bash. This works but scales badly — a command with broad tool access can drift over time to do things it wasn't designed for, and the drift is invisible until something breaks.

Better: grant the narrowest set that accomplishes the command's specific job. A refactoring command needs Read, Grep, Glob, Edit, Bash(pnpm typecheck:*). A test-running command needs Read, Bash(pnpm test:*). A review command needs Read, Grep, Glob (no write access at all — reviewers analyze; separate tools act).

The discipline: when writing a new command, start with the minimum tools you think it needs, run it, expand only when you hit a tool restriction that blocks a legitimate operation. Never expand past what the command's specific job requires.

description: written for the invocation surface

The description field appears in your slash-command autocomplete. It's the first (and often only) thing your teammates read to decide whether to invoke the command. Write it for that purpose.

Good: “Rename a symbol across all references; verify with typecheck; roll back on failure” — names what happens, names the verification, names the failure mode.

Bad: “Renames symbols” — leaves the reader guessing about scope, safety, and behavior.

The description should be under 100 characters (the autocomplete truncates), imperative in mood (matches how the command is invoked), and specific enough that a teammate reading it knows whether to use this command or a different one.

model: default sonnet, opus for reasoning

The model field selects which Claude model runs the command. Default to Sonnet for anything mechanical (running tests, applying edits, formatting output). Choose Opus for anything requiring judgment (review, analysis, threat modeling, migration planning).

The cost delta is small; the correctness delta on judgment-heavy work is large. If you're uncertain whether a command needs Opus, run it once with each and compare the output on a representative input. When Sonnet gets it right consistently, use Sonnet. When Sonnet misses things Opus catches, use Opus.

argument-hint: enables discoverability

The argument-hint field shows in the autocomplete alongside the description. Use it to name what the command expects. “[file path]” for a single-file command. “[old name] [new name] [optional: scope]” for a multi-arg command. “[URL]” for a URL-taking command.

Commands without argument-hint are invocable only by people who remember what they take — which is fine for solo work and fails as soon as your teammate tries to use the command you built. Argument-hint is small effort with large payoff for team distribution.

Task section: numbered, verifiable, rollback-aware

The Task section (the body of the command file) is the actual instructions Claude follows. Design principles:

  • Number the steps. Prose paragraphs are ambiguous; numbered lists are unambiguous.
  • Include verification. Every command that modifies code should verify (typecheck, test, or reference search) that the modification produced the expected result.
  • Include rollback. If verification fails, the command should undo its changes rather than leave a partial state.
  • Encode edge cases as rules. If the command has cases it shouldn't handle (never touch generated files; never modify test fixtures; never rename symbols in node_modules), list them explicitly.

The reference commands in the Slash Commands library all follow this pattern — look at any of them for a working example.

Naming and namespacing

Command names have to be memorable, consistent, and unambiguous within your team's command set. This gets harder as the command count grows.

Consistency conventions

Pick a naming convention and stick to it across all commands:

  • Verb-first, singular — /commit, /review, /deploy. Standard convention; works well for most commands.
  • Verb-object hyphenated — /test-unit, /review-security, /deploy-prod. Standard for commands with a specific object.
  • Prefix-based grouping — /db-migrate, /db-seed, /db-rollback. Groups related commands lexically; autocomplete becomes navigable.

Mixing conventions produces autocomplete lists that feel arbitrary. Pick one, apply it consistently, and your teammates will find commands by intuition after a week.

Namespacing for many commands

Claude Code supports namespacing via subdirectories: .claude/commands/db/migrate.md becomes /db:migrate (not /db-migrate). This is worth adopting once you exceed ~15 commands because it keeps the autocomplete organized.

Common namespaces: db/, test/, review/, deploy/, docs/, refactor/. Each namespace holds 3-8 related commands. Anything more granular becomes overhead; anything less granular defeats the point.

Reserving names

Some names are worth reserving even if you don't have the command yet. /help, /init, /reset, and similar short generic names should either be intentionally-used or intentionally-empty. Accidentally shadowing a common name with a specific command produces bad ergonomics.

Testing your commands

Commands are code that runs on your codebase. They deserve testing. Not the same kind of testing as your application code, but testing nonetheless.

Smoke testing on install

Every time you add or modify a command, invoke it on a known-safe target and verify the output. For a review command, run it on a file you already know has issues and confirm the issues get flagged. For a workflow command, run through the workflow on a test branch and confirm each step completes. Five minutes of smoke testing catches ~80% of the bugs that would otherwise surface at inconvenient moments.

Regression testing on modification

When you change an existing command (add a step, tighten a rule, change a tool), re-run the smoke test. If you don't, you'll discover the regression next time you invoke the command for real — which is much worse feedback than catching it immediately.

Team-level testing via fixtures

For commands your whole team relies on, consider committing test fixtures. A test-fixtures/ directory with sample files that represent common inputs your commands run against. Then a team convention: “when you modify /rename-symbol, run it against the fixtures in test-fixtures/refactoring/ and confirm the output matches the expected diff.” This is lightweight but catches drift across command changes over time.

Versioning and migration

Commands evolve. What was /deploy six months ago may need to become /deploy-staging and /deploy-prod as your setup matures. How you handle the transition determines whether teammates notice or fight the change.

Deprecation, not deletion

Never delete a command that teammates use. When you split /deploy into /deploy-staging and /deploy-prod, keep the old /deploy around for at least one iteration cycle — but modify it to print a deprecation warning and redirect the invoker:

NOTE: /deploy is deprecated. Use /deploy-staging for staging or /deploy-prod for production. Aborting to prevent ambiguous deploy.

Deprecation gives teammates time to migrate their habits. Deletion just breaks their workflow with no warning.

Breaking changes are announcements

If a command's behavior changes materially (different arguments, different behavior on the same arguments, different output format), that's a breaking change and worth announcing. Slack channel, PR description, or CLAUDE.md diff — whichever your team notices. Silent behavior changes to commands break trust; announced changes preserve it.

When to split, when to merge

The command count in your .claude/commands/ directory is a moving target. Sometimes commands should merge; sometimes they should split.

When to split

  • The command has cases that need different tool sets — if /deploy needs one set of tools for staging and a different set for prod (plus the prod-deploy-block hook), split into /deploy-staging and /deploy-prod.
  • The command has cases that need different models — if /review uses Sonnet for style review but needs Opus for security review, split into /review-style (Sonnet) and /review-security (Opus).
  • The description no longer fits under 100 characters — if you can't describe the command concisely, it's doing too many things.

When to merge

  • Two commands have the same tool set, model, and structure — and differ only in one hardcoded parameter. Add an argument and merge.
  • Neither command is invoked often — and the merged version would have clearer intent. Merge and reduce cognitive load.
  • The commands are always invoked in sequence — merge into a workflow command that runs both.

The split-vs-merge decision is judgment; there's no universal rule. But the failure modes point in different directions. Too many commands = teammates can't find the one they need. Too few commands = each command is doing too much. Aim for the middle by paying attention to the failure modes and adjusting.

Sharing across a team

Individual command files in .claude/commands/ committed to git is the base case. For teams above ~10 engineers, additional structure helps.

Ownership per command

Add an owner comment to each command file: <!-- owner: platform-team -->. When someone hits a bug or needs a change, they know who to ping. Ownership per command is more important than ownership per directory because commands cross feature boundaries.

Contribution conventions

Document how a teammate proposes a new command or modifies an existing one. Options:

  • PR-based — treat command changes like code changes. Review, discuss, merge. Highest friction; highest quality.
  • Notify-then-commit — commit changes directly but post a heads-up in the team channel. Lower friction; some coordination loss.
  • Anyone-goes — no formal process. Works for small teams; breaks down above ~5 engineers.

Pick one and be consistent. Mixing approaches (some commands need PR review, others don't) is where inconsistency accumulates.

Cross-repo sharing

For organizations with many repos that need similar commands, consider a shared-commands repo pulled in as a submodule or synced via a script. The .claude/commands/ in each repo becomes a mix of shared and repo-specific commands. Naming convention distinguishes: prefix shared commands (/org:review) vs. repo-local (/review). This scales further than per-repo command sets but requires more coordination.

Anti-patterns

Patterns worth avoiding, drawn from watching teams fail with commands over time.

The kitchen-sink command

A command called /do or /help-me that tries to figure out what the user wants based on context. Inevitably ambiguous, inevitably wrong at 20% of invocations, inevitably abandoned. Commands should have narrow, specific purposes; general-purpose helpers are just Claude Code without a command.

The undocumented workflow

A workflow command that doesn't document its steps in the file. When it breaks, nobody knows which step broke. Fix: add the numbered steps to the command file, even if it feels verbose. The verbosity pays off the first time something goes wrong.

The tool-unrestricted command

A command with allowed-tools: [all] or equivalent. Works today; drifts over months as Claude uses the granted tools in ways the command's original author didn't anticipate. Fix: narrow the tools to what the command specifically needs. Re-narrow annually as commands evolve.

The silent-fail command

A command that doesn't verify its output and doesn't report errors. Runs, produces something, teammates trust the output, mistakes propagate. Fix: every code-modifying command runs a verification step (typecheck, test, or reference search); every command reports failures explicitly rather than silently proceeding.

The name-collision command

A command whose name shadows a Claude Code builtin, an installed subagent name, or a hook script. Ambiguous invocation. Fix: check your naming against the built-in list and installed items before naming a new command. Prefix or namespace when in doubt.

What comes next

This guide covered command design. The related meta-guides:

And when you want to see many examples of well-designed commands, the Slash Commands library has 195 reference commands you can read and adapt. Every one of them follows the discipline this guide describes.

Next in the series →

Building Your Hook Layer: Security + Safety-Gates

Hooks are what commands sit on top of. The meta-guide on architecting the hook layer that makes commands (and everything else) safe.

Debugging Claude Code errors? See our sister site AI Error Hub for common error messages and fixes.
Visit AI Error Hub →

Frequently asked questions

Answers to the questions readers ask about this guide.

Fewer than you think. For most teams:

  • Solo developer — 5-10 commands.
  • Small team (2-5) — 10-20 commands.
  • Larger team (5-15) — 20-40 commands with namespacing.
  • Team-of-teams (15+) — 30-60 commands with strong namespacing and ownership.

If you're over 60 commands, look for merge candidates — you probably have several near-duplicates. If you're under 5, you're leaving value on the table — workflow commands like /commit and /pr-create pay for themselves the first week.

Copy first, adapt second. Reference commands encode discipline (verification, rollback, tool restrictions) that took thousands of invocations to get right. Starting from a reference command means starting with that discipline in place.

Adapt to your team's conventions after installing:

  • Tool aliases (your pnpm test may be yarn test).
  • Test command names (your test:fast may be test:quick).
  • Team-specific rules (your review commands may need to skip certain paths).

Writing from scratch is the last resort — reserved for commands genuinely unique to your team's workflow.

Reference environment variables in the command; never hardcode credentials. Commands are committed to git; hardcoded credentials are one bad commit from a compromise.

Pattern:

  1. Command references ${API_KEY}.
  2. Each teammate sets the env var in their shell or gitignored .env.local.
  3. Command's allowed-tools lets Bash read the env var but not print it.

This matches how MCP tokens are handled — same discipline for the same reason.

Not directly, but the effect is achievable. Two patterns:

  • Compose in prose — the command's Task section says “first run the equivalent of /commit, then run the equivalent of /pr-create”. Claude follows the composition.
  • Extract shared logic — if two commands share a lot of logic, extract into a shared file both reference. Less DRY but more explicit.

Direct command-invokes-command isn't supported and probably shouldn't be — the coupling would make debugging harder.

Two patterns:

  1. Dry-run mode — design the command with a --dry-run flag that reports what it would do without doing it. Test the dry-run mode; trust the real mode by extension.
  2. Sandbox environment — test the command against a throwaway repo, database, or cloud environment that can be reset if something goes wrong.

Never smoke-test destructive commands against production data. The five minutes saved isn't worth the incident risk.

Start with Sonnet. If the command produces good output consistently, keep Sonnet — you save cost with no correctness loss.

Switch to Opus if any of the following:

  • The command requires judgment (review, analysis, threat modeling, migration planning).
  • Sonnet output is inconsistent across runs on the same input.
  • Failure of the command is expensive (security review that misses a vulnerability; refactoring that breaks call sites).

Cost delta is small; correctness delta on judgment-heavy work is large.

Add arguments to the command; don't create a second nearly-identical command. Example: /deploy-prod with an optional --dry-run flag is one command; /deploy-prod and /deploy-prod-dry-run as separate commands is two commands with 90% duplication.

When arguments proliferate to more than 3-4, that's the signal to split. Under 3-4 arguments, single command with args is cleaner.

Common and useful pattern. Command's allowed-tools includes the specific MCP tools it needs: allowed-tools: Read, Grep, mcp__github__create_pull_request. The command orchestrates around the MCP call — setting up context before, handling errors after.

Discipline: never grant broad MCP access (mcp__github__*) when specific access (mcp__github__create_pull_request) accomplishes the job. The narrower grant scales better.

Share with