Test fixtures are the data your tests use. Handled well: tests read cleanly, code is refactor-friendly, setup is minimal. Handled poorly: fixture files sprawl into thousands of lines, every schema change breaks hundreds of tests, and reading a test file is 40 lines of setup for 1 line of behavior. This guide is the fixture discipline that worked on a 5-year-old codebase with 1,800 tests, refined and encoded into the subagent that now produces our new fixtures.
Why fixtures matter more than they seem
Fixtures are code you'll touch far more than production code. Every test file has fixture setup. Every refactor of a data model breaks fixture files across the codebase. Every new engineer learns the codebase partly through reading tests, which means partly through reading fixtures.
Bad fixtures make the codebase feel harder than it is. New engineers see 40-line test setups and conclude the domain must be very complex. Refactoring becomes expensive because every schema change requires manual fixture updates. Test files become read-only files that engineers avoid.
Good fixtures make the codebase feel simpler than it is. Tests read as intent: “when a locked user tries to log in, expect X.” Refactoring is fast because factories propagate schema changes automatically. Test files are living documentation.
The difference is not code volume; it's discipline.
The four fixture strategies
Every fixture strategy is one of four archetypes:
1. Factory functions. buildUser({ overrides }). Function returns an instance with sensible defaults plus caller's overrides. Flexible; composable; minimal setup. Preferred for most teams.
2. Builder pattern. UserBuilder().withEmail("x@y.com").withStatus("locked").build(). Chained calls; explicit; verbose. Common in Java, Kotlin; less so in JavaScript / Python.
3. Static fixture files. fixtures/users/admin.json. Files on disk; loaded via helper. Readable for large data; sprawls fast; brittle to schema changes.
4. Inline construction. const user = { id: 1, name: "Alice", ... }. Object literals in tests. Works for one-off cases; sprawls into copy-paste hell.
Real codebases mix these; usually one primary strategy plus others for specific cases. Consistency within a codebase matters more than optimal choice.
When to use each
Rough decision tree:
Factory functions if you're greenfield or refactoring and can pick one strategy. Best default for JavaScript/TypeScript, Python, Ruby.
Builder pattern if working in Java/Kotlin idiomatically, or if the domain has genuinely chained construction (e.g., queries, DSLs).
Static fixture files if you have large domain data that stays stable (product catalog, country list, geographic data). Not for test data that varies per test.
Inline construction for genuinely one-off test data that doesn't recur. If it recurs even twice, promote to factory.
The failure mode is mixing without discipline. “Sometimes we use factories; sometimes we inline; sometimes we load from JSON.” No consistency; harder to reason about; harder to change.
The factory pattern in depth
Since factories are our recommended default, worth going deep.
Factory function signature: buildX(overrides: Partial<X> = {}): X. Returns a valid X with sensible defaults for anything not overridden.
Defaults should be minimal-and-valid. Not every field populated; just enough that the returned instance is valid. Test-specific fields override.
Overrides use partial types. TypeScript's Partial<X>; Python's **kwargs; Ruby's attributes = {}. Overrides express what varies from the default.
Realistic data via Faker. Names, emails, dates, addresses from Faker. Not "test1" or "user@test.com". Realistic data catches issues that unrealistic data hides (like assumed email format).
Named traits for common variations. Instead of long override objects, name the pattern: buildUser({ trait: "locked" }). Traits encode intent; overrides encode variation.
Example progression:
Bad: const user = { id: 1, name: "Test User", email: "test@test.com", status: "locked", createdAt: new Date(), lastLoginAt: null, failedAttempts: 5, ...(15 more fields) }
Good: const user = buildUser({ trait: "locked" })
Both produce equivalent test data. Second is readable; encodes intent; propagates through schema changes.
Composition over inheritance
Some fixture libraries encourage inheritance (subclass factories). Some libraries encourage composition (compose smaller factories). Composition is more flexible and less brittle.
Example: an Order factory needs a User. Two approaches:
Inheritance: class UserOrderFactory extends OrderFactory. Rigid; adds a class for each combination.
Composition: buildOrder({ user: buildUser({ email: "x@y.com" }) }). Compose at call site; no new class needed. More flexible.
Composition scales better. Composition allows tests to express exactly the graph they need without predefined class hierarchies.
Handling relations and object graphs
Real domains have relations. User has Orders; Order has Items; Item references Product. How to handle in fixtures?
Rule 1: Factories compose factories. Order factory takes an optional user or creates one via buildUser(). Item factory takes an optional product or creates one. Composition ripples through.
Rule 2: Bidirectional relations set on one side only. If user.orders and order.user reference each other, set on one side (say, order.user) and derive the other. Prevents mismatch between sides.
Rule 3: Deep graphs are usually smell. If your test needs a user with 3 orders with 5 items each with a product with 2 variants, either the test is exercising too much or the domain is over-modeled. Simplify the test or the domain.
Rule 4: DB tests use factories + explicit persistence. Factory constructs; test explicitly persists. Persistence is separate from construction; factories should work for in-memory tests too.
Deterministic vs. random
Faker generates random data by default. Random has trade-offs.
Random by default. Catches order-dependence (different data each run breaks tests that depend on specific values). Prevents hardcoded-value assumptions.
Seed when specific values matter. Test asserts email format: seed Faker so the email is deterministic. Faker.seed(123) in beforeEach; deterministic within the test.
Explicit values when semantically meaningful. If the test cares about “locked user,” the status is explicit; the name can be random. Explicit overrides for meaningful fields; defaults for incidentals.
The failure mode is going all-deterministic (never catches assumption bugs) or all-random (occasionally breaks specific-value tests). Mixed is right; deliberate about which is which.
Fixture sprawl and how to prevent it
Sprawl is the primary fixture failure mode. Signs of sprawl:
Fixture files >500 lines. Multiple fixture files per module. Tests importing from many fixture locations. Same fixture pattern copy-pasted across tests. Fixture updates that require touching many files.
Prevention:
Extract to factory when a pattern appears twice. First use inline; second use = extract. Refactor at extraction time; don't accumulate.
Named traits for common combinations. Reduces override sprawl. buildUser({ trait: "admin_with_2fa" }) reads better than 8-line override object.
Delete fixtures for retired features. When a feature dies, fixture data for it should too. Dead fixtures accumulate.
Quarterly fixture audit. Read the fixture files; identify sprawl; refactor. Cheap maintenance; prevents runaway.
Sprawl compounds silently. A team without fixture discipline discovers this at year 3 when fixture maintenance becomes a bottleneck.
Fixtures for integration tests
Integration tests differ. Use real DB; use real HTTP; use fewer mocks. Fixtures shift accordingly.
Same factories. The same buildUser factory works for both unit and integration. Consistency across test types.
Explicit persistence. Integration test: const user = await persist(buildUser());. Factory constructs; test persists. Persistence layer separate.
Transactional isolation. Each test in a transaction; rollback in afterEach. No fixture cleanup needed; DB state consistent across tests.
Seed data for shared domain. If tests need product catalog etc., seed once per test run; not per test. Different fixture strategy for shared vs. per-test data.
What to do next
If your fixtures are mostly inline: pick a factory library for your language. Refactor one module's fixtures to factory pattern. Measure: test file line counts drop.
If your fixtures are factory-based but sprawling: audit one module. Extract traits for common overrides. Delete dead fixtures. Should drop file sizes by 30-50%.
If your fixtures are well-disciplined: install test-fixture-designer. New fixtures produced with same discipline; discipline compounds without human enforcement.
Fixtures are code you'll touch often. Investing in fixture discipline pays back weekly. The investment is small; the return compounds.