Generic test writing discipline: test quality, real assertions, anti-patterns, and rationalization resistance. Use when writing tests, adding test coverage, or fixing failing tests for any language or framework. Complements language-specific skills.
--- name: ia-writing-tests class: discipline description: >- Generic test writing discipline: test quality, real assertions, anti-patterns, and rationalization resistance. Use when writing tests, adding test coverage, or fixing failing tests for any language or framework. Complements language-specific skills. --- # Writing tests Produce tests that prove the requested behavior and fail when that behavior breaks. Follow user scope and the repository's actual contracts; this skill does not authorize implementation, external actions, or changes to acceptance criteria. ## Procedure 1. Discover the repository's runner, pinned wrapper, configuration, neighboring tests, and CI command before writing tests. Learn both focused and full-suite commands. Use project tooling rather than a global default; keep tracked scripts portable and apply personal wrappers only to the outer invocation. 2. Derive cases from user requirements, implemented behavior, and claims intended for the handoff. Map each acceptance criterion to a discriminating case that a plausible wrong implementation fails. Name tests by observable behavior; keep one behavior per test and make fixtures adversarial on the axis under test. 3. For bug fixes, write the reproducer, observe the intended failure, apply the smallest fix, and observe green. For new features, follow the selected test posture in vertical slices: default to tests-after (one small implementation slice, then its test); with an explicit test-first choice, write one test, then the minimal implementation that passes it. Repeat by slice; writing all tests first and all implementation afterward (horizontal slicing) is an anti-pattern that produces tests for imagined behavior. Refactor after green. Never alter a specification, assertion, fixture, snapshot, or expected output merely to make the implementation pass. 4. Prefer real internal objects, real temporary files, and real test databases. Mock only external boundaries, at the last owned adapter; framework-maintained fakes are appropriate where the framework recommends them. Check the real contract and side effects before mocking. 5. Assert consumer-visible outcomes rather than private structure, mock calls, or framework behavior. Include relevant boundary, invalid-input, concurrency, and failure cases. If error handling catches, logs, substitutes, or rolls back, assert its observable result as well as error visibility. 6. Prove absence/isolation assertions with a run-unique forbidden violation and observe that specific assertion fail. Confirm the control mutation actually landed and its build completed before evaluating it. Remove the control, restore the artifact, and verify green. 7. Run the narrow checks during edits and the applicable complete checks before handoff. Inspect passed/executed counts; an empty or skipped suite is not positive evidence. ## Select the detail needed - When choosing cases, fixture shape, mock boundaries, or unit/integration/E2E balance, read [test-design.md](./references/test-design.md). - For bug reproducers, mutation controls, or assertions that something must not happen, read [regression-proof.md](./references/regression-proof.md). - When reviewing generated tests, changing snapshots, testing async timing, or seeing mocks substitute for behavior, read [test-smells.md](./references/test-smells.md) and the applicable fix ladder in [anti-patterns-extended.md](./references/anti-patterns-extended.md). To audit whether existing tests detect regressions, use ia-test-audit. - For weak aggregates, vacuous assertions, race reproduction, or controls that may never reach their subject, read [oracle-smells.md](./references/oracle-smells.md). It routes to the longer catalogs when those failure shapes apply. - For container state, sandboxing, timeouts, environment relocation, or process/stream isolation, read [isolation-and-sandbox-traps.md](./references/isolation-and-sandbox-traps.md). - For wrappers, filters, conformance oracles, skip conditions, or misleading green summaries, read [false-pass-oracle-traps.md](./references/false-pass-oracle-traps.md). - When byte equivalence or a generated corpus is the contract, read [generated-corpus-techniques.md](./references/generated-corpus-techniques.md); compare against an independent implementation across value shapes and nesting. - When stuck, considering skipping tests, or assembling a test-work handoff, read [test-completion.md](./references/test-completion.md). Before arguing against a needed test, read [rationalization-table.md](./references/rationalization-table.md). ## Verify and report Verify public behavior and edge paths, independence between tests, reviewed expectation changes, and each claimed capability. Mentally remove a guard, flip a branch, or drop a side effect and confirm a test detects it. Prefer a unit suite fast enough for frequent use (under 30 seconds). Report what the tests exercised, actual results, failures, and omissions. Distinguish fixtures from live proof. Use the applicable PHP/Laravel or React/TypeScript skill for framework conventions; this generic skill does not replace them.
don't have the plugin yet? install it then click "run inline in claude" again.
this skill teaches test-writing discipline across any language or framework. tests prove behavior works; a test that can't fail is worthless, and a test that mocks the system under test is theater. use this skill when writing new tests, adding coverage to existing code, fixing failing tests, or reviewing test quality. pair it with language-specific skills (ia-php-laravel, ia-react-frontend, etc.) for framework conventions and tooling.
inputs: repository root, CI pipeline file, readme outputs: confirmed test command, runner invocation, test file locations
discover what the repository actually runs, not what seems default.
inputs: user requirements doc or issue, code diff, existing API/data contracts outputs: enumerated test cases mapped to each source with no coverage gaps
build coverage depth and breadth by sourcing from three independent streams, then verify every item maps to at least one test.
inputs: test case from step 2 outputs: test name and a one-line assertion statement
test names are first-level documentation. a reader should know what happens without opening the test body.
inputs: feature being tested, available test infrastructure (test DB, temp dirs, framework utilities) outputs: decision to use real object or mock, with rationale
mocks are a last resort, not first choice. every mock is an assumption about behavior that may drift from reality.
| use real objects for | use mocks/fakes for |
|---|---|
| database queries (use ephemeral test DB) | external HTTP APIs |
| internal services and classes | payment gateways |
| file system operations (use temp dirs) | email/SMS delivery |
| business logic and transformations | third-party SDKs with rate limits |
| queues, caches, time-based logic (use framework test doubles) | hard-to-reach error states |
exception: framework-provided test doubles (Laravel Queue::fake(), React test providers, vi.mock for API layers) are idiomatic and maintained alongside the framework -- use them. this rule targets hand-rolled mocks that drift.
if mocking is unavoidable, generate the mock from the real type/schema and include all fields consumed downstream, not just the ones the test author knows about. incomplete mocks hide bugs until production.
inputs: test name, decision on real/mock, test setup outputs: complete test with explicit assertion
each test verifies exactly one behavior.
inputs: test case, multiple similar tests outputs: test code with shared setup extracted, no hidden test intent
duplication in tests is acceptable, even desirable, when it makes intent obvious. extract shared setup only when it reduces noise without hiding what the test does.
inputs: feature being tested, error paths and boundary conditions outputs: test cases for empty/null, boundaries, errors, and error observability
for every feature, enumerate and test these categories:
edge cases: empty input, null/undefined, boundary values (0, 1, max, max+1), invalid types (string where number expected), concurrent access (if applicable), unicode and special characters, permission denied, timeout, network failure
silent failure coverage: hunt error paths where failures can be swallowed. for each one, add an assertion that proves the failure was observable:
use specific assertions: instead of expect(result).toBe(null) (passes for both "handled gracefully" and "silent drop"), prefer expect(logger.error).toHaveBeenCalledWith(expect.any(DatabaseError)) -- make the observable signal explicit.
inputs: test written (or test + code together for new features) outputs: passing test, clean refactor (if applicable)
for bug fixes (prove-it pattern):
for new features: write tests alongside implementation, not after. by the time the feature is done, tests exist and pass. when making a test pass, write the simplest code that satisfies it (not the abstraction that seems "right", not the feature that might be next). refactor only after the test is green.
inputs: all tests written outputs: checklist confirming coverage, naming, real objects, edge cases, speed, independence
before considering tests complete, check:
inputs: all tests, especially any written with LLM assistance outputs: corrected tests with smells removed
before committing any test (including self-written), scan for these six smells:
if you don't know the test runner for this project: discover it before writing tests (step 1). a globally installed binary often resolves to a different version than the project pins. always use the project-local wrapper if one exists.
if the test command in CI differs from the README: follow CI. CI is authoritative because it's what gates code.
if you're unsure whether to test an internal implementation detail or a user-visible outcome: test the outcome. tests that verify framework behavior (the ORM saves records, the router routes) are waste -- trust the framework. test the project's own logic (business rules, transformations, decisions).
if you're tempted to write a test-only method in production code (reset(), clearState(), setTestMode()): stop. if tests need to reset state, the code has a design problem. refactor to make state explicit and injectable.
if you must mock a method but don't know what the real method returns: check the API docs or type definition, then populate the mock with all fields consumed downstream, not just the ones you know about. incomplete mocks fail silently in production.
if a snapshot test is the only test for a feature: add behavioral assertions alongside snapshots. snapshots catch unintended changes but don't verify correctness.
if integration tests fail with row-count multipliers (expected 2 rows, got 8) yet pass on a fresh container: persistent infrastructure kept state from prior runs. reset infrastructure state between runs (ephemeral containers, fixture TRUNCATE, volume teardown) instead of relying on tests cleaning up after themselves.
if tests pass individually but fail together: use bisection to find the polluter. run tests one-by-one in isolation until you find the test with shared mutable state.
if a test is too complicated to write: simplify the interface being tested. if test setup is large, extract helpers that reduce noise without hiding intent. if you must mock everything, code is too coupled -- use dependency injection.
if you're about to skip, defer, or argue against writing a test: load the rationalization table (references/rationalization-table.md) first. thirteen common excuses with counter-truths. if you're still arguing, the argument is probably lost.
success looks like:
file locations:
you know the skill worked when:
reaching for a default test command: the bare global runner passes locally while CI invokes the project wrapper and fails. discover the runner, wrapper, and CI command first.
testing mock behavior instead of real behavior: test passes but production breaks. mocks assert correctly but actual system fails. replace mocks with real objects for internal code.
test-only methods in production code: methods like reset(), clearState(), setTestMode() exist only because tests need them. refactor to make state explicit and injectable.
snapshot tests as the only test: all tests are snapshots that get bulk-updated whenever anything changes. snapshots catch unintended changes but don't verify correctness. add behavioral assertions.
testing the framework: tests verify the ORM saves records, the router routes, or the framework does what its docs say. trust the framework. test the project's own logic.
incomplete mocks: mock includes only fields the test author knows about. downstream code consumes other fields and gets undefined. mock the complete data structure as it exists in reality.
mocking without understanding: before mocking, ask (1) what side effects does the real method have? (2) does this test depend on any of those side effects? (3) mock at the lowest level that removes the slow/external part.
persistent test infrastructure state contamination: integration tests fail with row-count multipliers (expected 2 rows, got 8) yet pass on a fresh container. reset infrastructure state between runs (ephemeral containers, TRUNCATE, volume teardown). never rely on tests cleaning up after themselves.
vacuous forall over an empty collection: forall-style assertion (every, all, .iter().all()) passes vacuously -- the factory never attached children. attach a realistic child set and confirm the predicate flips for at least one populated case.
constructing the object-under-test below the transform layer: the fix lives in an upstream transform (parser, normalizer, from_api_response), but the test builds the object via the leaf constructor with the already-correct value. feed the test the raw pre-transform input so the transform under test executes.
synchronous adapters hiding timing-dependent races: parallel requests through a zero-latency mock settle in the same microtask, so a dedup guard passes. under real wire latency, staggered arrivals miss the window. inject controllable latency; assert the guard holds for arrival-staggered bursts.
asserting only presence, never absence: payload tests assert expected fields exist but never that unexpected fields are absent. pin absence as well as presence: assert "proof_document_id" not in payload.
this skill covers generic test discipline. for framework-specific patterns, conventions, and tooling: