Enforces fresh verification evidence before any completion claim. Use when about to claim "tests pass", "bug fixed", "done", "ready to merge", handing off work, or before editing when a request has ambiguous scope.
--- name: ia-verification-before-completion class: discipline description: >- Enforces fresh verification evidence before any completion claim. Use when about to claim "tests pass", "bug fixed", "done", "ready to merge", handing off work, or before editing when a request has ambiguous scope. --- # Verification before completion Make completion claims only from fresh evidence for the actual claim. Follow the user's authorized scope; repository instructions supply applicable checks, not permission to mutate, publish, or weaken a requirement. ## Procedure 1. Resolve ambiguous scope before editing. Inspect the repository and state a safe assumption when one interpretation is clear. Ask only when materially different interpretations remain; do not edit the disputed scope while waiting. 2. Inspect `git status --porcelain` and preserve unrelated work. Before any suite or migration runs, including a baseline, confirm its database and service targets are disposable (a dedicated test database, container, or throwaway schema; never the dev database) ([proof-integrity.md](./references/proof-integrity.md), Pre-Verification Check). For dependency/framework upgrades, codegen, or migrations, capture the existing validation command set before writing and rerun it unchanged afterward. If that baseline is red, report before proceeding. Shared-module verification on a dirty tree needs an isolated base comparison. 3. **Identify** the command that proves the claim. For ship-level claims, check the full applicable chain: build, types, lint, tests, security scan, and diff review; stop on the first failure. Read project-declared gates and run the ones that apply to this action in their required order; do not invent gates. 4. **Run** the proof now. Earlier output, a subagent's report, confidence, and a renamed success phrase do not replace fresh execution. 5. **Read** complete output and exit status, including warnings, executed/passed counts, and missing artifacts. A suite that executes nothing is not proof. Confirm the intended binary, interpreter, source revision, and entry point actually ran. 6. **Verify** that evidence covers the requirements and relevant failure paths. An implemented safe positive capability must work through its intended entry point; a refusal-only path, stub, mock, or unreachable implementation is partial. 7. **Claim** only what the evidence establishes. Report the outcome, exercise command/URL/click path, failed or skipped checks, and material residual risks. State narrower proof scope and distinguish deterministic fixtures from live behavior. Never make an oracle easier to satisfy to obtain green. Review and justify semantic changes before regenerating expected output. Never hard-code the exercised subject or success path. A clean review is valid when it covers the relevant criteria; broaden checks only for a named remaining risk. ## Route by verification risk - For dirty worktrees, broad changes, strict input validation, or fixture-versus-live provenance, read [proof-integrity.md](./references/proof-integrity.md). For shared modules with unrelated edits, also read [isolated-verification.md](./references/isolated-verification.md). - For repository-wide sweeps or “every item” claims, read [scope-and-sweeps.md](./references/scope-and-sweeps.md). Enumerate every item in untracked/ignored scratch state, preserve explicit dispositions, re-enumerate after moves, and account for removals. Completion requires zero pending and zero blocked items. - For command wrappers, empty results, aggregate totals, installed binaries, or materialized revisions, read [verification-oracles.md](./references/verification-oracles.md). Require positive controls through the same invocation shape before interpreting an absence. - For frontend, backend, CLI, infrastructure, migration, package, schema, documentation, or scripted-sweep changes, read the matching row of [change-strategies.md](./references/change-strategies.md). It also covers adversarial probes, history rewrites, and stale reviews. - For integration boundaries, callbacks, or orphaned state, read [system-wide-test-check.md](./references/system-wide-test-check.md). - When classifying deliverables, handling failed checks, discussing branch scope, or encountering a pre-commit failure, read [claims-and-failures.md](./references/claims-and-failures.md). ## Failure and handoff rules Do not retry unchanged verification until it happens to pass. Fix authorized implementation failures and rerun; otherwise name the concrete blocker. “Pre-existing,” “environmental,” and “flaky” require evidence against the deliberately chosen base or independently established cause. Do not bypass a pre-commit failure caused by this work. The documented exception requires a reproduced base-branch failure and prior visibility to the user; this skill supplies no new bypass authority. Verify delegated work directly through the diff and relevant command. Check specification compliance separately from quality. Re-read requirements line by line: passing tests and meeting requirements are different claims. Refresh facts and coordinates when new commits or external state could invalidate a prior review. Keep reports decision-relevant, without empty status sections or fabricated certainty.
don't have the plugin yet? install it then click "run inline in claude" again.
extracted implicit decision logic into explicit if-else branches, added input contracts and edge cases, restructured procedure as numbered steps with clear input/output per step, formalized completion report structure, and added outcome signals with concrete verification markers.
never claim completion, handoff, or readiness without fresh verification evidence. run the proof command immediately before the claim, in the current session, and read the full output. "should pass", "looks correct", and "pretty sure it works" do not count. only command output with exit code 0 or explicit pass confirmation counts as evidence. use this skill whenever you are about to say "tests pass", "bug fixed", "done", "ready to merge", handing off to another agent, or before editing a request with ambiguous scope.
git status --porcelain)inputs: current working tree state
steps:
a. run git status --porcelain. check for uncommitted changes unrelated to the current task.
b. if dirty tree found: commit, stash, or explicitly acknowledge the unrelated changes before proceeding.
c. if the change touches a shared module or has broad blast radius (dependency bump, framework upgrade, codegen, migration), capture the baseline first: run the repo's full validation suite against the existing state before any write. record the exact command set and its output.
d. for delegated work: never trust the subagent's report. confirm via git diff that changes were actually made, then run the verification command directly yourself.
outputs: clean or documented working tree; baseline command set and output (if broad-radius change)
inputs: the original request phrasing, current file inventory applies when: the request uses spatial scope like "migrate my project", "refactor the codebase", "fix this across the app", "update everywhere" AND the user has not enumerated specific files. steps: a. run a breakdown command to surface the blast radius:
rg -l 'pattern' | cut -d/ -f1 | sort | uniq -c | sort -rn # files per top-level dir
rg -l 'pattern' | xargs dirname | sort -u # affected directories
b. present the result: "This touches N files across M subsystems".
c. ask the user (via AskUserQuestion in Claude Code or request_user_input in Codex) for scope confirmation: (a) everything, (b) just
skip this gate if: the request already lists specific file paths.
outputs: confirmed scope from user, or explicit scope bounds
inputs: the full set of items to cover (all files matching a pattern, all findings, all renames, etc.)
applies when: the task scope is "every item in a set" (repo-wide rename, "migrate everywhere", audit all files, resolve all findings).
steps:
a. create a ledger in session-scratch (git-ignored local directory, never tracked) with one row per item: pending, done, excluded (reason), or blocked (evidence).
b. after any path move, rename, or new item creation during the sweep, re-enumerate the set so new items enter coverage.
c. keep removed items in the ledger until explicitly accounted for: if an item silently vanishes, it looks the same as "finished" unless you mark it removed.
d. completion requires zero pending and zero blocked entries. "i covered a lot of them" is not a disposition.
outputs: ledger file with all items enumerated and zero pending/blocked entries
inputs: the claim you are about to make, the proof command from the Common Claims table below (or the command that proves it) steps:
| step | action | example |
|---|---|---|
| 1. identify | what command proves this claim. for ship-level claims (commit/push/PR-ready), run the full chain: build -> typecheck -> lint -> test -> security scan -> diff review, stopping on first failure. for a single claim, use the proof command from the Common Claims table. | pytest tests/, npm test, curl -s localhost:3000/health |
| 2. run | run it now, in this same message. output from an earlier turn is stale. | "i ran it earlier" fails this step |
| 3. read | read the complete output. check the exit code. do not scan for "passed" - read failure counts, warnings, errors. | scan all output, not just the summary line |
| 4. verify | does the output actually confirm the claim. | "42 passed, 0 failed" confirms tests pass. "41 passed, 1 failed" does not. |
| 5. claim | only now make the statement, with evidence visible. | "all 42 tests pass" with output visible above |
outputs: verification command output; explicit pass/fail determination
inputs: the type of change made match the strategy to the change type:
| change type | required verification |
|---|---|
| frontend (component, page, form) | start the dev server, exercise the feature in a browser, check the console. test happy path and one failure path. screenshot or describe the result. |
| backend handler / endpoint | curl the endpoint, check response shape and status code. hit at least one error path (invalid input, missing auth). |
| cli tool | run the binary with real inputs. check stdout, stderr, exit code. run from /tmp to catch "only works from source" bugs. |
| infra / iac (terraform, dockerfile, k8s) | terraform plan / docker build / kubectl apply --dry-run=server. review the diff before applying. |
| database migration | run migration up, down, then up again against production-shape data. |
| refactoring (no behavior change) | full test suite passes unchanged. public api surface diff shows no breakage. run grep on exported identifiers before and after. |
| library / package update | run the consumer's test suite against the new version. check for deprecation warnings. |
| schema change | old consumers parse the new shape (forward compat). new consumers handle old data still present (backward compat). test both directions. |
| documentation / prose | read the rendered output. confirm links, formatting, and content match intent. |
| config with no validator | validate syntax where possible (jq ., yamllint). otherwise read the file and confirm it matches the intended change. |
| non-runnable changes | git diff, confirm the diff matches intent. state explicitly: "no automated verification available , verified by reading the diff." |
if no row applies: fall back to the non-runnable row. reading code is not a strategy.
outputs: verification command output and result for the change type
inputs: the change (production logic or data-sensitive code) applies to: any change touching production logic exempt: docs, trivial typos, pure rename refactors pick at least one probe:
null, undefined, MAX_INT, 1-char unicode combining markoutputs: probe command, execution, and result
inputs: prior review(s) (agent or human), current HEAD
applies before shipping:
a. run git log --oneline <review-commit>..HEAD to list commits after the last review.
b. if new commits exist, verify the new changes do not invalidate prior conclusions: previously flagged issues still fixed, no new code contradicts the review.
c. if stale, request a new review or confirm the changes are safe.
outputs: staleness check result; decision to re-review or proceed
inputs: verification command output showing failure steps: a. do not claim completion. report the actual failure output to the user. b. do not retry hoping for a different result. c. return to implementation. fix the issue, then re-run from step 1 of the gate function. d. if failure is unrelated to current changes (pre-existing flaky test, environment issue): state it explicitly with evidence (show the failure also occurs on the base branch, or link to a known issue tracker).
outputs: explicit failure report with root cause or evidence of pre-existing issue
inputs: pre-commit hook rejection decision: is this failure caused by the current session's changes?
git commit --no-verify is forbidden. the hook is a verification checkpoint, not an obstacle.--no-verify permitted.outputs: either fixed root cause or documented pre-existing failure with evidence
a complete verification produces a structured report in this format:
## Completion report
**Status**: DONE / DONE_WITH_CONCERNS / BLOCKED / NEEDS_CONTEXT
**Changes made**
- path/to/file.ts: [one-line description of what changed and why]
- path/to/other.ts: [one-line description]
**Next**
- [the one command to paste, URL to open, or click-path that exercises this in one step -- or the decision now owed]
**Things I didn't touch (intentionally)**
- [thing noticed but out of scope, with one-line reason]
- [adjacent issue deferred, with one-line reason]
- [or: nothing noticed]
**Potential concerns**
- [any risk, uncertainty, or open question the reviewer should know about]
- [or: none]
**Verification evidence**
- [command]: [exit code / result summary]
- [command]: [exit code / result summary]
status codes: DONE means verification passed, zero concerns. DONE_WITH_CONCERNS means verification passed but risks exist (mandatory to explain in concerns section). BLOCKED means work cannot proceed (name the blocker). NEEDS_CONTEXT means you need info from the user (name the missing info).
changes made: list files actually touched and one line per file on what changed and why. prove scope was bounded.
next: one line only. the thing the reader does now (usually not the verification command itself; usually the command to exercise the feature, the URL to open, the decision needed, or the next file to edit). when result is not exercisable (docs, config, refactor), name what was read or checked.
things i didn't touch: mandatory section. if nothing was noticed, write "nothing noticed". goal is to prove scope was considered.
potential concerns: mandatory if DONE_WITH_CONCERNS. explain any risk, uncertainty, or open question. if DONE with no concerns, write "none".
verification evidence: list each command run, its exit code, and a one-line result summary. include enough detail that the reader can reproduce the verification.
the skill has worked when:
--no-verify without evidence of pre-existence.