codebeaver: Review Contract Spec (v1) β
The behavioral contract for the reviewer, distilled from docs/prompt-engineering-lessons.md. Everything normative lives in code or prompt files; this doc explains the design and where each decision came from.
Where things live β
| Artifact | Path |
|---|---|
| Finding schema (machine-readable source of truth) | packages/prompts/schema/finding.schema.json |
| Tier/depth dial (machine-readable) | packages/prompts/schema/tiers.json |
| TS types | packages/prompts/src/schema.ts, src/depth.ts |
| Prompt skeleton (mode-agnostic constitution) | packages/prompts/prompts/reviewer-core.md |
| Tier overlays | packages/prompts/prompts/reviewer-{low,medium,high}.md |
| Noise-control conventions (living assets) | packages/prompts/conventions/{exclusions,precedents,never-comment}.md |
| Golden eval set | eval/fixtures/ |
| Eval runner | eval/runner/ |
The finding schema β
One shape for every finding, mirroring the Anthropic ReportFindings + Copilot review schema patterns:
file+lineβpath:lineanchor, 1-indexed, new-file side. Never whole-file citations.summaryβ the claim only, one sentence, β€200 chars.failure_scenarioβ concrete inputs/state β wrong output or crash. This is the load-bearing field: a finding without a nameable failure scenario does not qualify (cleanup findings state the concrete cost instead).evidenceβ how the reviewer verified it: the guard checked and found absent, the caller traced, the upstream guarantee ruled out. Makes false positives auditable.categoryβ fixed slug enum for routing and grading.verdictβCONFIRMED(trigger constructible from the code) orPLAUSIBLE(real mechanism, rare/uncertain trigger). A three-state vote (with REFUTED) exists between finder and verifier roles; the public output only ever carries the two surviving states.- No severity labels. Severity is an ordering with a deterministic tie-break: correctness outranks cleanup when the cap forces a cut.
The depth dial β
One skeleton + one config knob, per the OpenAI # Juice lesson (reasoning effort is a single injected line over a frozen prompt β diff-able and testable):
- low β single pass, hunk-only scope, no verifier, cap 4. Precision-stanced.
- medium (default) β enclosing functions in scope, self-verify against the CONFIRMED/PLAUSIBLE definitions, cap 8. Precision stance.
- high β recall stance, finder/verifier split (8 angles Γ up to 6 candidates β per-candidate skeptic vote β sweep), cap 10. Activation policy: built into the dial, but only enabled in production after evals show a recall gap that precision mode cannot close. Complexity is bought with eval evidence.
A tier never changes prose elsewhere in the prompt. It changes: scope, verify bias, cap, and (at high) the orchestration protocol.
Precision/recall stance is implemented in the verifier's refutation bar β
Not in vague confidence words. Medium: refute anything without an actionable failure scenario. High: PLAUSIBLE-by-default, REFUTED only when constructible from the code (quote the line), and a whitelist of realistic states (races, rare-but-reachable error paths, missing optional fields, falsy-zero, off-by-one on unexcluded boundaries) that may never be dismissed as "speculative."
Noise control (the $20-bill bar) β
Three living convention files are inlined into the prompt at assembly time:
never-comment.mdβ categories that never qualify (style, naming, missing docs, hypothetical perf, praise).exclusions.mdβ specific finding patterns auto-excluded (theoretical DoS, log spoofing on non-rendered logs, trusted internal inputsβ¦), seeded from Anthropic's security-review HARD EXCLUSIONS.precedents.mdβ numbered standing rulings (P-01β¦) that settle recurring debates, seeded from security-review's PRECEDENTS. Owner process: a precedent is added/changed only by a PR to this file; the reviewer must not re-raise covered issues.
Uncertainty rule (strict tiers): if you're unsure whether something is a problem, do not mention it.
Empty state β
[] with no padding is a first-class, correct result. In human-readable mode: "No findings." + any residual risks or testing gaps, nothing else. No compliments about the code.
Read-only constitution β
The reviewer never modifies the repo. Allowed: read, search, run tests/builds that don't edit tracked files. Forbidden: edit, patch, format, commit. Tie-breaker: if the action "does the work" rather than "reviews the work," don't do it. The mode is not escapable by user phrasing β a request to "just fix it while you're there" is a request to hand the fix description to the caller, not to apply it.
Comment style β
The anti-slop block (docs/anti-slop-block.md) is inlined in reviewer-core.md as the style section for human-readable comment text: lead with the finding, no praise openers, no hedging fillers, no em dashes, name the fix directly when a fix section is requested.
Eval-first iteration rule β
No prompt change merges without an eval run (eval/runner/run.mjs). New bug classes enter as a fixture first (with expected findings and distractor non-findings), then the prompt change that catches them. New recurring false-positive debates enter as precedent rulings, not as prompt prose.