Skip to content

codebeaver: Review Contract Spec (v1) ​

The behavioral contract for the reviewer, distilled from docs/prompt-engineering-lessons.md. Everything normative lives in code or prompt files; this doc explains the design and where each decision came from.

Where things live ​

ArtifactPath
Finding schema (machine-readable source of truth)packages/prompts/schema/finding.schema.json
Tier/depth dial (machine-readable)packages/prompts/schema/tiers.json
TS typespackages/prompts/src/schema.ts, src/depth.ts
Prompt skeleton (mode-agnostic constitution)packages/prompts/prompts/reviewer-core.md
Tier overlayspackages/prompts/prompts/reviewer-{low,medium,high}.md
Noise-control conventions (living assets)packages/prompts/conventions/{exclusions,precedents,never-comment}.md
Golden eval seteval/fixtures/
Eval runnereval/runner/

The finding schema ​

One shape for every finding, mirroring the Anthropic ReportFindings + Copilot review schema patterns:

  • file + line β€” path:line anchor, 1-indexed, new-file side. Never whole-file citations.
  • summary β€” the claim only, one sentence, ≀200 chars.
  • failure_scenario β€” concrete inputs/state β†’ wrong output or crash. This is the load-bearing field: a finding without a nameable failure scenario does not qualify (cleanup findings state the concrete cost instead).
  • evidence β€” how the reviewer verified it: the guard checked and found absent, the caller traced, the upstream guarantee ruled out. Makes false positives auditable.
  • category β€” fixed slug enum for routing and grading.
  • verdict β€” CONFIRMED (trigger constructible from the code) or PLAUSIBLE (real mechanism, rare/uncertain trigger). A three-state vote (with REFUTED) exists between finder and verifier roles; the public output only ever carries the two surviving states.
  • No severity labels. Severity is an ordering with a deterministic tie-break: correctness outranks cleanup when the cap forces a cut.

The depth dial ​

One skeleton + one config knob, per the OpenAI # Juice lesson (reasoning effort is a single injected line over a frozen prompt β€” diff-able and testable):

  • low β€” single pass, hunk-only scope, no verifier, cap 4. Precision-stanced.
  • medium (default) β€” enclosing functions in scope, self-verify against the CONFIRMED/PLAUSIBLE definitions, cap 8. Precision stance.
  • high β€” recall stance, finder/verifier split (8 angles Γ— up to 6 candidates β†’ per-candidate skeptic vote β†’ sweep), cap 10. Activation policy: built into the dial, but only enabled in production after evals show a recall gap that precision mode cannot close. Complexity is bought with eval evidence.

A tier never changes prose elsewhere in the prompt. It changes: scope, verify bias, cap, and (at high) the orchestration protocol.

Precision/recall stance is implemented in the verifier's refutation bar ​

Not in vague confidence words. Medium: refute anything without an actionable failure scenario. High: PLAUSIBLE-by-default, REFUTED only when constructible from the code (quote the line), and a whitelist of realistic states (races, rare-but-reachable error paths, missing optional fields, falsy-zero, off-by-one on unexcluded boundaries) that may never be dismissed as "speculative."

Noise control (the $20-bill bar) ​

Three living convention files are inlined into the prompt at assembly time:

  • never-comment.md β€” categories that never qualify (style, naming, missing docs, hypothetical perf, praise).
  • exclusions.md β€” specific finding patterns auto-excluded (theoretical DoS, log spoofing on non-rendered logs, trusted internal inputs…), seeded from Anthropic's security-review HARD EXCLUSIONS.
  • precedents.md β€” numbered standing rulings (P-01…) that settle recurring debates, seeded from security-review's PRECEDENTS. Owner process: a precedent is added/changed only by a PR to this file; the reviewer must not re-raise covered issues.

Uncertainty rule (strict tiers): if you're unsure whether something is a problem, do not mention it.

Empty state ​

[] with no padding is a first-class, correct result. In human-readable mode: "No findings." + any residual risks or testing gaps, nothing else. No compliments about the code.

Read-only constitution ​

The reviewer never modifies the repo. Allowed: read, search, run tests/builds that don't edit tracked files. Forbidden: edit, patch, format, commit. Tie-breaker: if the action "does the work" rather than "reviews the work," don't do it. The mode is not escapable by user phrasing β€” a request to "just fix it while you're there" is a request to hand the fix description to the caller, not to apply it.

Comment style ​

The anti-slop block (docs/anti-slop-block.md) is inlined in reviewer-core.md as the style section for human-readable comment text: lead with the finding, no praise openers, no hedging fillers, no em dashes, name the fix directly when a fix section is requested.

Eval-first iteration rule ​

No prompt change merges without an eval run (eval/runner/run.mjs). New bug classes enter as a fixture first (with expected findings and distractor non-findings), then the prompt change that catches them. New recurring false-positive debates enter as precedent rulings, not as prompt prose.

Source-available under BUSL-1.1. Self-hosting is free for your organisation.