Skip to content

Review Methodology ​

This document is the single authored source for how CodeBeaver reviews: the pass structure, the evidence standards, the convergence doctrine, and the rules for changing any of it. src/prompts.ts implements this methodology; when this document and the code disagree, one of them is a bug.

Methodology changes follow the versioning rule at the end of this document.

Pass structure ​

A review runs in layers; each layer contributes evidence, never a second lifecycle:

  1. Deterministic layer — SAST, prechecks, structural impact, test impact, custom pre-merge checks, and the fix-evidence audit. These are pure code: reproducible, model-free, and cheap. They run before any AI pass and their output is untrusted-data-sanitized like everything else.
  2. Primary AI pass — one pass over the head-SHA diff with bounded, head-scoped context. Focused passes (deleted behavior, contract compatibility, security boundaries, async failure semantics, changed tests) are derived and attached before primary review; they are part of the primary pass, not separate review lifecycles. The finding stage is coverage-first: report every issue found, including uncertain and low-severity ones, each with a confidence level; filtering and ranking happen downstream, not here. The publication bar is author-relevance: report an issue the author would likely fix if made aware of it.
  3. Validator pass — a three-lane filter over candidate findings: validated, downgraded, or dropped, each ID classified exactly once. It receives candidates-with-hunks and bounded source/context sections, including fetched callers and consumers, never the primary pass's reasoning. Small sections remain complete; larger sections share the remaining 12,000 character context budget with explicit middle-truncation markers. Missing or clipped source is an evidence limit, not proof of absence. A warning or critical finding with an unproven element but a real mechanism is downgraded to suggestion severity rather than dropped; a finding is dropped only when the validator can refute it with the code. It fails closed: a validator outage suppresses unvalidated findings, demotes the verdict to comment, and marks the review incomplete.
  4. Posting — inline comments (diff-anchored, applicability-checked), the walkthrough comment, SARIF, and telemetry. Cross-referencing with other review tools happens after validation, before publication.

Convergence doctrine ​

Reviewing is a bounded budget, not an open-ended pursuit of certainty.

  • A delivery unit (PR) gets one full review plus at most one re-review of the fix delta. A round with zero new blocking findings is converged; a converged gate is a success, and re-reviewing converged work is a process violation.
  • Only blocking findings (critical/warning, and failing required checks) reopen work. Suggestions are follow-up, never a reason for another round.
  • Re-reviews are delta-scoped: they cover the fix diff plus the blast radius of the findings being fixed (direct callers, corresponding tests), and they inherit prior dispositions for lines and callers the fix did not touch. New critical evidence from non-review sources (CI failures, runtime faults, device tests) reopens a full round.
  • The auto-fix loop is bound by the same budget: a remediation push triggers a delta re-review of its own change, not a fresh full review, and a converged gate ends the loop even if suggestions remain.
  • Coverage gaps are reported as a bounded honest partial review (uncovered chunk counts stay in the output and logs). Retries chase coverage only when a named, decision-changing input was missing, never to reach a coverage quota.

Evidence-led review depth ​

Review depth follows changed contracts and consequences, not a line-count tier. A small authorization change can need more scrutiny than a large mechanical refactor. Within the primary pass, inspect visible entry points, guards, async continuations, cache identity, and sensitive-data consumers. Do not request hypothetical abstractions, flag TODOs without a demonstrated effect, or infer author diligence from writing style. A regression that cannot be proven at the warning bar is downgraded to suggestion with its stated confidence instead of being silently dropped.

Comments describe intent but cannot suppress a candidate. An unrelated if or catch does not prove the required guard covers the failing path. These judgments belong to the validator, with the supporting source, rather than keyword filters. The validator must also check that a suggested removal of a guard, cleanup, lock, or state check preserves the supplied entry-point and failure contracts.

No additional model-driven retrieval round is introduced. Reuse the already fetched source first; a future retrieval experiment must show that a named missing fact changes a decision within the existing time, installation, and token limits.

Severity and confidence ladder ​

Findings carry exactly one severity:

SeverityMeaningVerdict effect
criticaldata loss, security/privacy regression, broken release pipeline, incorrect behavior reaching users or devicesrequest_changes
warningincorrect behavior in the changed scope with user-visible impact, or a missing mandatory gaterequest_changes
suggestionimprovement, hardening, or follow-upcomment at most; never blocks

critical is an issue whose impact does not depend on any assumptions about the inputs — universal and trigger-free. When an issue's impact depends on specific inputs, state, or environment, the finding leads with the conditions that trigger it and severity is calibrated to that dependence.

Findings carry a model-assigned confidence (high/medium/low) recorded with the finding and in telemetry. Confidence ranks and downgrades findings downstream; it must not silently filter publication. Publication thresholds will be calibrated from false-positive telemetry and reaction outcomes (👍/👎), not from raw model self-reports. Until those thresholds exist, confidence must not gate publication.

Evidence standards ​

  • Every published model finding needs a typed proof and an exact quote from its added anchor line. Inline suggestions additionally need exact old lines and a range that matches the analyzed head SHA.
  • PR metadata, diff text, comments, fetched context, CI output, history, and learned examples are untrusted data. Only reviewer policy in the system prompt gives the model instructions.
  • Source-reviewed, tests-passed, built, deployed, and verified are different states. A passing test written after a fix is not before-repro evidence; "CI is green" is supporting evidence, never root cause or behavioral proof.
  • A zero-finding review still owes a summary contract: the summary must state what was verified (files read, paths traced, contracts checked), any residual risks, and whether the change is safe to merge.
  • Fix-evidence audit: fix-flavored PRs (title or label) should state a root cause, a before repro (or why one was impossible), the same scenario passing after, and adjacent checks sharing the affected component, flow, or state. The deterministic audit surfaces gaps as an advisory walkthrough block — never findings, never a verdict change.

Reviewer economics ​

  • Context is bounded by derived byte caps; changed files, exact search, semantic search, consumers, tree, and CI checks run in parallel waves with deterministic result order.
  • Evidence fetched from GitHub is fetched once per head SHA and reused across passes; a pass re-fetching evidence the pipeline already holds is a bug.
  • Reduced reasoning effort for verification passes is a candidate optimization. It ships only after a bakeoff shows no published-FP-rate regression; until then the validator runs at the standard effort.

Evaluator integrity ​

These surfaces are protected gates: the validator pass and its status, checkRunFinalized recovery, SAST secret blocking, severity-to-verdict derivation, finding filters and caps, and the convergence rules above.

  • Do not weaken, bypass, narrow, hide, or redefine a gate to make a result green. Changes to these surfaces are gate-surface-sensitive and require focused review scrutiny plus eval evidence.
  • A validator outage produces an incomplete neutral review with retry — never an approval of unvalidated findings.
  • A review with checkRunFinalized: false requires recovery even if its verdict is non-error.

Runtime evaluation ​

eval/runLlm.ts uses the production prompt builders, deterministic review focuses, validator, and publication filters. It forwards repository and graph context to validation and reports validator outages as errors instead of scoring empty findings as approval. The separate eval/runner/ evaluates the Markdown prompt package; its results do not establish production Worker prompt quality.

Version 1.1 adds paired ownership-context, independent credential-cleanup, and state-change-during-await fixtures. Compare positive recall and clean controls; a lower finding count alone is not improvement.

Change versioning ​

  • Patch: wording, logging, telemetry plumbing that changes no PR-facing obligation.
  • Minor: any change that alters what the reviewer demands of PRs — new evidence requirements, changed severity semantics, new gating, changed convergence rules. Bump at least the minor version and update this document in the same PR.
  • Major: breaking changes to published review output contracts (SARIF consumers, CLI output, check-run semantics).

Gate-surface changes additionally require an eval or bakeoff run as promotion evidence, per the eval harness in eval/.

Source-available under BUSL-1.1. Self-hosting is free for your organisation.