Runtime Reference β
Moved verbatim from AGENTS.md. Operational detail for the review pipeline; AGENTS.md keeps rules and commands only.
Contents β
- Bot Mode internals
- Runtime Estimates
- Per-Repo Config
- Review output and progress
- Auto-Fix Workflow
- KV Keys
- Self-Learning System
- TypeSafe Jev judgments
- AI Usage Cost Tracking
- Unchanged reviews and archived batches
Bot Mode internals β
Webhook event payloads parse through zod schemas (src/webhook-schemas.ts) after HMAC verification. A payload that fails its event schema logs webhook_payload_invalid with the event name and a bounded issue summary, and the handler returns HTTP 200 (Ignored: unrecognized payload) so GitHub does not redeliver a payload that can never parse; handlers therefore never consume fields from unvalidated casts.
CLI session auth for /api/cli/* routes goes through the shared requireCliSession helper (src/api-auth.ts): Bearer token, verified CLI_JWT_SECRET session. Missing or invalid sessions return the canonical 401 body { error: { code: "UNAUTHORIZED", retryable: false, message } }. Autofix and dependabot-repair routes also accept a GitHub token in place of a CLI session and are unchanged in that respect.
For TypeScript and JavaScript diffs, review context includes parsed declarations, resolved relative imports, a bounded local code graph, and repository-scoped Vectorize matches when the index is available. Exact graph and GitHub code-search matches rank before semantic matches. The prompt includes only files fetched at the current PR head SHA; the Vectorize result supplies candidate paths, not file contents. When that graph contains at least two relationships, the walkthrough can render a bounded Mermaid architecture-impact diagram. The Worker creates it from the parsed graph rather than model output, labels observed imports separately from inferred test relationships, and binds it to the analyzed head SHA and a graph fingerprint. The Worker derives focused passes for deleted behavior, contract compatibility, security boundaries, async failure semantics, and changed tests before primary review. PR metadata, diff text, source comments, fetched context, CI output, history, and learned examples are untrusted data; only the reviewer policy gives model instructions. Model findings need a typed proof and an exact quote from their added anchor line before publication. Review depth follows changed contracts, not diff size. Comments claiming intentional behavior and unrelated guards never automatically dismiss candidates. The validator receives bounded fetched source for callers and consumers as well as anchor files, and checks proposed guard/cleanup removals against those entry points. Inline replacements also need exact old lines and a range that matches the analyzed head SHA and one changed hunk; a failed applicability check leaves a prose-only finding. Findings whose fix is prose can still get a suggestion synthesized from the fix text when it is committable code (a wrapping code fence is unwrapped, fixes with residual fence markers or more than 40 lines are skipped), replaces the single added anchor line, and passes the same applicability check; at most 10 findings per review are synthesized, and rejected attempts leave the prose comment.
Context preparation uses bounded waves of GitHub reads and keeps result order deterministic. It probes at most twelve directory-scoped AGENTS.md files per review and paces GitHub code-search requests once every two seconds per App installation. Changed files, exact search, semantic search, consumer files, the repository tree, CI checks, and independent KV prompt inputs may run in parallel when their inputs are ready. The context byte cap is derived from the selected input-token budget so the diff and system prompt retain room in the same request. Symbol parsing is memoized per fetched path.
Synchronize events probe the head-scoped context and completed-result caches before collecting repository files. The probe rebuilds prompt inputs and accepts the cached result only when the PR metadata, merged config, system prompt, cache user prompt, and result fingerprint still match. Cached context records never contain GitHub credentials.
Hosted CLI records can satisfy the remote review gate only for full diff coverage without a profile/focus override. Legacy records lacking coverage evidence require a fresh review. Hosted CLI findings require validation regardless of primary response time unless the caller explicitly selects --skip-validation. The fix command reuses review recommendations without an independent improvement pass. Its PR URL mode uses the requested repository and PR head SHA. Agent and JSON command errors return one error object with a nonzero exit status; empty JSON results stay JSON. Re-review instructions retain the original base and diff scope. Local provider responses must validate the review object and derive the verdict from finding severity. Shared path globs treat punctuation literally; GitHub App authentication accepts both PKCS#1 and PKCS#8 RSA private keys. CLI filesystem tests must mock the home directory before importing configuration and use a unique temporary directory. Never delete or write the real user config during tests. Async CLI actions must be awaited by the top-level error handler.
Provider denials are terminal for that run. If every configured review model is rejected because credits are exhausted, the API secret is invalid, or the model is blocked (**AI review blocked**), finish the check run as neutral and post the provider details on the PR (CB-1731). NEUTRAL remains merge-allowed, but this is an incomplete review rather than a successful one; retry it when provider access is restored. Do not enqueue an orphan retry until billing/secrets/model access have been fixed. Malformed JSON, invalid schema, empty output, and exhausted output caps are recoverable review failures: the Worker advances through the configured chain and queues a retry if every route fails.
Admission and abuse controls β
- Tenant quota enforcement: the monthly plan cap is checked on both the queue path and the synchronous webhook fallback; a tenant over quota gets no review through either path. A missing tenantβinstallation mapping is re-provisioned on demand so an unmapped installation cannot review uncapped.
- Enqueue dedup:
saveQueuedReviewHeadis enforced before a job is enqueued; a delivery that loses the head race is dropped as a duplicate. - Comment-command rate limit: bot commands are capped per commenter via
RATE_LIMIT_MAX(default 20) perRATE_LIMIT_WINDOW(default 3600s). The chat catch-all only matches bot mentions outside code fences, inline code, and blockquotes. - Reply-learning trust: reply feedback enters self-learning only from identities GitHub associates with the repository (
OWNER,MEMBER,COLLABORATOR,CONTRIBUTOR) or collaborators verified via the API. - Provider error redaction: provider error bodies are credential-redacted before they reach logs, KV cooldown records, or CLI responses.
- Outbound URL policy: provider base URLs must be https (localhost http allowed for dev), the platform key is never attached to an unrecognized host, and
.codebeaver.yamlover 64KB is refused. - Session lifetime: CLI session tokens are capped at 7 days total (refresh chains carry
origIat).
Runtime Estimates β
Long-running review phases show a recalculated remaining time and UTC ETA in the replaceable PR progress comment. The comment also shows the current phase and a compact stage trail. The primary review mirrors phase text into the GitHub check run, but heartbeat updates only edit the PR comment. Estimates start with diff size, then use repo-local timing history after at least five matching records. Successful command timings are stored for 90 days under operation-timing:v1:*; failed, cancelled, skipped, and retry attempts never enter the estimate pool. Each operation uses an internal OperationPlan with ordered stages, timing source, current state, and stage start times. Queue and direct command paths share the same ProgressSession and editable operation comment.
Tracked phases include diff analysis, finding generation, validation, posting, autofix, autofix apply, generated tests/docstrings, conflict fixes, simplify, improve, describe, changelog, ask/chat, coverage, CI diagnosis, and focused recipes. A running estimate never becomes zero: overdue phases use longer-run history when available, otherwise an adaptive residual. The final walkthrough's Human review effort describes reviewer time, not bot runtime.
Per-Repo Config β
.codebeaver.yaml in each repo with orgβrepo inheritance. Review policy loads from the PR base SHA, not its head SHA, so a PR cannot alter its own reviewer rules. Controls: auto_review (including whether test-only PRs are reviewed and the test_review_mode: skip, lightweight, or full; test_review_mode takes precedence over legacy skip_test_only_prs, and the worker reviews test-only PRs in lightweight mode by default β deterministic checks only, no AI tokens), slop_detection, pre_merge_checks, labels, suggested_reviewers, walkthrough, supplementary_prompt, and blind_first_review. Blind-first mode withholds PR title and description from the primary model candidate pass. After validation, a bounded intent pass marks surviving findings as consistent, conflicting, or unknown using PR metadata and up to three linked issues; it cannot add findings or raise severity. Suggested-reviewer rules may scope reviewers with paths; when enabled, matching reviewers are shown on the walkthrough and auto_assign requests them from GitHub. When no enabled rule matches, the Worker falls back to matching .github/CODEOWNERS entries; auto_assign: true requests those owners too.
pre_merge_checks.operational_risk_notes defaults to true and adds advisory walkthrough rows for migration, secret/config, Wrangler binding, and Terraform changes. Each row gives a short blast-radius note and asks for a staging smoke before merge. It never blocks the check; set it to false for repositories with a different release process.
path_rules may include a stable id, severity, and enabled flag. Repository config can disable inherited rules with disabled_rule_ids. ast_rules run deterministic checks against added TypeScript and JavaScript AST nodes. Each rule has an id, message, node_type, optional dotted callee, plus optional paths and languages filters.
Built-in SAST findings remain authoritative. The reviewer passes a bounded, severity-sorted summary to the model only as context: it must not repeat or dispute those findings, but may inspect related code for effects or incomplete remediation.
TypeScript and JavaScript context collection follows static relative imports from changed files within the existing file and byte limits. It never evaluates source code or follows imports outside the repository.
Review output and progress β
- A normal review owns one batch comment per PR head SHA. A rerun of the same head edits that comment, including when the KV comment pointer is missing. A new head gets a new batch comment. Retry and failure notices put the latest state first and retain earlier provider attempts in a collapsed history.
- After the review gate accepts a new head, post
Preparing review contextbefore fetching its diff and repository context. Reuse that batch comment for later progress and terminal states so context failures are visible on the PR. - Long reviews and queued long commands update one active comment every 30 seconds. Running comments show the active stage, completed and pending stages, remaining time, ETA, start time, update time, and elapsed time. Heartbeats do not create new comments. Check runs change for phases and measured progress, not timestamp-only heartbeats.
- Terminal updates stop and drain the heartbeat before replacing the progress body. Queued review comments show an approximate count of waiting reviews ahead, an estimated start time, and the UTC time of the last queue check. A separate once-a-minute queue-status cron refreshes those comments; it does not edit every waiting PR when a single message dequeues.
- Queued review jobs retain that comment ID. If queue processing retries or fails, replace the queued body with the retry or failure state.
- A queue job records each state once per invocation and attempt, with final states recorded once per invocation. Cleanup may retry its GitHub and KV side effects without adding duplicate job events, history rows, or
review_job_finishedlogs. - Incomplete chunk reviews and transient queue failures use the KV-backed orphan queue after acknowledging the broker message. Completed chunk results are cached by PR head, so recovery resumes the missing chunks after a deploy, crash, or broker retry exhaustion.
- Orphan recovery is capped at ten attempts. Exhaustion finalizes the carried GitHub check as
neutral, records the retry as failed, and updates the batch comment. If check finalization fails, the KV entry remains for another recovery pass instead of leaving an in-progress check behind. - Due orphan retries are dispatched from the global KV prefix before the scheduled cleanup rotates through installed repositories. The repository scan remains a fallback for direct recovery and checks that need GitHub reads. A dispatched retry refreshes its own queued-head timestamp so a later stale queue scan does not replace it, but it never takes ownership from a newer invocation.
- A queued review waits only for a fresh in-flight marker. If a redelivery sees that marker, it schedules the same job back onto Queues with a 30-second delay, retaining the invocation, comment, and check-run IDs. When the marker becomes stale, the replacement recovers the old check instead of creating a second lifecycle. The retry status records
queue_deliverywhen the marker belongs to the same invocation, which identifies a possible Worker restart or cancellation before the application can log a catch/finally path. - In-flight heartbeats prove that the Worker is alive; they do not extend the review phase deadline. Recovery freshness uses the last completed phase, starting at
prepare_context, so a hung phase is reclaimed after the normal stale interval even while the heartbeat timer continues. REVIEW_JOBSis capped at six concurrent consumer invocations. Do not raise queue and chunk concurrency together without observing GitHub App rate-limit headroom under load.GITHUB_RATE_GOVERNORadmits reviews through an installation-scoped Durable Object. It starts atGITHUB_REVIEW_CONCURRENCY(four in production), ramps towardGITHUB_REVIEW_MAX_CONCURRENCY(eight in production) after clean intervals, halves its lane capacity after a GitHub limit, and keeps code search in a separate two-request lane. A governor outage fails open and logs the error.- A non-search GitHub API rate-limit response writes an installation-wide backoff through the reset time. Code-search limits use only the separate search cooldown. Queue consumers, orphan recovery, and scheduled repository scans defer that installation instead of multiplying failed API calls.
- Large-diff chunk calls use
REVIEW_CHUNK_CONCURRENCYwhen set, bounded by the available chunks. The default is two calls total. - Repository instructions keep the beginning and end of each file within the prompt budget. The Worker logs the retained and omitted character counts when it clips a file.
- A manual request for an active head is ignored, except when that head is in a delayed retry: it replaces the retry even if its retained in-flight marker is fresh. A request for a newer head is deferred until the active review releases its lock; scheduled cleanup is a fallback if that handoff is missed.
- Scheduled cleanup repairs a stale in-progress AI check when the same head already has a persisted, posted approval. It cancels and reruns only when that repair fails or no completed review state exists.
- Queued-review head pointers advance by request time. A retry reclaims its own same-head pointer so an out-of-order duplicate cannot suppress it.
- Manual
/reviewrequests use the same per-PR head and status records as automatic reviews. Requests for an active or queued head reuse the existing run, while a newer head is retained for review after the active run. Terminal pointers remain briefly so late deliveries cannot start duplicate work. - Maintainers can use
/review cancelor@bot cancelto cancel a queued, retrying, or active review. The command updates the existing batch status and check run when available; a later/reviewclears the cancellation marker before the active-head reuse check, so the next queued delivery can start. - Keep coding-agent instructions inside the current batch and only when open findings or failed checks exist. Do not tell agents to act on archived work.
- Cross-tool summary refreshes are best-effort. A failed KV, token, GitHub read, or comment update logs
cross_reference_update_failedwith the repository and PR number, while the webhook still completes.
Auto-Fix Workflow β
.github/workflows/auto-fix.yml β triggered by auto-fix label on any PR. Flow:
- Label
auto-fixapplied on an internal PR β GitHub Actions calls the Worker - Worker reviews the diff and writes fixes to a
codebeaver/autofix/...branch - The workflow checks out that branch and runs
autofix.verification.commands - Passing commands create a stacked PR targeting the original PR branch
- The workflow removes the label after it finishes; a failed run can be retried by applying it again
The workflow never checks out or writes to fork PRs. Failed verification leaves the autofix branch available for inspection and does not modify the author branch.
No custom secrets needed β auth uses secrets.GITHUB_TOKEN (GitHub-native, repo-scoped, auto-expiring). Optional: CODEBEAVER_URL (defaults to worker URL).
The caller repo must trigger on pull_request_target with labeled event:
on:
pull_request_target:
types: [labeled]Usage in other repos: add to their .github/workflows/:
on:
pull_request_target:
types: [labeled]
jobs:
auto-fix:
if: github.event.label.name == 'auto-fix'
uses: codebeaver/codebeaver/.github/workflows/auto-fix.yml@main
secrets: inheritCLI Autofix β
The codebeaver autofix CLI command delegates the entire fix pipeline (diff fetch, AI review, code generation, git commit) to the Worker:
codebeaver autofix --pr-url https://github.com/owner/repo/pull/N
codebeaver autofix --pr-url https://github.com/owner/repo/pull/N --jsonUnlike codebeaver fix (local review + suggestions), autofix prepares a verification branch via the Worker's GitHub installation token. The reusable workflow opens the follow-up PR only after the configured commands pass. Requires CLI auth (codebeaver auth login).
autofix:
verification:
commands: ["pnpm install --frozen-lockfile", "pnpm lint", "pnpm typecheck", "pnpm test"]
timeout_minutes: 20KV Keys β
in-flight:{repo}:{pr}, last-reviewed:{repo}:{pr}, findings:{repo}:{pr}, feedback:{repo}:{pr}:{findingId} (per-entry, where findingId is an FNV-1a hash of repo:file:normalized-issue; legacy blob at feedback:{repo}:{pr} still read for backward compat), dismissed-findings:{repo}:{pr}, auto-best-practices:{repo}, known-false-positives:{repo}, accuracy:{repo} (unified per-category telemetry: shown/accepted/rejected; fp-stats:{repo}:{category} is legacy and read-only), pause:{repo}, ratelimit:{repo}, ratelimit-config:{repo}, rate-limit:v1:{scope}:{principal}:{windowIndex} (fixed-window counters: issue-comment per commenter, auth-exchange/tenant-claim per IP; TTL is twice the window), conventions:{repo}, learnings:{repo}, orphan-queue:{repo}:{pr}, merged-pr-cleanup:v1:{repo}:{pr}, pr-history:{repo}, context/{owner}/{repo}:{sha}, reviewed-commits:{repo}, content-cache:v1:{ns}:{repo}:{sha}:{id} (immutable file contents, repo trees, and repo instruction files keyed by commit SHA, 30 days), embedding-cache:v1:{sha256(model:query)} (semantic query embeddings keyed by embedding model, 30 days), vector-index-manifest:v1:{repo} (indexed paths, used to delete vectors on app uninstall), cancel:{repo}:{pr}, ignore:{repo}:{pr}, skip-reason:{repo}:{pr}, closed-manual-review:{repo}:{pr} (30 days), review-status-comment:{repo}:{pr}, autofix-selection:{repo}:{pr}, autofix-prompt:{repo}:{pr}, review-log:0{reverseTimestamp}:{repo}:{pr}:{ms}, review-job-status:v1:{invocationId}, review-job-event:v1:{invocationId}:*, review-dlq:v2:{invocationId}, review-dlq-pr:v2:{repo}:{pr}:{invocationId}, review-dlq-recent:v1:{reverseTimestamp}:{invocationId}, review-dlq-terminal:v1:{repo}:{pr}, review-dlq-replay:v1:{id}, review-dlq:v1:{id} (legacy until review-dlq-migration:v1 is complete), delivery:{x-github-delivery} (5-minute webhook redelivery dedup). quality-report:latest holds the weekly quality snapshot as JSON (30 days). Tenant state lives under tenant:v1:{tenantId} (record), tenant-installation:v1:{installationId} (installationβtenant mapping), and tenant-reviews:v1:{tenantId}:{YYYY-MM} (per-tenant monthly counter enforced at review enqueue). Review-result caches are repository-scoped (one repo maps to one installation/tenant), not tenant-keyed. enqueueWebhookReview returns a tri-state outcome: queued (including already-reviewed/already-active heads), unavailable (queue missing or delivery failed; callers may use the synchronous direct path), and denied (monthly or concurrent tenant cap; callers must not fall back β the webhook returns 202 without a review). Concurrent-review leases are acquired on the tenant-keyed GitHubRateGovernor lane at enqueue and released when the job status record reaches a terminal state (completed, failed, cancelled, superseded) inside recordReviewJobStatus; the 30-minute lease TTL remains the backstop. Tenant ownership claims serialize through the same tenant-keyed DO (claimTenantThroughGovernor, blockConcurrencyWhile) with the KV claimTenant path as the binding-absent fallback. Hosted tenant portal state: oauth-state:v1:{nonce} (single-use login state, 10 minutes), web-session:v1:{sid} (server-side tenant session holding the login's GitHub token; the cb_session HttpOnly cookie carries a signed {sub, sid, kind: "tenant"} JWT β the token never reaches the browser, 8-hour TTL), installation-repos-cache:v1:{installationId}:{login} (5-minute repo listing cache; invalidated by installation_repositories webhooks via installation-repos-cache-logins:v1:{installationId}), and repo-disabled:v1:{installationId} (JSON array of repos with auto-review disabled; absent key or installation means enabled, so self-host is unaffected β checked at PR-review entry and for non-management @bot commands). Hosted billing + entitlement state: tenant-activerepos:v1:{tenantId}:{YYYY-MM} (JSON array of repos whose reviews ran this month, written by recordReviewLog, backs the plan repo limit β free covers 1 repo/month via PLAN_REPO_LIMITS, checked at PR-review entry with a posted skip reason), tenant-seats:v1:{tenantId}:{YYYY-MM} (distinct human committers this month, metered nightly by the 30 2 * * * cron over the tenant index; [bot]/automation authors excluded), marketplace-entitlement:v1:{login} (Marketplace purchase state from marketplace_purchase webhooks; plan mapping comes from MARKETPLACE_PLAN_IDS, absent config records events without changing tenants; cancellation suspends the matching tenant, payment resumes it β installation.unsuspend respects a cancelled entitlement), and sub-notice:v1:{tenantId} (30-day one-time comment flag when reviews are skipped for a suspended tenant). POST /api/tenants/state is the staff-only operator suspend/resume surface. Suspension is enforced on the auto-review path (PR entry) and on @bot action commands; label-triggered and CLI-initiated reviews are not yet suspension-gated (follow-up).
reverseTimestamp is this 13-character value: (9999999999999 - Date.now()).toString().padStart(13, "0"). KV prefix scans return recent usage first because lower keys sort earlier.
Scheduled cleanup compacts both learning-rule lists for every installed repository: it removes invalid or 90-day-stale rules, merges duplicate patterns, retains at most 50 rules per list, and logs retained and removed counts.
A Monday 06:00 UTC cron (0 6 * * 1) writes the weekly quality snapshot: runWeeklyQualityReport scans the most recent 500 review-log records and stores one quality-report:latest JSON object with the week window, review count, published-finding count, rejected-or-dismissed count, an estimated false-positive rate (rejected over generated where filter telemetry exists), average cost per review, and the top dropped-filter stages. Nothing is posted anywhere; the CLI reads it back through GET /api/cli/quality-report.
When the REPO_MEMORY Durable Object binding is enabled, feedback entries, learning rules, accuracy counters, PR history, and dismissed findings fold transactionally inside one RepoMemory instance per repository instead of the racy KV read-modify-write paths; the KV records remain the fallback and stay readable. Merge-time π/π reaction collection is off by default (learning.collect_merge_reactions); reply sentiment, dismissals, and applied autofixes feed learning regardless. REVIEW_CACHE also holds content-cache:v1:* entries (file contents, repo trees, repo instructions keyed by commit SHA, 30 days) and embedding-cache:v1:* query embeddings, so re-reviews and unchanged-SHA retries skip repeat GitHub and Workers AI reads. Review usage records carry embeddingCalls and embeddingCacheHits for semantic-context spend visibility.
Review logs may include optional runtime fields: totalDurationMs, queueWaitMs, sizeBucket, postingDurationMs, finalizationDurationMs, changedFilesCount, diffChars, and the provider-returned review model or models for chunked reviews.
Per-PR finding outcomes are stored for 90 days as independent records at finding-outcome:v1:{repo}:{pr}:{finding-hash}. The dashboard reports distinct findings seen during that period, verified fixes from later reviews, manually resolved conversations, and current open findings. A resolved conversation is not proof that the code was fixed, so it may overlap a fixed finding. When a forced rerun targets the same head SHA, the walkthrough is refreshed but locations already reported on that commit do not create duplicate inline threads. Stored findings retain those earlier locations so later same-head reruns stay quiet; they are not a record of only the threads posted by the latest rerun. Outcome summaries scan a bounded number of records; the API returns that limit with a capped result so the dashboard can show that its counts are partial.
Direct autofix exact-replacement lookups cache parsed {old,new,note} patches at autofix-patch:v1:{repo}:{sourceBlobSha}:{path}:{line}:{suggestionSha256} for 21 days. The key contains the Git blob SHA, requested path and line, plus a hash of the suggestion, never source content. Normal autofix reuses findings:{repo}:{pr} only when its headSha equals the current PR head; a missing or stale record, and any focused repair prompt, runs a fresh review instead. Completed AI analyses are cached for 14 days under hashed, tenant-scoped review-result:review-result-v4:* keys. The key fingerprints the PR number, model/provider, stable code/context inputs, deterministic scanner settings, review mode, and response schema; cache records never store raw prompts or diffs. The validator model is part of the fingerprint on purpose: changing validator configuration invalidates prior analyses so a result is never reused across validator configs. Complete zero-finding AI results are not written or reused, so a valid empty envelope from a degraded provider cannot suppress later recovery for the 14-day TTL. A synchronize event whose full review fingerprint is unchanged reuses the stored analysis inside the queued review instead of calling the models again. A changed prompt, context, scanner policy, partial result, or incomplete validator result runs the models again. The stored base SHA is audit data, not a cache gate. Successful large-diff file chunks use similarly hashed review-chunk:v1:* keys with the same TTL. App-cache hits record zero new provider usage plus avoided-call/token telemetry.
Every paid provider response must expose its usage. Primary and validator usage is part of the review analysis; linked-issue scope checks and custom pre-merge checks merge their usage into that same review record. Malformed HTTP-success responses, empty output, transient 429/5xx responses, and output-cap exhaustion get one bounded same-route retry when time remains. Direct DeepSeek retries high-reasoning empty and capped responses with thinking disabled. If every route exhausts its output cap before producing valid JSON, the review is recoverable and queues a retry. Preserve usage from every parseable provider response.
NVIDIA-hosted DeepSeek V4 review calls use the 8,192-token review-output cap. Z.ai Coding and direct DeepSeek V4 Flash review calls use 65,536 output tokens: the direct API's high-reasoning mode can consume a smaller allowance before it emits final JSON. REVIEW_INPUT_TOKENS_CONFIG controls route-specific prompt budgets separately; set million-context routes below their combined context limit so output and provider overhead still fit. Large reviews use one complete prompt when it fits the primary route, otherwise they batch whole diff hunks using the first available route's budget and never silently truncate a file. The rendered prompt also uses a 120,000-token latency limit, even when the provider accepts a larger request. If a batch falls back to a smaller route, it is repacked for that route before the call. Chunk caches include the PR head SHA and rendered context. The Worker caps high-thinking direct DeepSeek attempts at 360 seconds and Z.ai Coding attempts at 240 seconds, while reserving time for later routes. Its 900-second review wall matches the Cloudflare Queue consumer limit; a chunk chain remains capped at eight minutes, and the scheduler reduces a late wave's route share when fallback attempts no longer fit. REVIEW_MAX_TOKENS_CONFIG can override output caps by exact route, model, or provider; review-log usage line items store the cap sent to each call and the ordered fallback attempts.
REVIEW_TIMEOUT_CONFIG applies the same route, model, provider, and default precedence to per-model-attempt timeouts in milliseconds. JSON repair and response-format retries share that attempt budget. The review wall and later-route time reservation remain hard limits; review logs retain each attempted timeout. Providers are declared in the src/providers.ts registry (credentials, pinned models, request shaping, timeout floors, response-format handling). With REVIEWER_API_KEY_ZAI configured, the fixed chain runs Z.ai Coding first, then direct DeepSeek; REVIEW_PROVIDER_ORDER reorders known providers. Reviews with at most 20,000 changed-diff characters use a fast provider path. Larger diffs are split into focused batches capped at that size when hunk boundaries permit. Z.ai keeps thinking enabled with reasoning_effort: low, direct DeepSeek receives thinking.type: disabled, OpenRouter omits its high reasoning request, and fast attempts are capped at 120 seconds. Reasoning mode is selected per batch. The attempt logs record reasoning_path and reasoning_effort so the selected path is visible. review.depth overrides the automatic selection: fast forces the fast path (no reasoning, fast timeouts, fast-priority ordering) regardless of diff size and holds the finding cap at the configured noise.max_findings; thorough forces high reasoning on every attempt, disables the fast path, and raises the effective finding cap to at least 15; standard (default) keeps the size-based behavior. Validator calls disable direct DeepSeek thinking. A transient 429 in a multi-route chain advances to the next provider without a second outer retry; single-route setups retain one bounded retry. Terminal provider denials such as allowance errors are terminal for the current run, so later chunks and validator batches skip the denied provider. The Worker stores the denial in KV for 15 minutes β including a truncated HTTP status and provider error text as the denial reason β so new reviews skip that route until the marker expires. A route skipped by the marker reports Provider access cooldown active (last denial: β¦) instead of a bare provider 429, and attempt history classifies those skips as provider_denied, not rate_limited. Validator batches have a 180-second operation budget and reserve time across the ordered provider chain, giving each of the two default routes up to 60 seconds before moving on. Each validator attempt emits validator_model_attempt with provider, model, HTTP status, elapsed time, timeout budget, timeout flag, and success flag. Findings may also carry nitpick severity: pure style/preference nits that bypass the noise floor and inline posting entirely, publish only in a collapsed Nitpicks walkthrough digest (cap 10), and never affect the verdict. The validator returns compact decisions identified by the server-issued validation_id across three lanes: validated, downgraded (each with a reason), and dropped (each with a reason). Downgraded findings publish at suggestion severity with their original file, line, and evidence instead of being withheld, because the team cannot act on a finding that was never published. Findings carry a confidence field (high/medium/low) that the validator uses to rank and downgrade rather than silently filter. Every candidate ID must appear exactly once across validated, downgraded, or dropped; missing, unknown, and duplicated IDs make the batch retryable. The walkthrough filter-stats line reflects the lanes: N generated Β· N published Β· N removed by validation and policy filters, plus Β· N downgraded to suggestion when any finding was downgraded. An empty HTTP-success response, malformed envelope, transient 429/5xx, or output-cap exhaustion gets one bounded same-route retry when time remains. Direct DeepSeek retries high-reasoning empty and capped responses with thinking disabled. If all routes return either failure, empty output, or output-cap exhaustion, the Worker queues a review retry. Retry-After is honored up to five seconds without extending the route deadline. Chunk timeout reservations exclude routes already denied or capacity-exhausted in the current run. A provider timeout is local to the chunk that experienced it, so parallel chunks can still try that route; the timed-out route is skipped only for that chunk's remaining fallback chain. The chunk deadline also aborts the underlying provider request, so a timed-out route cannot keep running while its fallback is already processing. NVIDIA 503 ResourceExhausted worker-limit responses skip retries for that route and move to the next model; chunked reviews stop submitting further chunks there. Chunked reviews also disable a route for the current run after its first 401, 402, or 403 response, rather than repeating a terminal denial for every pending batch. Large reviews keep the rendered diff and context in one request only when they fit the primary route's configured input budget and the 120,000-token latency limit. Chunking starts when either limit is exceeded, or when recovery needs to retry incomplete work after a full request fails, returns empty output, or reaches its output cap. Chunk planning uses the first available route's input budget; fallback routes repack each batch for their own lower limit. For the pulls-files fallback, only files with patch bodies become chunks; disclose patchless files separately. Large-diff batches target 240,000 diff characters (β80k estimated tokens β sized so system prompt, chunk context, and batch stay inside ~128k model context windows) and are limited by the route budget, route timeouts, and review wall rather than an arbitrary provider-call count. A partial review fails the required check and queues a retry. Successful batches are cached, so the retry sends only failed or omitted batches to a provider. Chunk results must report exact reviewed, failed, deadline-omitted, and input-budget-omitted counts plus the current-path batches in each state. Large reviews allow up to 10 noncritical findings unless a higher repository cap is configured. Review records include the resulting coverage gap so omitted chunks count as a recall failure. After chunk results merge, a semantic dedup pass applies the suppression matcher (same file, normalized-issue token overlap β the same identity used for dismissed and previously-reported findings) across chunks: when two candidates match, the higher-severity one stays and severity ties keep the earlier chunk's finding. Dropped duplicates are counted in the crossChunkDeduped telemetry counter (cross_chunk_deduped on review_complete); the exact file:line:issue merge still runs first. Reviews also record filesWithoutFindings (bounded to 50; logged as files_without_findings on review_complete): changed files whose chunks or full review produced no findings, as pure observability for recall analysis. Retry metadata is head-scoped. Queue and orphan retries carry the current head SHA, progress comment, check run, invocation, installation, and request time; when a newer head arrives, its retry entry replaces the prior head's attempt count and artifacts.
Finding status lifecycle β
Inline finding comments carry a codebeaver:finding-status marker. When an author resolves the bot's thread, the comment marker flips to finding-status:resolved with a Marked resolved by @user note (bot-authored, marker-matched comments only; failures log and never enter the webhook error path). Push-driven re-reviews apply the same flip to findings the comparison shows as resolved. Resolution stamps resolvedAt/resolvedBy on the stored finding record; the dashboard shows a resolved badge.
Self-Learning System β
The reviewer learns from feedback to reduce false positives and reinforce real findings. Feedback flows through these sources:
- Review dismissal (
handleReviewDismissed): When a person dismisses a bot review, all findings are recorded as rejected signal (accuracy, per-entry feedback thumbs_down, dismissed-findings blob). Legacy fp-stats keys are read only. Bot-initiated stale-review dismissals are ignored. - Thread resolution (
handleReviewThreadEvent): When a review thread is resolved, the finding is recorded as accepted (thumbs_up). - Re-review resolved (
recordAcceptedFeedback): When a previously-flagged finding no longer appears in a new commit, it's recorded as accepted. - Autofix applied (
recordAcceptedFeedback): When the AI commits a fix for a finding, it's recorded as accepted. - Inline comment reply sentiment (
learnFromReply):pull_request_review_commentreplies to bot findings are classified from the webhook's exact comment and parent. Edits replace that reply's earlier signal; deleting the reply removes it. The original finding category is retained rather than learning from the reply text. - Reactions on merge (
collectFeedbackOnMerge): π/π reactions on bot comments are collected when the PR merges. GitHub provides reaction APIs but does not deliver a reaction webhook for this Worker, so this is not real-time feedback.
All feedback is stored as per-entry keys at feedback:{repo}:{pr}:{findingId} (findingId = FNV-1a hash of repo:file:normalized-issue); the legacy feedback:{repo}:{pr} blob remains read for backward compat. Legacy file/line-shaped keys are a compatibility fallback consulted only during reply deletion. After every new feedback record, processFeedbackLearning refreshes accepted and rejected patterns plus category accuracy. Merge handling collects reactions before that refresh. Accuracy markers store the counted reaction per feedback entry: edited replies move the count between accepted and rejected, and deleting a reply retracts it. The stored marker set is capped at 5,000 entries per repository to keep the accuracy record bounded.
Hard filter: loadDismissedFindings loads dismissed findings for a PR and runReviewAnalysis strips any AI finding matching a dismissed file and normalized issue before posting β line-agnostic, so code edits don't keep a dismissed issue suppressed and a different issue at the same line still surfaces. Built-in SAST findings are never suppressed by dismissals.
Prompt-level signal: loadFpRates, loadAccuracyMetrics, loadBestPractices, and loadFalsePositives inject per-category stats and learned rules into the review system prompt on every run. Each injection logs learning_rule_applied (matched best-practice and false-positive counts), and after an incomplete-free review completes, recordLearningRuleApplications bumps timesMatched/lastMatchedAt on exactly the rules whose patterns appeared in the prompt β at most one write per rule store per review.
Rule lifecycle: new rules start pending; manual approval stays available via the dashboard and CLI. A pending rule whose recurrence count reaches RULE_AUTO_APPROVE_COUNT (5) is promoted to approved at the next extraction refresh; rejected is sticky and never auto-approves.
Eval fixture skeletons: author-confirmed misses (a reply classified as a dispute, or a same-head review dismissal) are appended to eval-fixture-candidates:{repo} with {file, line, issue, evidence, rejectionSource, headSha, recordedAt}; issue/evidence are truncated to 200 chars and the list is capped at 20 entries FIFO. GET /api/cli/learning/fixtures?repo=β¦ (CLI JWT auth) returns the list. This is data-plane only and never affects review outcomes.
Learning rules use the newest feedback timestamp as lastSeen; reprocessing old feedback does not keep an old rule alive.
TypeSafe Jev judgments β
TYPESAFE_API_KEY (operator-provided secret, absent by default) enables bounded judgments from TypeSafe's Jev System One model β pinned jev-1.13.0, raw HTTP in src/jev.ts, 4-second AbortController timeout, billed per input token. jevSystemOne returns null on unconfigured, timeout, HTTP error, or malformed payload and logs jev_request_failed; it never throws into a review path.
| Judgment | Trigger | Shape | Fallback |
|---|---|---|---|
reply_sentiment | reply to a bot finding (learnFromReply) | one Choice agrees/disputes/other; confidence < 0.5 records no signal | word-list regex (negation guard included) |
intent_reconciliation | blind-first gate after publication filters | one Choice per finding over shared PR-intent state; every finding must get a valid answer | validator-chain LLM call |
slop_detection | review start when slop_detection.enabled and regex did not already flag | Score over 3 levels; slop at score β₯ 1.5 union regex hit | regex-only baseline |
evidence_check | after publication filters | one Choice per finding (cap 20, evidence_quote required) judging quote-vs-issue; writes finding.evidence_check + jev_evidence_check_summary | no check recorded |
State is bounded untrusted text only (reply β€ 2k chars, issue β€ 600β800, evidence quote β€ 400, PR body β€ 12k) β never credentials or full files. hasNonLatinLetters skips every judgment for predominantly non-Latin text because TypeSafe documents non-English scripts as weaker.
Results are advisory only: sentiment feeds learning, intent status labels findings, slop applies a label, evidence checks are recorded on findings. None can change a verdict, suppress a finding, or bypass the validator. Usage from the three in-review judgments (intent, slop, evidence check) merges into the review usage record as intent_reconciliation_jev, slop_detection_jev, and evidence_check_jev line items (provider typesafe), so their tokens and cost appear in the walkthrough AI usage: block and the usage breakdown; reply-sentiment calls happen on the comment trigger and are not part of the review summary. Telemetry: jev_request_completed, jev_request_failed, jev_sentiment_disagreement, jev_slop_disagreement (regex-vs-Jev verdict splits for threshold tuning). Review-log records persist per-finding evidence_check results and a jevDisagreement counter for review-side slop disagreements; reply-sentiment disagreements stay in learning logs because they fire outside the review run.
AI Usage Cost Tracking β
Completed reviews store OpenRouter usage from successful model responses in the review log. The stored summary includes prompt tokens, completion tokens, total tokens, cached/reasoning tokens when OpenRouter returns them, total cost, and per-call line items for review, chunk, fallback, validator, and autofix calls. Run results group provider calls by endpoint and returned model. Each group reports cached and non-cached input tokens, output tokens, call count, cache hits, and cost, followed by a run total. Application-cache hits never count as provider calls. Chunked and retried calls remain separate inputs to that aggregation.
Direct DeepSeek and Z.ai calls are priced from published list rates when the response carries no provider cost: cached input tokens bill at the cache-hit rate and the remaining input at the cache-miss rate, derived from prompt_cache_miss_tokens when reported and from the prompt total otherwise. Z.ai figures are API list-price equivalents for the subscription endpoint, with its cached-input rate applied to reported cached tokens. The review system prompt keeps static policy text ahead of per-repo blocks and places per-commit CI results in the final section so provider-side implicit prompt caching can reuse the stable prefix across requests; inserting per-run content ahead of the static text defeats that caching.
Stats responses carry cachedUsageReviews and hitRate per period β usage-carrying runs with at least one application cache hit, over all usage-carrying runs β so a silent cache collapse (for example after a prompt change rotates cache keys) is visible in the dashboard rather than only in per-run telemetry.
Before the diff reaches a model, the pipeline applies repo ignore_paths plus a built-in model-facing exclusion list (additional lockfiles, minified assets, source maps, and coverage/vendor/out trees) to the AI payload only. SAST scans the raw diff, so these defaults never narrow the secret scanner.
The embedding model is configurable via REVIEW_EMBEDDING_MODEL and REVIEW_EMBEDDING_DIMENSIONS (defaults @cf/baai/bge-base-en-v1.5, 768 dimensions). Switching models requires a Vectorize index with matching dimensions and a full repository re-index; embedding cache keys include the model, so previous-model entries simply expire. Dimension mismatches fail closed with an embedding_dimension_mismatch warning instead of upserting mismatched vectors. Walkthroughs and completed check runs show the provider-returned review model, which can differ from the configured primary model after fallback routing.
The PR walkthrough shows AI usage: <tokens> tokens Β· <cost> when telemetry is enabled. Dashboard stats aggregate total/average spend and tokens from review-log:* records. Completed review records also persist removedFindings: up to 20 dropped-candidate identities (issue text truncated to 200 chars) tagged with the stage that dropped each (validatorAutoDropped, validatorDropped, malformed, evidence, proof, dedupe, lint, nullGuard, loopback, dismissed, noise, cap), collected by the publication filters and noise controls.
Unchanged reviews and archived batches β
Completed clean approvals can reuse analysis when the PR's review-input fingerprint matches, including its full diff, repository tree, prompt, context, model, and scanner policy. A failed tree lookup limits reuse to the exact commit. An empty result becomes reusable only after publication and check finalization succeed. Partial reviews, validator outages, and failed pre-merge checks do not qualify. Current deterministic and pre-merge checks still run on reuse. A cached rerun on the same head reuses the existing GitHub review only when its commit, state, and body match; a dismissed approval requires a new review. Older batch summaries show an outdated notice and link to the current review. Repeated archival preserves the original body. Outdated does not mean fixed. Partial and validator-incomplete runs preserve prior findings and blocking reviews; they cannot mark missing findings resolved. Cached reruns recover missing review IDs from GitHub before posting another approval.