Architecture β
CodeBeaver has three user-facing programs backed by one Cloudflare Worker:
- a GitHub App webhook service;
- the
@codebeaver/cliterminal package; - an authenticated React dashboard.
The Worker owns GitHub installation auth, review orchestration, model calls, deterministic checks, persisted review state, learning, CLI APIs, and dashboard APIs.
System map β
flowchart LR
GH["GitHub App webhooks"] --> W["Cloudflare Worker"]
CLI["CodeBeaver CLI"] --> W
DASH["Dashboard SPA"] --> W
ADMIN["API Admin session service"] --> DASH
W --> GHA["GitHub REST and GraphQL APIs"]
W --> MODEL["OpenRouter-compatible models"]
W --> KV["Cloudflare KV: REVIEW_CACHE"]
W --> VEC["Cloudflare Vectorize"]
W --> EMBED["Workers AI embeddings"]The dashboard is a static single-page app. Cloudflare Access owns the browser session; the app stores only identity-scoped query data and calls the Worker for data. The Worker validates the Access assertion through API Admin before dashboard queries or cache hydration begin. Production calls API Admin through the API_ADMIN service binding.
Workspace β
src/
index.ts Worker entry, shared routing, PR event gates
webhook-routes.ts HMAC verification and webhook dispatch
api-routes.ts CLI, auth, and dashboard API routing
review.ts Review orchestration and context gathering
ai.ts Provider calls, usage, retries, validation
findings.ts Built-in SAST and persisted finding state
config.ts Org and repository config loading
learning.ts Feedback ingestion and learned rules
walkthrough.ts GitHub output formatting
github.ts GitHub auth and API helpers
cleanup.ts Scheduled cleanup and orphan recovery
commands/ Pull-request comment handlers
codebase-context.ts CLI indexing and search endpoints
vector-context.ts Vector IDs, chunks, search, merge refresh
symbol-context.ts Static declarations, imports, and code graph
structural-impact.ts Affected consumers and tests
stats-api.ts Dashboard data endpoints
logger.ts Structured logs and request correlation
packages/cli/
src/index.ts Command registry and output formatting
src/config.ts OAuth token and user config storage
src/local-review.ts Direct provider mode
src/finding-bundle.ts Stable v1 bundle and agent handoff
skills/SKILL.md Packaged agent instructions
apps/dashboard/
src/routes/ Login, overview, reviews, users, learning
src/lib/api.ts Worker API client and response types
src/lib/query-client.ts Persisted TanStack Query cacheWorker routing β
The default export in src/index.ts wraps both scheduled and HTTP handlers with request metadata. Every HTTP request receives a request ID. Unhandled failures return a generic JSON error containing that ID.
Public and operator routes β
| Route | Method | Auth | Purpose |
|---|---|---|---|
/ | GET | none | Redirect to dashboard login |
/ and unmatched POST paths | POST | GitHub HMAC | Webhook receiver |
/health | GET | none | Verify required secrets and GitHub App JWT creation; report deployed SHA and resolved review-route hosts |
/diag?repo=&pr= | GET | debug bearer token | Read-only installation and PR diagnostic |
/test-review?repo=&pr= | GET | debug bearer token | Run a live forced review with GitHub side effects |
DEBUG_ALLOWED_REPOS can limit the two debug routes. DEBUG_INSTALL_ID selects the GitHub App installation they use.
CLI routes β
All CLI routes are POST requests under /api/cli/:
auth/exchange auth/refresh
review describe improve
changelog test autofix
ask coverage sarif
index-repo search-context
learning learning/delete
learning/approve learning/rejectReview, code generation, indexing, and repository questions validate the CLI JWT and then resolve the GitHub App installation for the named repository. Local-only CLI commands do not call these routes.
Dashboard routes β
/api/auth/session
/api/stats
/api/stats/reviews
/api/stats/users
/api/stats/review/:owner/:repo/:pr
/api/stats/learning/rulesStats routes require an admin session. The learning-rules endpoint uses GET, PATCH, and DELETE for listing, status changes, and deletion.
The dashboard host is protected by Cloudflare Access. Access supplies the cf-access-jwt-assertion header, and the Worker validates it through API Admin before serving stats. The dashboard does not start a separate GitHub OAuth flow or store browser tokens; the legacy OAuth routes were removed. CLI JWTs are rejected on dashboard endpoints outside local development, and dashboard data endpoints return 503 while the authentication backend is unavailable.
Webhook dispatch β
src/webhook-routes.ts verifies the request before parsing and dispatching it. A five-minute delivery marker prevents immediate GitHub redeliveries from double-running reviews or feedback handlers.
| GitHub event | Work |
|---|---|
pull_request | Review gates, lifecycle cleanup, merge feedback, vector refresh, label triggers |
issue_comment | Command registry and PR chat |
pull_request_review | Cross-reference updates and same-head dismissed-review feedback; stale-head dismissals are neutral |
pull_request_review_comment | Cross-reference updates and reply sentiment |
pull_request_review_thread | Resolved or reopened finding feedback |
merge_queue_entry | Merge-queue deterministic gate |
installation | created provisions a tenant named from the installation account login (recording the account id for rename-proof claim verification; redeliveries refresh the name/id on pending tenants only); deleted deactivates the tenant's installation and clears its vector index |
Bot-authored issue comments are ignored so help text and status comments cannot trigger command loops.
Review sequence β
sequenceDiagram
participant G as GitHub
participant W as Worker
participant Q as Cloudflare Queue
participant K as KV
participant A as AI provider
G->>W: pull_request webhook
W->>W: HMAC, config, eligibility gates
W->>K: record latest head and queue status
W->>G: queued batch comment
W->>Q: compact review job
W-->>G: 202 Accepted
Q->>K: in-flight and reviewed-SHA checks
Q->>G: progress comment and check run
Q->>G: fetch PR, diff, files, CI, search candidates
W->>W: filter paths, SAST, AST, symbols, graph, impacted tests
W->>A: primary review
A-->>W: findings and usage
W->>A: validator and optional custom checks
A-->>W: accepted findings and usage
W->>G: fetch linked issues when referenced
W->>A: optional linked-issue scope check
W->>W: dismissals, dedup, severity caps, noise controls, verdict
W->>G: walkthrough, inline review, labels, check conclusion
W->>K: findings, outcomes, usage log, last-reviewed SHA
W->>K: release in-flight stateGate and prepare β
handlePullRequest handles PR lifecycle events before a model call. It loads config using the repository installation, checks automatic-review policy, consumes the rate limit, clears superseded review state on new commits, and queues the review. Review-triggering lifecycle and label events return after the queue message and queued batch comment are written. If a queue binding is unavailable, the existing direct path remains as a fallback; a tenant cap denial returns without running a review instead of falling back.
After the gate accepts a new head, the Worker posts Preparing review context before it fetches the PR at its current head SHA, loads the filtered diff, and gathers bounded repository context. The same batch comment carries later progress and any terminal failure, so a preparation failure does not leave the PR silent. If GitHub refuses the whole diff for a large PR, the review code can build context from paginated file patches rather than treating the PR as empty.
Deterministic layer β
Built-in SAST and configured AST rules inspect added lines. SAST output is authoritative. The model sees a bounded severity-sorted summary to inspect related effects, but it is told not to repeat or dispute the deterministic result.
Deterministic prechecks inspect PR metadata and changed paths. Configured pre-merge checks can require tests, description length, docstrings, secret blocking, or model-evaluated repository rules.
Context ranking β
The Worker ranks changed files before collecting context. Small reviews favor files with secrets, public boundaries, persistence, and high-risk change paths. Large reviews rank the current chunk the same way, then fetch that batch at its exact head SHA. Each batch gets up to 12,000 context characters and an invocation gets up to 80,000; changed code is placed before imports and search matches. Missing paths are recorded in the review result. The Worker also derives a deterministic review plan for deleted behavior, public contracts, security boundaries, async failure semantics, and changed tests before it calls a model.
The Worker sends one complete prompt when the primary route's configured input budget can hold it and the rendered input remains within the 120,000-token latency limit. Otherwise, it splits the diff using the first available route's budget, so a large-context primary does not inherit the smaller fallback's setup overhead. Large-diff batches target 160,000 diff characters and remain bounded by the route budget, route timeouts, and review wall rather than an arbitrary call count. If a primary batch falls through to a smaller route, that batch is repacked at the fallback route's budget before the call. A successful batch is cached with its head and rendered context, while failed and terminally omitted batches remain explicit. A later retry sends only work that can still make progress. An oversized hunk is split at line boundaries with overlapping context; a line that cannot fit any route becomes a terminal coverage error instead of an unbounded retry.
For changed JavaScript and TypeScript, parsed declarations produce up to four GitHub search terms. Exact search paths come first; repository-scoped Vectorize results follow by score. GitHub code searches are paced once every two seconds per App installation, so a busy installation may defer lower-priority terms instead of exhausting GitHub's smaller search-rate bucket. The Worker fetches up to four of those candidate files and 24,000 characters as affected consumers. It then selects likely tests from search results plus filename matches in the first 5,000 repository-tree paths, again capped at four files and 24,000 characters. Declarations and the graph become prompt evidence after the files are fetched; they are not a separate ranking tier.
The same bounded graph supplies the walkthrough's architecture-impact prose and optional Mermaid diagram. Solid diagram edges are parsed imports; dashed edges are likely-test matches. The renderer uses stable path-derived IDs, caps output at 12 nodes and 12 edges, reports omitted relationships, and includes the analyzed head SHA plus a graph fingerprint. It emits no diagram for fewer than two relationships and never asks a model to write Mermaid syntax.
The context fetcher reads changed files in waves of four and processes each wave in the original ranked order. Exact search runs beside Vectorize, and consumer files run beside the repository-tree read. CI, learning data, conventions, external instructions, and review history load in one bounded parallel group. The changed-file context cap is reduced when the selected input-token budget is small, leaving room for the diff and system prompt. A prepared context snapshot is written before provider calls, so a retry can reuse GitHub reads after it rechecks the current config, PR metadata, and rendered prompts.
The queue path does not run a full cache probe inside the webhook request. A queued review checks the head-scoped context snapshot after it starts. A complete result cache is still available to explicit fast-path callers, while a missing or changed prompt falls through to normal context collection.
Vectorize supplies paths only. The Worker reads selected content from GitHub at the captured head SHA, so a semantic result never injects source text by itself.
Every GitHub read or write uses the installation attached to the repository. If an older queued job has no installation ID, the Worker resolves /repos/{owner}/{repo}/installation before it runs. It never chooses the first App installation. Findings keep the SHA that was analyzed; if the PR head changes before posting, the old check is marked superseded and a fresh-head review is queued.
Other prompt inputs may include:
- repo and directory
AGENTS.mdinstructions; - configured external repository files;
- CI check output and annotations;
- recent PR summaries;
- previous findings on this PR;
- feedback-derived false-positive and practice rules;
- branch stack context.
All inputs are bounded by file, byte, record, or time limits. The target branch's base SHA supplies .codebeaver.yaml for the review policy. PR metadata, diffs, source comments, repository context, CI output, history, and learned examples are untrusted data. They cannot change the review policy. Model findings need an exact quote from their added anchor line and a typed proof before publication.
Linked issues are handled later. After validation, the Worker fetches referenced issues and can run a separate chat-model scope check; issue text is not part of the primary review prompt.
When blind_first_review is enabled, the primary prompt omits the PR title and description. After proof validation, a bounded intent pass compares surviving findings with the PR text and up to three linked issues. It can mark a finding consistent, conflicting, or unknown. It cannot add findings, raise severity, or excuse security and contract failures. The result-cache fingerprint includes the mode.
When shadow_recall is enabled, its private model call starts after the public analysis is ready but does not hold the check run or final walkthrough open. Its telemetry and usage patch the matching review log after the call completes.
Model calls β
The service has six model purposes:
| Purpose | Default model source | Separate API credentials supported |
|---|---|---|
| primary | REVIEW_MODEL | yes, through primary provider settings |
| validator | primary | yes |
| autofix | deepseek/deepseek-v4-pro via OpenRouter, deepseek-v4-pro direct | yes |
| chat | primary | yes |
| describe | primary | no separate credential fields |
| changelog | primary | no separate credential fields |
Each paid response must include usage. The review record merges prompt, completion, reasoning, cached, and cache-write tokens plus provider cost and per-call line items. Direct DeepSeek does not return a billed amount, so the Worker estimates it from the configured model and DeepSeek's published rates.
Every model purpose can use an ordered endpoint chain. When REVIEWER_API_KEY_ZAI is configured, the production chain is Z.ai Coding (glm-5.3-flash), then direct DeepSeek for primary review, validation, autofix, chat, description, changelog, and other model-backed checks. Providers are declared in src/providers.ts; REVIEW_PROVIDER_ORDER overrides the default sequence. Provider denials such as invalid credentials, blocked models, or exhausted credits move to the next configured endpoint; transient failures also use the configured retry and fallback recovery. An exhausted chain reports the last denial. Malformed HTTP-success responses, empty output, and output-cap exhaustion get one bounded same-route recovery when the route has time left. Direct DeepSeek also retries high-reasoning empties and capped responses with thinking disabled. If every route returns one of those recoverable failures, the Worker records a recoverable failure and queues a new review instead of treating the provider responses as final. Validator and autofix chains use one deadline across their configured endpoints.
Large reviews keep the complete rendered prompt together whenever it fits the primary route's configured input budget. Chunking starts only when that hard budget is exceeded, or when recovery needs to retry incomplete work after a full request fails, returns empty output, or reaches its output cap. Chunk planning uses the first available route's input budget. Fallback routes with smaller input windows repack each batch to their lower limit. When the Worker rebuilds a diff through the pulls-files endpoint, only files with a GitHub patch body become chunks. The bounded prompt can show an excerpt, but chunk scheduling uses the complete set of those patches. Files without a patch body are disclosed in the review details instead of becoming empty model calls. Every chunk can try the full endpoint chain and receives its changed file plus a bounded slice of the fetched repository context, including related definitions, consumers, and likely tests; the 900-second review wall remains the final bound (a single chunk is capped at 480 seconds). If repository instructions consume most of the chunk budget, the Worker keeps a bounded beginning and end of those instructions so a small diff stays grouped instead of becoming one batch per changed line. A partial review records provider failures, deadline omissions, and chunks that exceed every route's input budget as separate counts. It fails the required check and requests another review without marking that commit as complete. Successful chunk results stay in the chunk cache, so the retry calls providers only for work that did not complete. The review log and walkthrough include the current-path batches in each partial state.
Validation and filters β
The second pass receives findings, repository rules, and a bounded selection of fetched context for the files it is checking. After it returns, the Worker:
- removes malformed or non-English output;
- suppresses findings already covered by CI lint output;
- suppresses guarded null-check false positives;
- suppresses HTTPS-only findings for explicit loopback development URLs;
- applies per-PR dismissals;
- deduplicates multiple findings at the same line;
- caps configured paths without downgrading critical findings;
- applies minimum severity, quiet mode, and count limits;
- removes locations already posted for the same head SHA.
Each validator candidate has an ephemeral validation_id. The validator must return that ID for a confirmed candidate, so a wording edit cannot make a real finding disappear during reconciliation. Validator prompts cap their estimated input at 240,000 tokens; oversized hunks are trimmed around the reported line before the request. If every validator route fails or returns invalid output, the review ends as comment with an incomplete result, keeps warnings and critical findings for the retry, and may post preliminary suggestions that already have a changed-line hunk and concrete evidence. The walkthrough labels those suggestions unvalidated, the check concludes as neutral, and the retry is still required.
The derived verdict is approve, comment, or request_changes. src/github.ts maps it to success, neutral, or failure using the check-run config and pre-merge blockers.
When those filters remove every model candidate, the walkthrough says so instead of repeating a raw model summary that still claims findings remain. Deterministic SAST output stays separate in its own section.
Review logs retain counts for each filter stage, suggestions held during a validator outage, verified and downgraded code replacements, the active finding cap, and the number of uncovered diff chunks. A code replacement is committable only after its exact old lines, range, path, head SHA, and changed hunk match. A failed check removes the code block but keeps the prose finding. The dashboard shows those counts with a private recall sample when one is enabled; that sample is for measuring missed findings and does not post feedback or alter a verdict.
Posting and state β
The Worker posts or updates a single status comment, creates inline review comments, writes the GitHub review state, updates the check run, suggests or assigns reviewers, applies configured labels, and may update the PR description.
Later reviews compare normalized findings with stored state. A finding can be:
- open on the current head;
- not reported because it disappeared on a later review;
- verified fixed only after a successful autofix path;
- manually resolved through a GitHub review thread.
Manual resolution is not treated as proof of a code fix, so the dashboard may count a finding as both resolved and later reported as not reported or verified fixed.
State in KV β
REVIEW_CACHE holds bounded operational and product state. Common prefixes:
| Prefix | Contents |
|---|---|
in-flight: | Review lock, head SHA, status comment, check run, heartbeat, latest phase, review-run ID |
queued-review:v1: | Latest requested head and invocation while a review is queued or active |
review-job-status:v1: | Retained queue lifecycle record for one invocation |
last-reviewed: | Last completed head SHA for deduplication |
findings: | Current PR finding set and GitHub review identifiers |
finding-outcome:v1: | Per-finding open, not-reported, verified-fixed, and resolved records |
feedback: | Human feedback keyed by stable file-and-issue finding identity; legacy line-keyed entries remain readable |
dismissed-findings: | Per-PR hard suppression signal |
known-false-positives: | Learned false-positive patterns |
auto-best-practices: | Learned accepted practice rules |
accuracy: | Unified per-category feedback telemetry (shown/accepted/rejected) with processed markers; the separate fp-stats: keys are legacy read-only |
review-log: | Reverse-time review, usage, model, total/model/posting/finalization timing, cache, and bounded failure-stage/detail records |
review-result:review-result-v4: | Completed model-analysis cache, reusable after an unchanged branch update |
review-chunk:v1: | Successful large-diff chunk cache |
autofix-patch:v1: | Parsed exact-replacement patch cache |
orphan-queue: | Reviews waiting for recovery |
github-app-rate-limit:v1: | Installation-wide GitHub API backoff until GitHub's reset time |
pause: / ignore: / cancel: | Repository or PR control state |
pr-history: | Recent repository review summaries |
tenant:v1: / tenant-installation:v1: | Tenant records and installationβtenant mapping |
tenant-reviews:v1: | Per-tenant monthly review counters for plan caps |
content-cache:v1: | Immutable blobs keyed by repository and commit SHA: changed/consumer/test file contents, repository trees, and repo instruction files; 30-day TTL |
embedding-cache:v1: | Semantic search query embeddings keyed by query hash; 30-day TTL |
vector-index-manifest:v1: | Indexed paths per repository, used to delete vectors on app uninstall |
Durable Object WebhookDeliveryDedup | Webhook delivery claim, renewal, and completion state |
Durable Object RepoMemory | Per-repository feedback entries, learning rules, accuracy counters, PR history, and dismissed findings when the binding is enabled; transactional folds replace the racy KV read-modify-write paths, which remain as fallback |
Prepared context, completed analysis, and chunk caches expire after 14 days. The completed-analysis cache is scoped to one PR, its stable review inputs, and deterministic scanner policy; the stored base SHA is audit data. Webhook deliveries enqueue before any full cache probe. An exact, complete result remains available to explicit fast-path callers, while a changed prompt, context, scanner policy, partial analysis, or incomplete validator result goes through the normal queue. Autofix patch entries expire after 21 days. Finding outcomes and learning rules use 90-day retention. Other operational markers have shorter TTLs suited to locking, deduplication, or recovery.
Semantic index isolation β
Repository vector namespaces are SHA-256 hashes of the owner and repository. Vector IDs are 64-byte-safe hashes derived from repository, path, chunk, and content identity. Search queries use that namespace, so results cannot cross repositories even without a metadata index.
Indexable files are filtered for generated, binary, secret-bearing, and unsupported paths. Merged PR handling reindexes changed paths and deletes vectors for removed or renamed paths.
Learning loop β
Feedback records enter processFeedbackLearning, which updates category accuracy and proposes or refreshes repository rules. Rules store their newest real feedback timestamp. Reprocessing old data cannot keep a stale rule alive.
Scheduled cleanup scans installed repositories, drops invalid or 90-day-stale rules, merges duplicate patterns, and caps each learned list at 50. Dashboard and CLI operators can approve, reject, or delete rules.
Scheduled recovery β
The full cleanup cron runs every five minutes. A separate one-minute tick refreshes queued status comments and dispatches due KV-backed retries. scheduledCleanup:
- scans stale checks on open PRs plus the newest page of closed PRs; closed PRs get a terminal check update and never a replacement review;
- dispatches due KV-backed retries globally, then drains the selected orphan queues with awaited review calls;
- removes expired or inconsistent state;
- compacts learned rules;
- keeps review status from being stranded after a Worker termination or deploy.
The direct KV scan runs every minute. Due retry entries are dispatched before the repository rotation, so a retry does not wait for its repository to be selected or for the five-minute cleanup interval. If GitHub cannot finalize an abandoned check run, recovery keeps the marker and retries later rather than starting a duplicate review. The slower marker-missing scan checks five repositories every 30 minutes so it cannot consume the GitHub App rate limit used by live reviews.
An in-flight review marker is considered stale after 20 minutes. This is longer than the 15-minute Queue consumer wall because provider calls can hold an awaited network operation past the interval heartbeat; cleanup must not cancel that live review and start a duplicate. AI analysis stops at 12 minutes, leaving one minute for context preparation and two minutes for GitHub posting, check finalization, and recovery-state cleanup.
REVIEW_JOBS has six concurrent consumer invocations. REVIEW_PRIORITY_JOBS has two reserved invocations for PRs in codebeaver/codebeaver, so reviewer changes are not delayed by the normal backlog. Each invocation handles one message. Large reviews can still process their own independent file chunks in parallel, with the provider-call cap kept separate from queue concurrency.
Reviews also pass through the GITHUB_RATE_GOVERNOR Durable Object, keyed by GitHub App installation. Production starts four review leases and can ramp to six after clean intervals. A non-search GitHub rate-limit result halves the active lane capacity and keeps the installation in cooldown. Repository code search uses a separate two-request lane and its own cooldown, so search bursts cannot consume every review lease or pause normal GitHub reads.
When GitHub reports a non-search rate limit, the Worker stores the reset time for that installation. Queue consumers, orphan recovery, and scheduled repository scans defer their work until then. A busy installation pauses together instead of each PR burning another failed API call.
Manual review requests for an active head are ignored, except when the queued head is retrying: that request replaces the delayed retry even if its retained in-flight marker is fresh. A request for a newer head is stored separately; the finishing run releases it, and scheduled cleanup requeues it if a timing race leaves it behind. Every manual request clears a prior cancellation marker before this active-head check, so a retry after /review cancel cannot be queued with cancellation still set.
Cloudflare CPU limits do not measure time spent waiting on network I/O. External requests still use explicit abort timers so a hung provider or GitHub call cannot hold an invocation forever.
Trust boundaries β
- Webhook bodies are untrusted until HMAC verification passes.
- Diff text, PR metadata, issue content, repository instructions, and external context are data, not executable instructions.
- GitHub reads and writes use the target repository's App installation.
- Logs redact and bound metadata; credentials and raw private material must never enter log fields.
- Debug routes use a separate bearer token and optional repository allowlist.
- The dashboard accepts only sessions validated by API Admin.
- Autofix refuses fork writes and dangerous paths, then requires repository-owned verification before opening a PR.
Known boundaries β
CodeBeaver supports GitHub only. Semantic context stays inside one repository. Stack awareness describes the current head/base relationship but does not maintain a full cross-repository dependency graph. Published model and deterministic findings share a versioned stored record with origin, validation state, head SHA, and evidence references. The authenticated review-detail endpoint returns that JSON and emits SARIF 2.1.0 with ?format=sarif; unvalidated model output never enters either export.
Review jobs and live progress β
Webhook-triggered reviews are sent to the codebeaver-jobs Cloudflare Queue. The consumer handles one message at a time and can run beyond the 30-second post-response lifetime available to waitUntil(). Queue messages contain only repository coordinates, installation ID, PR number, head SHA, operation, and an invocation ID. They never carry tokens or diff text.
The review comment is a batch keyed by PR head SHA. It is created before the model work starts and edited as phases move forward. The worker sends a comment heartbeat every 30 seconds after work begins; it does not create heartbeat comments. The body shows a compact stage trail plus recalculated remaining time and UTC ETA. A check run receives phase and measured-count updates, while a plain heartbeat only edits the batch comment.
Long comment commands use the isolated codebeaver-commands queue when that binding is present. Their queue message carries an invocation ID and GitHub coordinates only. Question text, recipe selection, and other command inputs are held in operation-status:v1:* in KV, where the consumer reads them just before execution. Each command begins with one replaceable operation comment so an interrupted job can report its retry or failure state without starting a new timeline thread. It says Waiting to start while queued, then receives the same live stage, remaining-time, UTC ETA, and terminal updates. Successful command sessions also persist total and per-stage durations under operation-timing:v1:*, filtered by operation, size bucket, and file-count band. The consumer requires five matches before using that history. When the queue binding is unavailable, the direct command path starts the same session and comment before invoking the handler, so failures can stop the heartbeat and replace that comment with a terminal result.
Queued reviews use their own compact KV index, scoped to the GitHub App installation. Before the consumer starts, the batch comment reports an approximate count of queued reviews ahead, a coarse start estimate, and Queue checked in UTC. Reviewer PRs show that they are in the priority lane. A separate once-a-minute cron refreshes waiting review comments in batches of at most 50. The number is an estimate: consumer concurrency, retries, newer heads, and rate-limit backoff can change execution order.
If a queued review has already made an AI attempt and is waiting for a retry, closing or merging the PR cancels only that retry. The existing retry comment and provider details remain visible, with a note explaining why the retry stopped. A review that never started keeps the shorter queued-cancellation message.
Unchanged reviews and archived batches β
Completed clean approvals can reuse analysis when the PR's review-input fingerprint matches, including its full diff, repository tree, prompt, context, model, and scanner policy. A failed tree lookup limits reuse to the exact commit. An empty result becomes reusable only after publication and check finalization succeed. Partial reviews, validator outages, and failed pre-merge checks do not qualify. Current deterministic and pre-merge checks still run on reuse. A cached rerun on the same head reuses the existing GitHub review only when its commit, state, and body match; a dismissed approval requires a new review. Older batch summaries show an outdated notice and link to the current review. Repeated archival preserves the original body. Outdated does not mean fixed. Partial and validator-incomplete runs preserve prior findings and blocking reviews; they cannot mark missing findings resolved. Cached reruns recover missing review IDs from GitHub before posting another approval.