Skip to content

Architecture ​

CodeBeaver has three user-facing programs backed by one Cloudflare Worker:

  • a GitHub App webhook service;
  • the @codebeaver/cli terminal package;
  • an authenticated React dashboard.

The Worker owns GitHub installation auth, review orchestration, model calls, deterministic checks, persisted review state, learning, CLI APIs, and dashboard APIs.

System map ​

mermaid
flowchart LR
    GH["GitHub App webhooks"] --> W["Cloudflare Worker"]
    CLI["CodeBeaver CLI"] --> W
    DASH["Dashboard SPA"] --> W
    ADMIN["API Admin session service"] --> DASH
    W --> GHA["GitHub REST and GraphQL APIs"]
    W --> MODEL["OpenRouter-compatible models"]
    W --> KV["Cloudflare KV: REVIEW_CACHE"]
    W --> VEC["Cloudflare Vectorize"]
    W --> EMBED["Workers AI embeddings"]

The dashboard is a static single-page app. Cloudflare Access owns the browser session; the app stores only identity-scoped query data and calls the Worker for data. The Worker validates the Access assertion through API Admin before dashboard queries or cache hydration begin. Production calls API Admin through the API_ADMIN service binding.

Workspace ​

text
src/
  index.ts                 Worker entry, shared routing, PR event gates
  webhook-routes.ts        HMAC verification and webhook dispatch
  api-routes.ts            CLI, auth, and dashboard API routing
  review.ts                Review orchestration and context gathering
  ai.ts                    Provider calls, usage, retries, validation
  findings.ts              Built-in SAST and persisted finding state
  config.ts                Org and repository config loading
  learning.ts              Feedback ingestion and learned rules
  walkthrough.ts           GitHub output formatting
  github.ts                GitHub auth and API helpers
  cleanup.ts               Scheduled cleanup and orphan recovery
  commands/                Pull-request comment handlers
  codebase-context.ts      CLI indexing and search endpoints
  vector-context.ts        Vector IDs, chunks, search, merge refresh
  symbol-context.ts        Static declarations, imports, and code graph
  structural-impact.ts     Affected consumers and tests
  stats-api.ts             Dashboard data endpoints
  logger.ts                Structured logs and request correlation

packages/cli/
  src/index.ts             Command registry and output formatting
  src/config.ts            OAuth token and user config storage
  src/local-review.ts      Direct provider mode
  src/finding-bundle.ts    Stable v1 bundle and agent handoff
  skills/SKILL.md          Packaged agent instructions

apps/dashboard/
  src/routes/              Login, overview, reviews, users, learning
  src/lib/api.ts           Worker API client and response types
  src/lib/query-client.ts  Persisted TanStack Query cache

Worker routing ​

The default export in src/index.ts wraps both scheduled and HTTP handlers with request metadata. Every HTTP request receives a request ID. Unhandled failures return a generic JSON error containing that ID.

Public and operator routes ​

RouteMethodAuthPurpose
/GETnoneRedirect to dashboard login
/ and unmatched POST pathsPOSTGitHub HMACWebhook receiver
/healthGETnoneVerify required secrets and GitHub App JWT creation; report deployed SHA and resolved review-route hosts
/diag?repo=&pr=GETdebug bearer tokenRead-only installation and PR diagnostic
/test-review?repo=&pr=GETdebug bearer tokenRun a live forced review with GitHub side effects

DEBUG_ALLOWED_REPOS can limit the two debug routes. DEBUG_INSTALL_ID selects the GitHub App installation they use.

CLI routes ​

All CLI routes are POST requests under /api/cli/:

text
auth/exchange       auth/refresh
review              describe             improve
changelog           test                 autofix
ask                 coverage             sarif
index-repo          search-context
learning            learning/delete
learning/approve    learning/reject

Review, code generation, indexing, and repository questions validate the CLI JWT and then resolve the GitHub App installation for the named repository. Local-only CLI commands do not call these routes.

Dashboard routes ​

text
/api/auth/session
/api/stats
/api/stats/reviews
/api/stats/users
/api/stats/review/:owner/:repo/:pr
/api/stats/learning/rules

Stats routes require an admin session. The learning-rules endpoint uses GET, PATCH, and DELETE for listing, status changes, and deletion.

The dashboard host is protected by Cloudflare Access. Access supplies the cf-access-jwt-assertion header, and the Worker validates it through API Admin before serving stats. The dashboard does not start a separate GitHub OAuth flow or store browser tokens; the legacy OAuth routes were removed. CLI JWTs are rejected on dashboard endpoints outside local development, and dashboard data endpoints return 503 while the authentication backend is unavailable.

Webhook dispatch ​

src/webhook-routes.ts verifies the request before parsing and dispatching it. A five-minute delivery marker prevents immediate GitHub redeliveries from double-running reviews or feedback handlers.

GitHub eventWork
pull_requestReview gates, lifecycle cleanup, merge feedback, vector refresh, label triggers
issue_commentCommand registry and PR chat
pull_request_reviewCross-reference updates and same-head dismissed-review feedback; stale-head dismissals are neutral
pull_request_review_commentCross-reference updates and reply sentiment
pull_request_review_threadResolved or reopened finding feedback
merge_queue_entryMerge-queue deterministic gate
installationcreated provisions a tenant named from the installation account login (recording the account id for rename-proof claim verification; redeliveries refresh the name/id on pending tenants only); deleted deactivates the tenant's installation and clears its vector index

Bot-authored issue comments are ignored so help text and status comments cannot trigger command loops.

Review sequence ​

mermaid
sequenceDiagram
    participant G as GitHub
    participant W as Worker
    participant Q as Cloudflare Queue
    participant K as KV
    participant A as AI provider

    G->>W: pull_request webhook
    W->>W: HMAC, config, eligibility gates
    W->>K: record latest head and queue status
    W->>G: queued batch comment
    W->>Q: compact review job
    W-->>G: 202 Accepted
    Q->>K: in-flight and reviewed-SHA checks
    Q->>G: progress comment and check run
    Q->>G: fetch PR, diff, files, CI, search candidates
    W->>W: filter paths, SAST, AST, symbols, graph, impacted tests
    W->>A: primary review
    A-->>W: findings and usage
    W->>A: validator and optional custom checks
    A-->>W: accepted findings and usage
    W->>G: fetch linked issues when referenced
    W->>A: optional linked-issue scope check
    W->>W: dismissals, dedup, severity caps, noise controls, verdict
    W->>G: walkthrough, inline review, labels, check conclusion
    W->>K: findings, outcomes, usage log, last-reviewed SHA
    W->>K: release in-flight state

Gate and prepare ​

handlePullRequest handles PR lifecycle events before a model call. It loads config using the repository installation, checks automatic-review policy, consumes the rate limit, clears superseded review state on new commits, and queues the review. Review-triggering lifecycle and label events return after the queue message and queued batch comment are written. If a queue binding is unavailable, the existing direct path remains as a fallback; a tenant cap denial returns without running a review instead of falling back.

After the gate accepts a new head, the Worker posts Preparing review context before it fetches the PR at its current head SHA, loads the filtered diff, and gathers bounded repository context. The same batch comment carries later progress and any terminal failure, so a preparation failure does not leave the PR silent. If GitHub refuses the whole diff for a large PR, the review code can build context from paginated file patches rather than treating the PR as empty.

Deterministic layer ​

Built-in SAST and configured AST rules inspect added lines. SAST output is authoritative. The model sees a bounded severity-sorted summary to inspect related effects, but it is told not to repeat or dispute the deterministic result.

Deterministic prechecks inspect PR metadata and changed paths. Configured pre-merge checks can require tests, description length, docstrings, secret blocking, or model-evaluated repository rules.

Context ranking ​

The Worker ranks changed files before collecting context. Small reviews favor files with secrets, public boundaries, persistence, and high-risk change paths. Large reviews rank the current chunk the same way, then fetch that batch at its exact head SHA. Each batch gets up to 12,000 context characters and an invocation gets up to 80,000; changed code is placed before imports and search matches. Missing paths are recorded in the review result. The Worker also derives a deterministic review plan for deleted behavior, public contracts, security boundaries, async failure semantics, and changed tests before it calls a model.

The Worker sends one complete prompt when the primary route's configured input budget can hold it and the rendered input remains within the 120,000-token latency limit. Otherwise, it splits the diff using the first available route's budget, so a large-context primary does not inherit the smaller fallback's setup overhead. Large-diff batches target 160,000 diff characters and remain bounded by the route budget, route timeouts, and review wall rather than an arbitrary call count. If a primary batch falls through to a smaller route, that batch is repacked at the fallback route's budget before the call. A successful batch is cached with its head and rendered context, while failed and terminally omitted batches remain explicit. A later retry sends only work that can still make progress. An oversized hunk is split at line boundaries with overlapping context; a line that cannot fit any route becomes a terminal coverage error instead of an unbounded retry.

For changed JavaScript and TypeScript, parsed declarations produce up to four GitHub search terms. Exact search paths come first; repository-scoped Vectorize results follow by score. GitHub code searches are paced once every two seconds per App installation, so a busy installation may defer lower-priority terms instead of exhausting GitHub's smaller search-rate bucket. The Worker fetches up to four of those candidate files and 24,000 characters as affected consumers. It then selects likely tests from search results plus filename matches in the first 5,000 repository-tree paths, again capped at four files and 24,000 characters. Declarations and the graph become prompt evidence after the files are fetched; they are not a separate ranking tier.

The same bounded graph supplies the walkthrough's architecture-impact prose and optional Mermaid diagram. Solid diagram edges are parsed imports; dashed edges are likely-test matches. The renderer uses stable path-derived IDs, caps output at 12 nodes and 12 edges, reports omitted relationships, and includes the analyzed head SHA plus a graph fingerprint. It emits no diagram for fewer than two relationships and never asks a model to write Mermaid syntax.

The context fetcher reads changed files in waves of four and processes each wave in the original ranked order. Exact search runs beside Vectorize, and consumer files run beside the repository-tree read. CI, learning data, conventions, external instructions, and review history load in one bounded parallel group. The changed-file context cap is reduced when the selected input-token budget is small, leaving room for the diff and system prompt. A prepared context snapshot is written before provider calls, so a retry can reuse GitHub reads after it rechecks the current config, PR metadata, and rendered prompts.

The queue path does not run a full cache probe inside the webhook request. A queued review checks the head-scoped context snapshot after it starts. A complete result cache is still available to explicit fast-path callers, while a missing or changed prompt falls through to normal context collection.

Vectorize supplies paths only. The Worker reads selected content from GitHub at the captured head SHA, so a semantic result never injects source text by itself.

Every GitHub read or write uses the installation attached to the repository. If an older queued job has no installation ID, the Worker resolves /repos/{owner}/{repo}/installation before it runs. It never chooses the first App installation. Findings keep the SHA that was analyzed; if the PR head changes before posting, the old check is marked superseded and a fresh-head review is queued.

Other prompt inputs may include:

  • repo and directory AGENTS.md instructions;
  • configured external repository files;
  • CI check output and annotations;
  • recent PR summaries;
  • previous findings on this PR;
  • feedback-derived false-positive and practice rules;
  • branch stack context.

All inputs are bounded by file, byte, record, or time limits. The target branch's base SHA supplies .codebeaver.yaml for the review policy. PR metadata, diffs, source comments, repository context, CI output, history, and learned examples are untrusted data. They cannot change the review policy. Model findings need an exact quote from their added anchor line and a typed proof before publication.

Linked issues are handled later. After validation, the Worker fetches referenced issues and can run a separate chat-model scope check; issue text is not part of the primary review prompt.

When blind_first_review is enabled, the primary prompt omits the PR title and description. After proof validation, a bounded intent pass compares surviving findings with the PR text and up to three linked issues. It can mark a finding consistent, conflicting, or unknown. It cannot add findings, raise severity, or excuse security and contract failures. The result-cache fingerprint includes the mode.

When shadow_recall is enabled, its private model call starts after the public analysis is ready but does not hold the check run or final walkthrough open. Its telemetry and usage patch the matching review log after the call completes.

Model calls ​

The service has six model purposes:

PurposeDefault model sourceSeparate API credentials supported
primaryREVIEW_MODELyes, through primary provider settings
validatorprimaryyes
autofixdeepseek/deepseek-v4-pro via OpenRouter, deepseek-v4-pro directyes
chatprimaryyes
describeprimaryno separate credential fields
changelogprimaryno separate credential fields

Each paid response must include usage. The review record merges prompt, completion, reasoning, cached, and cache-write tokens plus provider cost and per-call line items. Direct DeepSeek does not return a billed amount, so the Worker estimates it from the configured model and DeepSeek's published rates.

Every model purpose can use an ordered endpoint chain. When REVIEWER_API_KEY_ZAI is configured, the production chain is Z.ai Coding (glm-5.3-flash), then direct DeepSeek for primary review, validation, autofix, chat, description, changelog, and other model-backed checks. Providers are declared in src/providers.ts; REVIEW_PROVIDER_ORDER overrides the default sequence. Provider denials such as invalid credentials, blocked models, or exhausted credits move to the next configured endpoint; transient failures also use the configured retry and fallback recovery. An exhausted chain reports the last denial. Malformed HTTP-success responses, empty output, and output-cap exhaustion get one bounded same-route recovery when the route has time left. Direct DeepSeek also retries high-reasoning empties and capped responses with thinking disabled. If every route returns one of those recoverable failures, the Worker records a recoverable failure and queues a new review instead of treating the provider responses as final. Validator and autofix chains use one deadline across their configured endpoints.

Large reviews keep the complete rendered prompt together whenever it fits the primary route's configured input budget. Chunking starts only when that hard budget is exceeded, or when recovery needs to retry incomplete work after a full request fails, returns empty output, or reaches its output cap. Chunk planning uses the first available route's input budget. Fallback routes with smaller input windows repack each batch to their lower limit. When the Worker rebuilds a diff through the pulls-files endpoint, only files with a GitHub patch body become chunks. The bounded prompt can show an excerpt, but chunk scheduling uses the complete set of those patches. Files without a patch body are disclosed in the review details instead of becoming empty model calls. Every chunk can try the full endpoint chain and receives its changed file plus a bounded slice of the fetched repository context, including related definitions, consumers, and likely tests; the 900-second review wall remains the final bound (a single chunk is capped at 480 seconds). If repository instructions consume most of the chunk budget, the Worker keeps a bounded beginning and end of those instructions so a small diff stays grouped instead of becoming one batch per changed line. A partial review records provider failures, deadline omissions, and chunks that exceed every route's input budget as separate counts. It fails the required check and requests another review without marking that commit as complete. Successful chunk results stay in the chunk cache, so the retry calls providers only for work that did not complete. The review log and walkthrough include the current-path batches in each partial state.

Validation and filters ​

The second pass receives findings, repository rules, and a bounded selection of fetched context for the files it is checking. After it returns, the Worker:

  • removes malformed or non-English output;
  • suppresses findings already covered by CI lint output;
  • suppresses guarded null-check false positives;
  • suppresses HTTPS-only findings for explicit loopback development URLs;
  • applies per-PR dismissals;
  • deduplicates multiple findings at the same line;
  • caps configured paths without downgrading critical findings;
  • applies minimum severity, quiet mode, and count limits;
  • removes locations already posted for the same head SHA.

Each validator candidate has an ephemeral validation_id. The validator must return that ID for a confirmed candidate, so a wording edit cannot make a real finding disappear during reconciliation. Validator prompts cap their estimated input at 240,000 tokens; oversized hunks are trimmed around the reported line before the request. If every validator route fails or returns invalid output, the review ends as comment with an incomplete result, keeps warnings and critical findings for the retry, and may post preliminary suggestions that already have a changed-line hunk and concrete evidence. The walkthrough labels those suggestions unvalidated, the check concludes as neutral, and the retry is still required.

The derived verdict is approve, comment, or request_changes. src/github.ts maps it to success, neutral, or failure using the check-run config and pre-merge blockers.

When those filters remove every model candidate, the walkthrough says so instead of repeating a raw model summary that still claims findings remain. Deterministic SAST output stays separate in its own section.

Review logs retain counts for each filter stage, suggestions held during a validator outage, verified and downgraded code replacements, the active finding cap, and the number of uncovered diff chunks. A code replacement is committable only after its exact old lines, range, path, head SHA, and changed hunk match. A failed check removes the code block but keeps the prose finding. The dashboard shows those counts with a private recall sample when one is enabled; that sample is for measuring missed findings and does not post feedback or alter a verdict.

Posting and state ​

The Worker posts or updates a single status comment, creates inline review comments, writes the GitHub review state, updates the check run, suggests or assigns reviewers, applies configured labels, and may update the PR description.

Later reviews compare normalized findings with stored state. A finding can be:

  • open on the current head;
  • not reported because it disappeared on a later review;
  • verified fixed only after a successful autofix path;
  • manually resolved through a GitHub review thread.

Manual resolution is not treated as proof of a code fix, so the dashboard may count a finding as both resolved and later reported as not reported or verified fixed.

State in KV ​

REVIEW_CACHE holds bounded operational and product state. Common prefixes:

PrefixContents
in-flight:Review lock, head SHA, status comment, check run, heartbeat, latest phase, review-run ID
queued-review:v1:Latest requested head and invocation while a review is queued or active
review-job-status:v1:Retained queue lifecycle record for one invocation
last-reviewed:Last completed head SHA for deduplication
findings:Current PR finding set and GitHub review identifiers
finding-outcome:v1:Per-finding open, not-reported, verified-fixed, and resolved records
feedback:Human feedback keyed by stable file-and-issue finding identity; legacy line-keyed entries remain readable
dismissed-findings:Per-PR hard suppression signal
known-false-positives:Learned false-positive patterns
auto-best-practices:Learned accepted practice rules
accuracy:Unified per-category feedback telemetry (shown/accepted/rejected) with processed markers; the separate fp-stats: keys are legacy read-only
review-log:Reverse-time review, usage, model, total/model/posting/finalization timing, cache, and bounded failure-stage/detail records
review-result:review-result-v4:Completed model-analysis cache, reusable after an unchanged branch update
review-chunk:v1:Successful large-diff chunk cache
autofix-patch:v1:Parsed exact-replacement patch cache
orphan-queue:Reviews waiting for recovery
github-app-rate-limit:v1:Installation-wide GitHub API backoff until GitHub's reset time
pause: / ignore: / cancel:Repository or PR control state
pr-history:Recent repository review summaries
tenant:v1: / tenant-installation:v1:Tenant records and installation→tenant mapping
tenant-reviews:v1:Per-tenant monthly review counters for plan caps
content-cache:v1:Immutable blobs keyed by repository and commit SHA: changed/consumer/test file contents, repository trees, and repo instruction files; 30-day TTL
embedding-cache:v1:Semantic search query embeddings keyed by query hash; 30-day TTL
vector-index-manifest:v1:Indexed paths per repository, used to delete vectors on app uninstall
Durable Object WebhookDeliveryDedupWebhook delivery claim, renewal, and completion state
Durable Object RepoMemoryPer-repository feedback entries, learning rules, accuracy counters, PR history, and dismissed findings when the binding is enabled; transactional folds replace the racy KV read-modify-write paths, which remain as fallback

Prepared context, completed analysis, and chunk caches expire after 14 days. The completed-analysis cache is scoped to one PR, its stable review inputs, and deterministic scanner policy; the stored base SHA is audit data. Webhook deliveries enqueue before any full cache probe. An exact, complete result remains available to explicit fast-path callers, while a changed prompt, context, scanner policy, partial analysis, or incomplete validator result goes through the normal queue. Autofix patch entries expire after 21 days. Finding outcomes and learning rules use 90-day retention. Other operational markers have shorter TTLs suited to locking, deduplication, or recovery.

Semantic index isolation ​

Repository vector namespaces are SHA-256 hashes of the owner and repository. Vector IDs are 64-byte-safe hashes derived from repository, path, chunk, and content identity. Search queries use that namespace, so results cannot cross repositories even without a metadata index.

Indexable files are filtered for generated, binary, secret-bearing, and unsupported paths. Merged PR handling reindexes changed paths and deletes vectors for removed or renamed paths.

Learning loop ​

Feedback records enter processFeedbackLearning, which updates category accuracy and proposes or refreshes repository rules. Rules store their newest real feedback timestamp. Reprocessing old data cannot keep a stale rule alive.

Scheduled cleanup scans installed repositories, drops invalid or 90-day-stale rules, merges duplicate patterns, and caps each learned list at 50. Dashboard and CLI operators can approve, reject, or delete rules.

Scheduled recovery ​

The full cleanup cron runs every five minutes. A separate one-minute tick refreshes queued status comments and dispatches due KV-backed retries. scheduledCleanup:

  • scans stale checks on open PRs plus the newest page of closed PRs; closed PRs get a terminal check update and never a replacement review;
  • dispatches due KV-backed retries globally, then drains the selected orphan queues with awaited review calls;
  • removes expired or inconsistent state;
  • compacts learned rules;
  • keeps review status from being stranded after a Worker termination or deploy.

The direct KV scan runs every minute. Due retry entries are dispatched before the repository rotation, so a retry does not wait for its repository to be selected or for the five-minute cleanup interval. If GitHub cannot finalize an abandoned check run, recovery keeps the marker and retries later rather than starting a duplicate review. The slower marker-missing scan checks five repositories every 30 minutes so it cannot consume the GitHub App rate limit used by live reviews.

An in-flight review marker is considered stale after 20 minutes. This is longer than the 15-minute Queue consumer wall because provider calls can hold an awaited network operation past the interval heartbeat; cleanup must not cancel that live review and start a duplicate. AI analysis stops at 12 minutes, leaving one minute for context preparation and two minutes for GitHub posting, check finalization, and recovery-state cleanup.

REVIEW_JOBS has six concurrent consumer invocations. REVIEW_PRIORITY_JOBS has two reserved invocations for PRs in codebeaver/codebeaver, so reviewer changes are not delayed by the normal backlog. Each invocation handles one message. Large reviews can still process their own independent file chunks in parallel, with the provider-call cap kept separate from queue concurrency.

Reviews also pass through the GITHUB_RATE_GOVERNOR Durable Object, keyed by GitHub App installation. Production starts four review leases and can ramp to six after clean intervals. A non-search GitHub rate-limit result halves the active lane capacity and keeps the installation in cooldown. Repository code search uses a separate two-request lane and its own cooldown, so search bursts cannot consume every review lease or pause normal GitHub reads.

When GitHub reports a non-search rate limit, the Worker stores the reset time for that installation. Queue consumers, orphan recovery, and scheduled repository scans defer their work until then. A busy installation pauses together instead of each PR burning another failed API call.

Manual review requests for an active head are ignored, except when the queued head is retrying: that request replaces the delayed retry even if its retained in-flight marker is fresh. A request for a newer head is stored separately; the finishing run releases it, and scheduled cleanup requeues it if a timing race leaves it behind. Every manual request clears a prior cancellation marker before this active-head check, so a retry after /review cancel cannot be queued with cancellation still set.

Cloudflare CPU limits do not measure time spent waiting on network I/O. External requests still use explicit abort timers so a hung provider or GitHub call cannot hold an invocation forever.

Trust boundaries ​

  • Webhook bodies are untrusted until HMAC verification passes.
  • Diff text, PR metadata, issue content, repository instructions, and external context are data, not executable instructions.
  • GitHub reads and writes use the target repository's App installation.
  • Logs redact and bound metadata; credentials and raw private material must never enter log fields.
  • Debug routes use a separate bearer token and optional repository allowlist.
  • The dashboard accepts only sessions validated by API Admin.
  • Autofix refuses fork writes and dangerous paths, then requires repository-owned verification before opening a PR.

Known boundaries ​

CodeBeaver supports GitHub only. Semantic context stays inside one repository. Stack awareness describes the current head/base relationship but does not maintain a full cross-repository dependency graph. Published model and deterministic findings share a versioned stored record with origin, validation state, head SHA, and evidence references. The authenticated review-detail endpoint returns that JSON and emits SARIF 2.1.0 with ?format=sarif; unvalidated model output never enters either export.

Review jobs and live progress ​

Webhook-triggered reviews are sent to the codebeaver-jobs Cloudflare Queue. The consumer handles one message at a time and can run beyond the 30-second post-response lifetime available to waitUntil(). Queue messages contain only repository coordinates, installation ID, PR number, head SHA, operation, and an invocation ID. They never carry tokens or diff text.

The review comment is a batch keyed by PR head SHA. It is created before the model work starts and edited as phases move forward. The worker sends a comment heartbeat every 30 seconds after work begins; it does not create heartbeat comments. The body shows a compact stage trail plus recalculated remaining time and UTC ETA. A check run receives phase and measured-count updates, while a plain heartbeat only edits the batch comment.

Long comment commands use the isolated codebeaver-commands queue when that binding is present. Their queue message carries an invocation ID and GitHub coordinates only. Question text, recipe selection, and other command inputs are held in operation-status:v1:* in KV, where the consumer reads them just before execution. Each command begins with one replaceable operation comment so an interrupted job can report its retry or failure state without starting a new timeline thread. It says Waiting to start while queued, then receives the same live stage, remaining-time, UTC ETA, and terminal updates. Successful command sessions also persist total and per-stage durations under operation-timing:v1:*, filtered by operation, size bucket, and file-count band. The consumer requires five matches before using that history. When the queue binding is unavailable, the direct command path starts the same session and comment before invoking the handler, so failures can stop the heartbeat and replace that comment with a terminal result.

Queued reviews use their own compact KV index, scoped to the GitHub App installation. Before the consumer starts, the batch comment reports an approximate count of queued reviews ahead, a coarse start estimate, and Queue checked in UTC. Reviewer PRs show that they are in the priority lane. A separate once-a-minute cron refreshes waiting review comments in batches of at most 50. The number is an estimate: consumer concurrency, retries, newer heads, and rate-limit backoff can change execution order.

If a queued review has already made an AI attempt and is waiting for a retry, closing or merging the PR cancels only that retry. The existing retry comment and provider details remain visible, with a note explaining why the retry stopped. A review that never started keeps the shorter queued-cancellation message.

Unchanged reviews and archived batches ​

Completed clean approvals can reuse analysis when the PR's review-input fingerprint matches, including its full diff, repository tree, prompt, context, model, and scanner policy. A failed tree lookup limits reuse to the exact commit. An empty result becomes reusable only after publication and check finalization succeed. Partial reviews, validator outages, and failed pre-merge checks do not qualify. Current deterministic and pre-merge checks still run on reuse. A cached rerun on the same head reuses the existing GitHub review only when its commit, state, and body match; a dismissed approval requires a new review. Older batch summaries show an outdated notice and link to the current review. Repeated archival preserves the original body. Outdated does not mean fixed. Partial and validator-incomplete runs preserve prior findings and blocking reviews; they cannot mark missing findings resolved. Cached reruns recover missing review IDs from GitHub before posting another approval.

Source-available under BUSL-1.1. Self-hosting is free for your organisation.