Operations â
This guide covers local development, configuration secrets, CI, production deployment, health checks, logs, recovery, and live evals.
Runtime targets â
| Service | Production URL | Source |
|---|---|---|
| GitHub webhook Worker | https://codebeaver.workers.dev | src/ |
| Staff dashboard and browser APIs | https://codebeaver.dev | apps/dashboard/, src/ |
| CLI package | @codebeaver/cli on GitHub Packages | packages/cli/ |
The Worker uses Cloudflare Workers, KV, Vectorize, and Workers AI. Model completion calls use an OpenRouter-compatible API.
Prerequisites â
Use the same major versions as CI:
- Node.js 22;
- pnpm 11;
- Wrangler 4;
- GitHub CLI for PR and workflow operations;
- Cloudflare credentials for deploy or remote bindings;
- production secret-manager access for secret work.
Install dependencies at the repository root:
pnpm install --frozen-lockfileThe pnpm workspace contains the Worker root, CLI package, and dashboard app.
Local checks â
Run the checks that match CI before pushing:
pnpm lint
pnpm typecheck
pnpm test
pnpm check:dashboardOr run both groups:
pnpm checkScripts:
| Script | Work |
|---|---|
pnpm lint | Biome checks Worker, eval, CLI, and dashboard source |
pnpm typecheck | Root TypeScript build check |
pnpm build | Wrangler dry-run build to dist/ |
pnpm test | Build Worker, then run Vitest |
pnpm check:worker | Lint, typecheck, build-backed tests |
pnpm check:dashboard | Dashboard typecheck, tests, and production build |
pnpm test:eval | Run eval suite with EVAL=1 |
pnpm test must build first. The Miniflare Worker suite imports dist/index.js.
Worker-core coverage gates are 65% statements, 60% branches, 65% functions, and 65% lines. The command wrappers under src/commands/ have focused tests; the gate covers the review, queue, webhook, learning, provider, and reporting modules.
Full-stack E2E probe â
node scripts/e2e-local.mjs boots the real dev stack â the Worker twice (once with ENVIRONMENT:local, once with ENVIRONMENT:production) and the Vite dashboard â and probes the dashboard API surface over HTTP the way a real client would: bearer session bootstrap, data-endpoint auth and shape contracts, handler-reach past the auth gate, the learning ruleIndex rejection, SARIF output, the legacy auth routes staying dead, production-mode fail-closed behavior, and the Vite /api proxy. It prints a PASS/FAIL line per probe and exits non-zero on any failure. Run it before merging changes to src/api-auth.ts, src/stats-api.ts, src/api-routes.ts, or anything under apps/dashboard/src/lib.
Prerequisites: pnpm install, and .dev.vars at the repo root with the boot secrets (REVIEW_MODEL, REVIEWER_API_KEY, GITHUB_APP_PRIVATE_KEY in PKCS#8 form wrapped in double quotes so the multiline value survives dotenv parsing, GITHUB_APP_WEBHOOK_SECRET). A PKCS#1 (BEGIN RSA PRIVATE KEY) key fails the health probe's JWT mint.
Flags: --worker-port (default 8811), --prod-port (8813), --dashboard-port (3311), --secret (local test JWT signing secret, shared with the booted Worker), --worker-secret (override the Worker's secret to exercise auth-failure paths), --skip-dashboard, --skip-production, --quiet. The script kills only the processes it starts; pick free ports when other dev servers are running.
Do not bypass Lefthook with --no-verify. Repository changes go through a feature branch, pull request, passing CI, automated review, and squash merge.
Run locally â
Start the Worker:
pnpm devStart the dashboard in another terminal:
pnpm --dir apps/dashboard devThe dashboard dev server listens on port 3001. VITE_CODEBEAVER_API_BASE_URL selects a Worker URL for local development. CODEBEAVER_DEV_API_TARGET selects the Vite /api proxy target (default http://127.0.0.1:8787), and CODEBEAVER_DEV_BEARER_TOKEN can inject a local-only bearer token. VITE_CLOUDFLARE_ACCESS_TEAM_DOMAIN names the Access team used by the production logout redirect.
Post-deploy dashboard smoke against the Access-protected origin needs long-lived Cloudflare Access service-token secrets in the GitHub production environment (synced from the operator's secret manager):
| Secret | Use |
|---|---|
CF_ACCESS_CLIENT_ID | CF-Access-Client-Id header for HTML and asset smoke |
CF_ACCESS_CLIENT_SECRET | CF-Access-Client-Secret header for HTML and asset smoke |
DASHBOARD_SMOKE_COOKIE | Optional CF_Authorization=<JWT> staff cookie for authenticated /api/auth/session and /api/stats only |
Create the Access service token with a Service Auth policy on the CodeBeaver Access app in your Cloudflare Zero Trust tenant. Until those secrets sync, the workflow can fall back to DASHBOARD_SMOKE_COOKIE for the Access edge. Never print credentials or response bodies.
The workflow gates the session/stats checks on DASHBOARD_SMOKE_COOKIE being set, not on it being valid, so an expired cookie left in the environment still fails the deploy. Once the service token is live, either refresh the cookie on the Access session schedule or remove it from the GitHub production environment to skip those checks. The workflow also prefers the service token whenever both CF_ACCESS_* secrets are present and does not fall back to the cookie if the token is rejected, so apply the Access Service Auth policy before or with the secret sync.
The CLI can target a local Worker:
codebeaver --url http://127.0.0.1:8787 doctorOnly loopback HTTP URLs are accepted. Other service URLs must use HTTPS.
Wrangler local development still needs usable bindings and secrets for any path you exercise. /health remains available when required Worker config is missing, but it reports degraded status.
Worker configuration â
Required strings â
Environment validation requires these values before ordinary requests run:
| Name | Use |
|---|---|
REVIEW_MODEL | Primary review model |
REVIEWER_API_KEY | Primary model provider token |
REVIEWER_API_KEY_ZAI | Dedicated Z.ai Coding token; enables the registry provider chain (Z.ai first, then direct DeepSeek) |
TYPESAFE_API_KEY | Optional TypeSafe Jev judgment-model key; enables reply-sentiment, intent-pass, and slop judgments |
OPENROUTER_MANAGEMENT_KEY | Optional OpenRouter management key used only for dashboard credit status |
DEEPSEEK_BILLING_API_KEY | Optional DeepSeek API key used only for dashboard balance status |
GITHUB_APP_ID | GitHub App numeric ID |
GITHUB_APP_PRIVATE_KEY | GitHub App RSA private key |
GITHUB_APP_WEBHOOK_SECRET | GitHub webhook HMAC secret |
CLI_JWT_SECRET | CLI session signing secret, at least 32 characters |
REVIEWER_API_BASE_URL is optional. When REVIEWER_API_KEY_ZAI is present, the Worker uses Z.ai Coding (glm-5.3-flash), then direct DeepSeek (deepseek-v4-flash). Providers are declared in src/providers.ts, and REVIEW_PROVIDER_ORDER reorders known providers. Without the Z.ai key, the legacy configurable chain remains available.
TYPESAFE_API_KEY is optional. When present, the Worker calls TypeSafe's Jev System One model (jev-1.13.0, https://api.typesafe.ai/v1/systemone) for four bounded judgments: reply sentiment on bot findings, the blind-first intent pass, slop scoring, and advisory evidence checks on published findings. Every Jev call fails open â unconfigured, timeout, HTTP error, or malformed payload falls back to the existing regex or validator-chain path, and the failure is logged as jev_request_failed.
The dashboard reads provider billing status from /api/provider-status. OpenRouter uses OPENROUTER_MANAGEMENT_KEY and DeepSeek uses DEEPSEEK_BILLING_API_KEY; neither credential is used for model calls or returned to clients. Successful reads are cached for 15 minutes.
CLI and dashboard auth â
| Name | Use |
|---|---|
CLI_GITHUB_ORG | Restrict CLI OAuth users to an org when set; the tenant-claim flow is exempt (it verifies installation-admin status instead) |
CLI_GITHUB_OAUTH_CLIENT_ID | GitHub OAuth app client ID |
GH_APP_CLIENT_SECRET | GitHub OAuth app client secret |
API_ADMIN_BASE_URL | Local/test dashboard session validator fallback; defaults to https://api-admin.codebeaver.dev |
ENVIRONMENT | Set to production so a missing API_ADMIN service binding returns 503 instead of using HTTP fallback |
Model overrides â
| Purpose | Model | API token | Base URL |
|---|---|---|---|
| first fallback | REVIEWER_FALLBACK_MODEL | REVIEWER_FALLBACK_API_KEY | REVIEWER_FALLBACK_API_BASE_URL |
| second fallback | REVIEWER_SECONDARY_FALLBACK_MODEL | REVIEWER_SECONDARY_FALLBACK_API_KEY | REVIEWER_SECONDARY_FALLBACK_API_BASE_URL |
| validator | REVIEW_MODEL_VALIDATOR | REVIEWER_API_KEY_VALIDATOR | REVIEWER_API_BASE_URL_VALIDATOR |
| validator fallback | REVIEWER_FALLBACK_MODEL_VALIDATOR | REVIEWER_FALLBACK_API_KEY_VALIDATOR | REVIEWER_FALLBACK_API_BASE_URL_VALIDATOR |
| validator second fallback | REVIEWER_SECONDARY_FALLBACK_MODEL_VALIDATOR | REVIEWER_SECONDARY_FALLBACK_API_KEY_VALIDATOR | REVIEWER_SECONDARY_FALLBACK_API_BASE_URL_VALIDATOR |
| autofix | REVIEW_MODEL_AUTOFIX | REVIEWER_API_KEY_AUTOFIX | REVIEWER_API_BASE_URL_AUTOFIX |
| autofix fallback | REVIEWER_FALLBACK_MODEL_AUTOFIX | REVIEWER_FALLBACK_API_KEY_AUTOFIX | REVIEWER_FALLBACK_API_BASE_URL_AUTOFIX |
| autofix second fallback | REVIEWER_SECONDARY_FALLBACK_MODEL_AUTOFIX | REVIEWER_SECONDARY_FALLBACK_API_KEY_AUTOFIX | REVIEWER_SECONDARY_FALLBACK_API_BASE_URL_AUTOFIX |
| chat | REVIEW_MODEL_CHAT | REVIEWER_API_KEY_CHAT | REVIEWER_API_BASE_URL_CHAT |
| describe | REVIEW_MODEL_DESCRIBE | primary token | primary base URL |
| changelog | REVIEW_MODEL_CHANGELOG | primary token | primary base URL |
Without Z.ai, autofix defaults to deepseek-v4-pro across the ordered provider chain, other purposes default to deepseek-v4-flash, and validator defaults to Pro. Z.ai mode pins GLM-5.3-Flash first and DeepSeek V4 Flash on the direct DeepSeek fallback for every purpose. Direct DeepSeek responses do not report a billed amount, so the Worker estimates it from the configured model and records cache-hit tokens separately.
The required REVIEWER_API_KEY is the direct DeepSeek fallback credential in the registry chain. REVIEWER_FALLBACK_API_KEY may override it for the direct route. The Worker never sends the primary key to another provider.
Production reviews use glm-5.3-flash through Z.ai Coding first when REVIEWER_API_KEY_ZAI is configured, then deepseek-v4-flash through direct DeepSeek. Without the Z.ai key, production reviews use the legacy configurable chain. The Worker never sends the Z.ai or direct DeepSeek credential to another provider. Every model purpose follows the active fixed order. In Go-only mode, autofix uses deepseek-v4-pro by default. A high-thinking direct DeepSeek attempt can use up to 360 seconds, while Z.ai Coding retains a 240-second ceiling. Time remains reserved for later routes and GitHub finalization. The Worker sends one complete prompt when the changed diff is at most 20,000 characters, fits that route budget, and the rendered prompt is at most 120,000 tokens. Larger diffs use focused batches targeted at 240,000 diff characters; so do prompts that need recovery after a failed full attempt. Each batch is routed by its own size instead of inheriting the whole PR's reasoning mode. For a changed diff or focused batch of 20,000 characters or fewer, the Worker takes the fast path: direct DeepSeek disables thinking, OpenRouter does not request high reasoning, and Z.ai keeps thinking enabled with low reasoning. The attempt is capped at 120 seconds. review_model_attempt logs include the selected reasoning path and effort. Z.ai Coding and direct DeepSeek receive json_object responses immediately. The Worker does not send those routes an unsupported JSON Schema request first. Z.ai Coding review requests set thinking.type to enabled, nested thinking.clear_thinking to false, and reasoning_effort to max on high paths or low on fast paths. Standard direct DeepSeek review requests explicitly set thinking.type to enabled and reasoning_effort to high. Standard OpenRouter DeepSeek requests explicitly set reasoning.effort to high. Validator requests disable direct DeepSeek thinking so the 4,096-token validator output budget is reserved for the JSON verdict. If a multi-route chain receives HTTP 429, it advances to the next provider immediately. Provider allowance denials such as Monthly usage limit reached are terminal for that run, so later chunks and validator batches skip the denied provider. Other transient 429s can still receive one bounded retry in a single-route setup. The Worker stores that denial in KV for 15 minutes, so a new review in the same window skips the known-unavailable route before making a request. The marker expires on its own, allowing a billing fix to recover.
REVIEW_MAX_TOKENS sets the per-call output cap when REVIEW_MAX_TOKENS_CONFIG has no route, model, provider, or default cap. The same cap applies to every individual model request in the review chain. REVIEW_MAX_TOKENS_CONFIG selects a cap in this order: exact routes, models, providers, default, then REVIEW_MAX_TOKENS. Provider names are the values in review usage, such as Z.ai, DeepSeek, and OpenRouter. Production Z.ai Coding and direct DeepSeek V4 Flash review calls use 65,536 output tokens. DeepSeek counts thinking content inside max_tokens, and recent review attempts exhausted 32,768 before emitting JSON; the NVIDIA-hosted route remains at 8,192.
REVIEW_INPUT_TOKENS and REVIEW_INPUT_TOKENS_CONFIG set the maximum prompt size for each review route, with the same precedence. This is separate from REVIEW_MAX_TOKENS: input budgets must leave room for the requested output and provider overhead. The Worker sends the complete rendered diff when the primary route's budget covers it and the rendered prompt is at most 120,000 tokens. Larger prompts use focused batches even when the provider's hard input budget allows one request. Otherwise it groups whole diff hunks into batches; it never clips a changed file to force a batch under budget. Configure the million-context DeepSeek V4 routes below their advertised context window, for example 900,000 prompt tokens, and keep smaller fallback routes at a lower limit.
Chunk planning uses the first available route's budget, rather than the smallest fallback budget, so a large-context primary can review larger batches with less repeated prompt setup. If that route fails, each batch is repacked for the fallback route before retrying. The fallback never receives a batch that exceeds its own input budget. If repository instructions consume most of the chunk budget, the Worker keeps a bounded beginning and end of those instructions and preserves room for diff content. It does not turn a small diff into one batch per changed line.
{
"default": 24000,
"models": { "glm-5.3-flash": 900000, "deepseek-v4-flash": 900000 },
"routes": {
"openrouter/deepseek/deepseek-v4-flash": 900000
}
}{
"default": 4096,
"providers": { "integrate.api.nvidia.com": 2048 },
"models": { "deepseek-ai/deepseek-v4-pro": 3072 },
"routes": {
"integrate.api.nvidia.com/deepseek-ai/deepseek-v4-flash": 1024
}
}Review-log usage line items retain the cap sent to each provider call. Completed and failed reviews also retain the ordered model attempts with timeout, HTTP status, and error details.
REVIEW_TIMEOUT_MS sets the per-model-attempt timeout in milliseconds when REVIEW_TIMEOUT_CONFIG has no route, model, provider, or default timeout. REVIEW_TIMEOUT_CONFIG uses the same JSON shape and precedence as output caps. JSON repair, malformed envelopes, transient HTTP errors, and empty responses share that attempt's timeout budget. A route gets one bounded retry for a 429, 5xx, malformed completion envelope, or empty body before the chain moves on. DeepSeek retries an empty or output-capped high-reasoning response once with thinking disabled. The retry honors a provider Retry-After value up to five seconds and never extends the route deadline. If every route returns one of those recoverable failures, the Worker queues a review retry. The configured timeout is still limited by the route ceiling, the 720-second AI analysis budget, and the time reserved for later fallback routes. The Queue consumer has a 900-second wall, so the Worker keeps 60 seconds for context preparation and 120 seconds for posting the result, finalizing the GitHub check, and clearing recovery state. Two chunks run concurrently by default; when a second wave would exceed the Queue wall, the scheduler divides its remaining time between the ordered routes instead of silently dropping the fallback. When no explicit route, model, or provider timeout is set, Z.ai Coding uses a 240000 millisecond minimum, while direct DeepSeek uses 360000 milliseconds. Fast primary calls remain capped at 90000 milliseconds; explicit route timeouts can lower this cap. Larger reviews and validator calls retain their existing budgets. Keep hosted review routes at 45 seconds or longer unless a provider documents a lower bound. A 15-second route timeout is too short for a large prompt and turns normal provider latency into fallback traffic. After changing provider access or timeouts, rerun the affected PR on a new head or clear its old result cache; complete zero-finding AI results are intentionally not cached anymore. For chunked reviews, the Worker gives Z.ai Coding up to four minutes and direct DeepSeek high-thinking work up to six minutes, with an eight-minute total chunk chain. A response-body read uses the same absolute route deadline as the request and rejects a body above 4 MiB. The 30-second body-idle timer only records a diagnostic warning, so a healthy slow first byte is not cut off by a second hidden timeout. The chain remains bounded by the 720-second AI analysis budget inside the 900-second Queue consumer wall. It reserves each eligible later route's configured timeout before assigning time to the current route; terminally denied or capacity-exhausted routes are removed for the current run. Small primary reviews (at most 20,000 diff characters and 120,000 estimated prompt tokens) try direct DeepSeek before the subscription routes in fixed-provider mode. This spends direct API credit earlier to reduce latency; it does not change the validator route order or any evidence gate. Legacy configured chains retain their preferred order except for temporary timeout recovery.
Primary review timeouts observed in two distinct runs within five minutes move a route behind healthy alternatives for two minutes. The next review after expiry tries its normal position again. A successful request clears the timeout history. Routes are never removed by this policy: if healthier alternatives fail, a cooling route can still recover the review. Health state is scoped by hashed endpoint, model, credential, and fast/standard review class. KV accounting is advisory and may miss concurrent observations; storage errors preserve normal routing. Logs use review_route_timeout_cooldown and review_route_deprioritized.
A provider timeout belongs to the current chunk, so another parallel chunk can still try that route; the timed-out route is skipped only for that chunk's remaining fallback chain. The chunk deadline aborts the underlying provider request before fallback processing continues. When those configured reservations do not fit in the remaining window, it falls back to an equal split so every eligible route still gets a chance. Provider-attempt logs record the route, actual elapsed time, timeout budget, HTTP status, and whether the call completed, which separates a slow provider from a denied one. Successful review calls also log response-header latency, first-byte latency, body duration, body size, body chunk count, and a 30-second body-idle warning. The Worker keeps the non-streaming JSON contract so malformed partial output cannot be published.
Large-diff chunks use two concurrent provider calls total by default. Set REVIEW_CHUNK_CONCURRENCY to a positive integer to lower or raise that limit; the Worker caps it at the number of pending chunks for the review.
Set REVIEW_OFF_PEAK_REVIEW_DELAYS=true to delay webhook-triggered reviews that arrive inside provider peak windows (DeepSeek peaks at 01:00-04:00 and 06:00-10:00 UTC Mon-Fri; the Z.ai coding-plan peak is 06:00-10:00 UTC Mon-Fri) until the window ends, capped at the one-hour Queues delay limit. Manual command reviews are never delayed, and the flag does nothing outside peak windows.
Validator batches run serially when a review has more than ten findings. This keeps capacity-limited validator routes from receiving a burst of requests and making every batch fail together. Each batch has a 180-second operation budget. The three-route chain receives up to 60 seconds per route by default, with unused time moving to the next route and explicit REVIEW_TIMEOUT_CONFIG entries still taking precedence. Results are merged in input order, and a failed batch still makes the review incomplete and retryable. Each validator route attempt logs its provider, model, HTTP status, elapsed time, timeout budget, timeout flag, and success flag as validator_model_attempt. The validator returns one compact decision per server-issued validation_id. It does not copy file, line, severity, or proof fields back to the Worker. A missing, unknown, or duplicated ID makes the batch invalid and retryable.
Validator and autofix endpoint chains share one operation deadline. A malformed HTTP-success response, empty content, or invalid response schema advances to the next configured endpoint rather than being accepted as a completed call.
NVIDIA's 503 ResourceExhausted response for a reached worker request limit skips retries for that route and moves to the next configured model. Chunked reviews stop submitting more chunks to that route after the first such response. A 401, 402, or 403 does the same: the Worker waits for the first route probe before sending concurrent chunks, then skips a denied route for the rest of that review.
{
"default": 60000,
"providers": { "integrate.api.nvidia.com": 30000 },
"models": { "deepseek-ai/deepseek-v4-pro": 45000 },
"routes": {
"integrate.api.nvidia.com/deepseek-ai/deepseek-v4-flash": 15000
}
}For the production Z.ai Coding â DeepSeek review chain, keep Z.ai at 240000 and direct DeepSeek at 360000. The REVIEW_TIMEOUT_CONFIG secret overrides the REVIEW_TIMEOUT_MS Worker variable, so changing wrangler.toml alone does not replace a shorter explicit route entry.
Service controls â
| Name | Use | Default |
|---|---|---|
BOT_HANDLE | Mention used by PR commands | @codebeaver |
CROSS_REFERENCE_BOTS | Comma-separated review bots to cross-reference | built-in bot list |
LOG_LEVEL | trace, debug, info, warn, error, or fatal | info |
Debug routes â
| Name | Use |
|---|---|
DEBUG_AUTH_TOKEN | Separate bearer token for /diag and /test-review |
DEBUG_INSTALL_ID | GitHub App installation used by debug routes |
DEBUG_ALLOWED_REPOS | Comma-separated repository allowlist; * allows every installed repo. Default-deny: without it no repo is allowed |
Keep the debug token separate from the webhook secret. DEBUG_ALLOWED_REPOS is default-deny: with no allowlist configured, the debug endpoints reject every repository.
Cloudflare bindings â
wrangler.toml declares:
[[kv_namespaces]]
binding = "REVIEW_CACHE"
[[vectorize]]
binding = "CODEBASE_VECTORS"
index_name = "codebeaver-codebase"
[ai]
binding = "AI"
remote = trueThe Worker uses a 300,000 ms CPU limit and a five-minute cron. Network wait does not count as Worker CPU, so provider and GitHub requests also have explicit abort timers.
Create Vectorize once if the account does not have the index:
wrangler vectorize create codebeaver-codebase --dimensions=768 --metric=cosineSecret ownership â
Production values are provided by the operator through a secret manager (Infisical in the reference deployment). Map them to their destinations:
| Destination | Contents |
|---|---|
| Cloudflare Worker secrets and vars | Worker runtime secrets and vars |
| GitHub Actions production environment | Deploy tokens and smoke credentials |
Typical production smoke secrets for GitHub Actions:
| Name | Owner |
|---|---|
CF_ACCESS_CLIENT_ID / CF_ACCESS_CLIENT_SECRET | Access service token for the deploy dashboard smoke (preferred) |
DASHBOARD_SMOKE_COOKIE | Optional staff Access JWT for authenticated session/stats smoke; deprecate once service tokens are live |
CLOUDFLARE_API_TOKEN / CLOUDFLARE_ACCOUNT_ID | Wrangler deploy |
Change values in your secret manager, run the configured sync to the destination, then verify the deployed service. Do not commit .env files, provider tokens, GitHub private keys, or debug credentials.
If production secrets are owned by a Terraform configuration, remember that a merged Terraform PR does not apply itself. After a reviewed production plan has explicit approval, run the manual apply from that checkout, then verify /health, the active provider routes, and the deployment logs. Keep the apply and the approval tied to the same reviewed plan.
Keep REVIEW_MODEL synchronized between the Cloudflare Worker path and the GitHub Actions production environment. Live eval calls the provider directly, so a mismatch tests another model than production.
CI and pull requests â
.github/workflows/ci.yml runs on non-draft PRs targeting main:
- dependency install with the frozen lockfile;
- high-severity dependency audit with a dated exception;
- Worker lint, typecheck, build, and tests;
- dashboard typecheck, tests, and build.
CI, Workflow Lint, and Eval also accept workflow_dispatch for recovering a missed GitHub pull-request event. Run each workflow against the PR head, then confirm their completed checks are attached to that commit before merging.
The auto-merge workflow enables GitHub native squash auto-merge with branch deletion for every eligible, non-draft PR from this repository. Required checks and a human CHANGES_REQUESTED review still block the merge.
The production deploy workflow triggers automatically on every push to main (and manually via workflow_dispatch). It reruns a frozen install, Worker lint and root typecheck, the Wrangler dry build, root Vitest, and dashboard checks before deploying the Worker and dashboard and publishing the CLI â so a squash merge to main is itself the production deploy decision. The manual dispatch path below is for recoveries and operator-initiated redeployments.
The same workflow also builds and deploys the docs site (apps/docs, VitePress over the docs/ markdown) to the Cloudflare Pages project cb-codebeaver-docs and smoke-checks https://cb-codebeaver-docs.pages.dev/. The initial Pages project is created once per account with wrangler pages project create cb-codebeaver-docs --production-branch main. A custom domain (for example docs.codebeaver.dev) is attached through the Cloudflare dashboard or wrangler pages project domain commands and needs the zone's DNS.
PR title format without a Linear issue:
<area>: <imperative description>Documentation-only work may name the artifact as the area. PR bodies use What changed, Why, Validation, and Notes.
Deploy â
Production deploy normally follows a squash merge to main:
push main
-> lint, typecheck, build, test
-> deploy Worker
-> deploy dashboard
-> check dashboard URL
-> check Worker /health
-> publish CI-suffixed CLI versionManually run the same workflow:
gh workflow run deploy.yml --ref main
gh run watchThe Worker deploy uses root wrangler.toml. The dashboard runs its own build and apps/dashboard/wrangler.toml; Cloudflare serves dist/ with single-page fallback.
The post-deploy health check waits for propagation, then retries for about 150 seconds. It requires:
{
"status": "ok",
"jwt": "valid",
"secrets": "ok",
"build": { "sha": "<deployed-git-sha>" },
"review": {
"directDeepSeek": "configured"
}
}The CLI publish job runs only after deploy succeeds. It creates a CI prerelease version for that workflow run and publishes to GitHub Packages.
Health and diagnostics â
Health â
curl --fail https://codebeaver.workers.dev/health/health checks that required provider, webhook, and CLI signing strings exist and that the Worker can create a GitHub App JWT. Its build.sha identifies the deployed commit. review.routes exposes only model IDs and provider hostnames, while review.directDeepSeek confirms that the deployment can reach its direct DeepSeek fallback. It does not call the model provider, read a repository, or verify every Cloudflare binding.
Installation diagnostic â
Use /diag before any write test:
curl --fail \
-H "Authorization: Bearer $DEBUG_AUTH_TOKEN" \
"https://codebeaver.workers.dev/diag?repo=OWNER/REPO&pr=123"It resolves the configured installation token and reads the PR title. A failure response identifies whether installation-token creation or the GitHub API read failed.
Live forced review â
/test-review is a write operation. It can create or update comments, check runs, reviews, labels, learning state, and PR description content.
curl --fail \
-H "Authorization: Bearer $DEBUG_AUTH_TOKEN" \
"https://codebeaver.workers.dev/test-review?repo=OWNER/REPO&pr=123"Use it only on a named PR where those side effects are intended. Prefer /diag, GitHub API reads, and logs first.
Logs and error reporting â
Stream production Worker logs:
pnpm tailOr request JSON output directly:
wrangler tail codebeaver --format jsonwrangler tail is live-only. It does not recover logs emitted before the tail session started.
The outer handler logs request_start, request_failed, and request_end with request ID, method, path, status, and duration. Queue messages add correlation_id (the invocation ID), optional origin_request_id, job kind, attempt, and repository/PR context to every log. Each review adds a review_run_id to its start, phase, and failure logs, and stores its latest phase with the recovery marker in KV. review_job_finished records completed, retrying, cancelled, superseded, and failed queue states with the duration, retry decision, and check-run ID. A state is emitted once per invocation and attempt; final states remain once per invocation. Repeated writers still refresh status and retry cleanup side effects without appending another event, history row, or completion log. review_complete is the terminal success record: it includes the GitHub check ID and finalization state, validator status, model, wall-clock timings, queue wait, partial/incomplete state, every publication-filter count, the active finding cap, coverage gap, and suggestion-validation totals. Search by the correlation or run ID to follow one review across Worker logs and dashboard data.
The prepare_context phase includes context_timings_ms for config and diff loading, changed-file reads, structural search, consumer/tree reads, test reads, and prompt inputs. Compare these fields across size buckets before changing concurrency or route budgets. Synchronize cache hits still rebuild the independent prompt inputs; a mismatch is logged and falls through to the full context path. Shadow recall starts while the public result is being posted, so its provider time is included in the final usage record without delaying context or primary analysis.
validatorStatus is validated, cancelled, provider_exhausted, invalid_response, or not_run. Review records count suggestions held during a validator outage. Partial chunk logs include reviewed and failed counts, plus separate deadline and input-budget omission counts. Review records also store the resulting coverage gap and active finding cap so omitted work is visible in recall metrics. A partial review completes the required check as neutral and queues a retry when the outage looks transient; permanent provider denials (credits / model access) also finalize as neutral so merges are not blocked by infra alone (CB-1731). Cached completed chunks do not call a provider again. Configuration fetch failures other than a missing config file log as config_load_failed with repository, ref, and scope. Partial log records also retain the current-path batches in each state. Reviews that wait at least five minutes emit review_queue_wait_alert; partial chunk runs with failed, omitted, or deferred batches emit review_chunk_retry_alert. Both events include the repository, PR, review run, and bounded counts for alert filtering. Provider responses that send headers and then stop producing body bytes are aborted when the route deadline expires and appear as timed-out model attempts. They cannot hold a queue review open indefinitely, and a slow response is no longer aborted by the diagnostic idle warning.
When at least one structured inline suggestion is generated, suggestion_validation_summary records generated, verified, downgraded, and reason counts. A failed head-file read also emits one suggestion_source_fetch_failed warning per affected file. These events contain repository, PR context, head SHA, path, and a bounded error message; they never contain replacement code. Repository instruction and configured-context failures log as repo_instruction_fetch_failed, external_context_fetch_failed, or external_context_fetch_timeout; expected missing files use the corresponding *_missing event. KV fallback failures log as learnings_load_failed or conventions_load_failed. GitHub responses with an error or a low remaining budget log github_api_rate_budget with the endpoint, status, remaining count, and reset time. GitHub check-run and reaction rate limits log github_rate_limit_backoff_recorded after writing the installation backoff. Review admission and code-search saturation log review_deferred_for_github_admission or github_search_admission_deferred; the shared governor state lives in the GITHUB_RATE_GOVERNOR Durable Object.
Logs redact and bound metadata at every nesting level. Each invocation also has a 96 KiB budget for info logs, leaving room below Cloudflare's 256 KiB log cap for platform metadata and late warnings or errors. When the budget is spent, the Worker emits one info_log_budget_exhausted warning and drops later info events; warnings and errors still pass through. Never add raw tokens, cookies, private keys, diffs, or unrestricted provider responses to log fields.
Feedback learning reads both Durable Object and KV fallback records, deduplicated by PR number and finding ID. Edited replies use the newest reply ID and capture time. The Durable Object migrates its former finding-only primary key to (pr, id) transactionally on startup. Existing rows keep their stored PR; feedback already overwritten or dropped by the old key cannot be reconstructed. Reply deletion removes its Durable Object signal and retracts its accuracy count.
Dashboard review totals report full execution time, separate model time and queue wait, completion, retries, feedback coverage, and cost per published or accepted finding. Capped scans and failed learning reads carry partial or unavailable metadata; they are not shown as zero activity. Finding outcomes distinguish open, not_reported, and verified_fixed. A missing finding in a later review is not_reported; verified_fixed is reserved for a successful autofix or an explicit accepted-fix path.
The queue dashboard reads retained job events under review-job-event:v1:<invocation-id>:, current health alerts, and captured dead-letter records under review-dlq:v2:<invocation-id>. Each status transition emits an event with request time, receive time, queue wait, attempt, duration, and a typed error class when one is known. The API reads review-dlq-recent:v1:* in reverse-time order, removes duplicate and dangling indexes, and returns at most the newest 100 actionable records. During the v1 migration it unions indexed records with review-dlq:v1:*; both legacy IDs and opaque v2 IDs remain replayable.
Closing a PR writes review-dlq-terminal:v1:<repo>:<pr> before pruning its DLQ records. Unmerged closes remove every job for that PR. Merged closes remove review and command jobs but retain merged_pr_cleanup while merged-pr-cleanup:v1:<repo>:<pr> exists. Successful feedback collection and indexing remove the cleanup dead letter and record completion in the terminal marker. Reopening removes the marker before a new review is queued. Draft conversion leaves DLQ work alone.
Review incident checks â
For a report that a review is stuck or missing:
- Read the PR head SHA,
AI Code Reviewcheck runs, bot reviews, walkthrough, and inline threads. - Search Worker logs for that repository, PR number, and request window.
- Check for
review_job_enqueued,review_job_received,review_job_cancelled_before_start,review_started, provider response or denial, posting steps,request_failed, and final check update. - Inspect
review-job-status:v1:<invocation-id>,in-flight:<repo>:<pr>,queued-review:v1:<repo>:<pr>,orphan-queue:<repo>:<pr>,merged-pr-cleanup:v1:<repo>:<pr>, andlast-reviewed:<repo>:<pr>in KV. Allow for KV's eventual consistency. A job record remains for 90 days and shows queued, received, started, retrying, completed, cancelled, superseded, or failed. An in-flight record retains the latest review phase and review-run ID when recovery is needed. - For a full timeline, inspect
review-job-event:v1:<invocation-id>:. For exhausted queue delivery, inspect/api/queue's actionabledeadLetterJobslist or thereview-dlq:v2:,review-dlq-pr:v2:, andreview-dlq-recent:v1:records. Legacy records can remain underreview-dlq:v1:until reconciliation finishes. Replay only after checking the stored attempt count and failure class. - Call
/healthand then/diagif installation auth is suspect. - Verify each AI finding against the current head before treating it as a real blocker.
A completed check with conclusion failure can mean the service worked and found a configured blocker. Separate service health from review correctness.
Keep diagnosis read-only until the evidence is collected and a write is authorized. Webhook redelivery, /test-review, a /review comment, KV edits, and thread resolution all change state or start work.
If a stale cancelled check blocks merging, an authorized operator can run /review and wait for a new completed check. Do not admin-merge around it.
Scheduled recovery â
The full cleanup cron runs every five minutes. A one-minute tick refreshes queued comments and dispatches due KV-backed retries, so an incomplete review does not wait for the next cleanup rotation. Full cleanup requeues active queue statuses that have gone stale, rescues queued review jobs that have remained in queued state for ten minutes without reaching a consumer, scans installed repositories, recovers stale in-flight records and check runs, drains remaining retry work with awaited review calls, compacts learned rules, and clears stale operational state. Dispatching an orphan retry refreshes its existing queued-head timestamp so a later stale-queue scan does not replace it; a newer invocation still keeps ownership. It also retries terminal-PR DLQ pruning and reconciles at most 20 distinct legacy DLQ PRs per run, with one GitHub state read per PR. Open-PR v1 records gain v2 indexes. Closed review and command records are deleted, while merged cleanup remains only when its cleanup KV entry still exists. Failed GitHub or KV reads leave the record untouched for the next run; review-dlq-migration:v1 is written only after every v1 record was deleted or indexed.
Recovered jobs keep their invocation, comment, and check-run IDs. Queued-job recovery replaces the stale head pointer only after Cloudflare accepts a replacement message. Merge webhooks put feedback collection and Vectorize updates in merged-pr-cleanup:v1:*; the cron retries failed entries with backoff instead of leaving that work in a 30-second waitUntil() window.
If a queue redelivery arrives while the same cleanup invocation is already running in a Worker isolate, the duplicate is acknowledged and the active invocation keeps the KV lease. A stale lease remains eligible for the scheduled sweep, so a crashed invocation can still be retried without turning normal overlap into a dead-letter job.
Each run records scheduled_cleanup_phase timing events for GitHub App authentication, installation and repository scans, recovery, queue processing, and learning-rule compaction. Use elapsed_ms and phase_elapsed_ms to identify a slow external call before changing the schedule or timeout.
Review code also records a heartbeat in the in-flight entry. Heartbeat comment writes are serialized and drained before the completed walkthrough replaces the progress text. Running comments show a recalculated remaining time and UTC ETA; the estimate may move as the current phase runs longer or shorter than history. Heartbeats prove that the Worker is alive; they do not extend the review phase deadline. Recovery freshness uses the last completed phase, starting at prepare_context, so a hung phase is reclaimed after the normal stale interval even while the heartbeat timer continues. Command timing records live under operation-timing:v1:<repo>:<operation>:* for 90 days and include total/stage durations, size bucket, file count, and completion time. Only successful sessions are written. Queued command comments use the same lifecycle and retain their invocation timing in the 90-day job record. A push records the newest requested head at queued-review:v1:*; older queued jobs are dropped and a newer head waits for an active review.
If the Worker deploys during a review, its lock becomes stale after 20 minutes. A 20-minute marker is longer than the 15-minute Queue consumer wall because a provider call can delay the heartbeat while its network promise is awaited; a shorter marker could cancel a live review and start a duplicate. A queue redelivery that sees the fresh marker schedules the same invocation back onto Queues with a 30-second delay. That delivery recovers the old check as soon as the marker becomes stale. The five-minute cleanup remains a fallback for a consumer that never receives another delivery. The in-flight marker records the queue invocation ID and latest phase, and a same-invocation redelivery records queue_delivery so a runtime cancellation or restart is visible in the job timeline. Check the status comment and logs before manually redelivering the webhook. Each PR head has one batch comment. A retry or terminal provider failure updates that comment in place, even after its KV pointer is lost, and keeps earlier provider attempts in a collapsed history below the current status.
Evals â
Deterministic fixtures run in CI for changes to review prompts, schemas, verdict logic, finding quality, and eval files.
Run deterministic eval tests locally:
pnpm vitest run eval/deterministic.test.ts eval/score.test.ts eval/vitest.eval.tsRun live-model fixtures:
export REVIEWER_API_KEY=your-provider-key
export REVIEW_MODEL=provider/model
EVAL=1 pnpm vitest run eval/vitest.eval.tsThe model bakeoff repeats each fixture once by default for API compatibility. Set EVAL_REPETITIONS=3 (the live bakeoff default) to measure run-to-run stability. The report retains the first run in fixtures and all runs in runs; stability reports verdict flips, finding identity overlap and churn, severity churn, malformed/error rate, and min/max/median duration and cost. Set EVAL_MODELS to a comma-separated list when comparing a smaller set of models. Repeated live runs cost provider tokens, so use the default one-run mode for quick checks and three runs when selecting or changing a model or prompt.
eval/env.ts accepts OPENROUTER_API_KEY as an alternative token. Live runs carry the fixed provider chain into the eval env: REVIEWER_API_KEY_ZAI, REVIEW_PROVIDER_ORDER, and the fallback route settings, mirroring buildReviewModelChain. The live GitHub Actions job uses the production environment. If no token is present, it reports a notice and leaves the deterministic gate as the result.
Live fixtures call the model provider directly. They do not pass through the deployed Worker, use GitHub App context, or write PR feedback. The live job runs after pushes to main and is visible in Actions, but only deterministic fixtures block the evaluation gate: model output can vary despite retries, and a single provider miss must not mark an otherwise healthy deployment as failed.
The eval-llm-gate job is the required live exception. It runs the curated subset from GATE_FIXTURE_IDS (eval/loadFixtures.ts) through eval/vitest.eval-gate.ts: the clean/noise precision fixtures with antiFindings and canonical single-defect recall fixtures. 025 stays out until its known state-changes-during-await miss is fixed; the non-blocking full run still reports it. It uses the same secrets, retries, and assertions as the full live run (eval/llmAssertions.ts), so a red eval-llm-gate means a real recall or precision regression on seeded fixtures â fix the regression or, only if a fixture itself proved unstable, curate it out of the subset in the same PR that documents why. The full eval-llm job stays continue-on-error and skipped on pull_request events; the gate job follows the same event filter and blocks the Evaluation Gate on pushes and workflow dispatch.
Live runs can measure whether self-learning rules help. Set EVAL_LEARNING_RULES=1 and EVAL_LEARNING_RULES_FILE=<path> to inject a KV-rule-store-shaped JSON file ({"bestPractices": [...], "falsePositives": [...]}, rule shape in src/learning-model.ts) into the eval system prompt exactly like production; the compare runner (runLearningBakeoff, flag handling in eval/bakeoff.ts) runs the same model set with and without the file and reports per-fixture verdict/finding deltas beside each arm's stability metrics. See eval/README.md.
Rollback â
Cloudflare retains Worker deployment history. For a bad deploy:
- stop manual review triggers that would increase impact;
- identify the last healthy deployment and the failing request IDs;
- roll back through the Cloudflare deployment history or redeploy the last healthy
maincommit through the normal workflow; - verify
/health, one read-only/diag, dashboard load, and Worker logs; - re-run affected reviews with
/reviewafter service health is restored.
Do not point production at an unreviewed local branch as a shortcut. The repository requires the post-merge production deployment to match main.
Queue operations â
Provision codebeaver-jobs, codebeaver-jobs-dlq, codebeaver-priority-jobs, codebeaver-priority-jobs-dlq, codebeaver-commands, and codebeaver-commands-dlq before deploying a Worker that binds REVIEW_JOBS, REVIEW_PRIORITY_JOBS, and COMMAND_JOBS. The normal consumer processes one message per invocation with a maximum concurrency of six. The priority consumer has two reserved slots for PRs in codebeaver/codebeaver. Watch queue retries, DLQ depth, GitHub comment PATCH failures, stale invocation rejections, and GitHub API volume after deploy. While a review waits, its PR comment shows an approximate number of reviews ahead, a coarse estimated start time, and the last UTC queue check. A separate once-a-minute cron refreshes at most 50 waiting comments per run. Those numbers are estimates, not a FIFO guarantee: eight review consumers can run in parallel; the command lane has two more slots. Retries, newer PR heads, and installation-wide rate-limit backoff can reorder work. The installation governor starts at four reviews and ramps to eight after clean intervals; GitHub throttling halves its capacity. Keep chunk concurrency at two while evaluating this increase. Compare queue wait, GitHub throttling, and AI-provider failures under demand before raising either limit again. Transient review jobs retry after one minute and then five minutes. Provider billing, credential, and model-access failures are terminal. Long-running commands use two isolated consumer slots. If COMMAND_JOBS is not configured, command jobs fall back to REVIEW_JOBS for compatibility. Non-search GitHub rate-limit responses are not ordinary queue retries: the Worker records the reset or Retry-After time in an installation-wide KV backoff and records a durable orphan-recovery entry for the affected review. Code-search limits use only the separate search cooldown. Queue consumers and scheduled recovery defer that installation until the reset, then cron resumes the review. A redelivery for the same fresh in-flight head waits instead of being acknowledged as complete; it schedules a delayed Queue delivery instead. Once that marker is stale, the queue starts its replacement review directly. Orphan recovery stops after ten attempts. At that point it closes the carried GitHub check as neutral, updates the batch comment, and removes the retry entry. If GitHub rejects that final update, the entry stays in KV for another attempt so an in-progress check is never abandoned silently.
For a queued command, inspect operation-status:v1:* with the invocation ID from Worker logs. The associated PR operation comment is the public source of truth for its last phase, retry state, and terminal result.
Linear read access â
Store LINEAR_REPOSITORY_CREDENTIALS in the operator's secret manager and sync it to the Worker as a secret. It is a JSON object keyed by lowercase exact GitHub owner/repo names, with no wildcard or first-entry fallback:
{
"owner/repo": {
"apiKey": "<Linear API key>",
"workspace": "<Linear workspace URL slug>",
"teamIds": ["<allowed Linear team UUID>"]
}
}Use a read-only Linear API key scoped to the required teams when provisioning credentials. This operator-owned mapping grants access; repository YAML cannot select another credential, workspace, or team. Unmapped repositories do no Linear work. Only enable a mapping when the repository's readers may see that team's task content. Cached records contain bounded task data, not credentials, and are partitioned by repository, credential/read scope, and PR metadata. Credentials must never appear in logs or PR comments.
The current query shape is checked against Linear's GraphQL API. Look for linear_context_unavailable, linear_advisory_unavailable, and linear_cache_read_failed/linear_cache_write_failed to diagnose missing context. These optional failures never make the review fail or approve. A slow advisory may be omitted when the review posts. No Linear write scope is required for draft URLs.