An AI pull-request reviewer, running as a GitHub App. A verified webhook lands on a BullMQ queue; an orchestrator plans the review, runs specialised agents in two parallel waves with 12 permission-scoped tools, filters false positives in a reflection pass, and posts line comments back to the PR.
- 11
- specialised agents, in parallel waves
- 12×
- fewer GitHub API calls per review
- 7 → 68
- tests
One PR, from GitHub’s webhook to a posted review — queued, planned, inspected by two waves of agents, and filtered before it speaks.
PR Sentinel01 / 07
GitHub fires a webhook, and its signature is verified before anything else runs.
HMAC over the raw body; tokens and secrets are encrypted at rest with AES-256-GCM.
Never fetch what an agent already has
up to 12× fewer calls
- Problem
- Six agents were each allowed to read the same PR comments and reviews, so one review could make a dozen near-identical GitHub calls before reading a file.
- Approach
- Immutable git data cached by commit SHA, fetched once ever; mutable PR data served from webhook-synced tables. Code search stays live — its index moves.
- Outcome
- Cache hits are byte-identical to a live call, so review quality isn’t traded for budget.
Pause and resume, instead of finding the ceiling in production
soft + hard floor
- Problem
- Even after caching, a busy installation could run its shared GitHub budget down in the middle of a review.
- Approach
- A Redis-backed budget tracker from GitHub’s own headers. Below a soft floor no new wave starts; the review checkpoints and re-enqueues itself for the reset.
- Outcome
- The UI says “paused — resuming in ~12m” instead of looking stuck or failed.
A rate-limit floor that could deadlock a review
9 / 10 left, still paused
- Problem
- A review paused and re-paused forever with 90% of its budget left: a hard-coded minimum floor of 50 was larger than a category whose whole limit was 10.
- Approach
- Floors became pure ratios of the reported limit, and the rate-limit category now goes into the log so the next surprise is diagnosable.
- Outcome
- Found live in end-to-end testing — exactly what the design doc promised the code already did.
One revoked token took down the whole server
500 → 401
- Problem
- Express 4 doesn’t forward rejected promises from async handlers, so a single expired OAuth token crashed the entire process.
- Approach
- An async-handler wrapper on every uncovered route, and a specific 401 telling the user to sign in again.
- Outcome
- One bad request now fails one request.
A circuit breaker — and a dashboard that stopped lying
3 fails → 60 s open
- Problem
- When the model provider degraded, seven agents each burned three retries back to back — and an all-failed run showed as a “confirmed clean pass”.
- Approach
- A process-wide breaker that opens after three consecutive full failures; the empty state now checks agent executions, not just finding count.
- Outcome
- No retry storms during outages, and “inconclusive” is reported as inconclusive.