A·MIST Résumé ↗
← All work

Case study 02AI agents · GitHub App · In progress2026

An AI pull-request reviewer, running as a GitHub App. A verified webhook lands on a BullMQ queue; an orchestrator plans the review, runs specialised agents in two parallel waves with 12 permission-scoped tools, filters false positives in a reflection pass, and posts line comments back to the PR.

Role
Solo — backend, agents and dashboard
Runs as
GitHub App · Express API + BullMQ worker · Next.js 16 dashboard
Models
Vercel AI SDK via OpenRouter — a model tier per agent
Source on GitHub
11
specialised agents, in parallel waves
12×
fewer GitHub API calls per review
7 → 68
tests
Field notes

11specialised agents around one orchestrator, scheduled into parallel waves by a planner.rest
12permission-scoped tools — each agent may only call the ones its role needs.rest
12×fewer GitHub calls per review, by caching immutable git data against the commit SHA.rest
7 → 68tests, with webhooks verified by HMAC and tokens encrypted with AES-256-GCM.rest
The architecture

One PR, from GitHub’s webhook to a posted review — queued, planned, inspected by two waves of agents, and filtered before it speaks.

PR Sentinel01 / 07

GitHub fires a webhook, and its signature is verified before anything else runs.

HMAC over the raw body; tokens and secrets are encrypted at rest with AES-256-GCM.

Decisions & challenges · 05

  1. 01ShippedD6

    Never fetch what an agent already has

    up to 12× fewer calls

    Problem
    Six agents were each allowed to read the same PR comments and reviews, so one review could make a dozen near-identical GitHub calls before reading a file.
    Approach
    Immutable git data cached by commit SHA, fetched once ever; mutable PR data served from webhook-synced tables. Code search stays live — its index moves.
    Outcome
    Cache hits are byte-identical to a live call, so review quality isn’t traded for budget.
  2. 02ShippedD7

    Pause and resume, instead of finding the ceiling in production

    soft + hard floor

    Problem
    Even after caching, a busy installation could run its shared GitHub budget down in the middle of a review.
    Approach
    A Redis-backed budget tracker from GitHub’s own headers. Below a soft floor no new wave starts; the review checkpoints and re-enqueues itself for the reset.
    Outcome
    The UI says “paused — resuming in ~12m” instead of looking stuck or failed.
  3. 03ShippedD18

    A rate-limit floor that could deadlock a review

    9 / 10 left, still paused

    Problem
    A review paused and re-paused forever with 90% of its budget left: a hard-coded minimum floor of 50 was larger than a category whose whole limit was 10.
    Approach
    Floors became pure ratios of the reported limit, and the rate-limit category now goes into the log so the next surprise is diagnosable.
    Outcome
    Found live in end-to-end testing — exactly what the design doc promised the code already did.
  4. 04ShippedD19

    One revoked token took down the whole server

    500 → 401

    Problem
    Express 4 doesn’t forward rejected promises from async handlers, so a single expired OAuth token crashed the entire process.
    Approach
    An async-handler wrapper on every uncovered route, and a specific 401 telling the user to sign in again.
    Outcome
    One bad request now fails one request.
  5. 05ShippedD23

    A circuit breaker — and a dashboard that stopped lying

    3 fails → 60 s open

    Problem
    When the model provider degraded, seven agents each burned three retries back to back — and an all-failed run showed as a “confirmed clean pass”.
    Approach
    A process-wide breaker that opens after three consecutive full failures; the empty state now checks agent executions, not just finding count.
    Outcome
    No retry storms during outages, and “inconclusive” is reported as inconclusive.
Next case studySmartFormFlowA multi-tenant event platform that replaced paper, spreadsheets and manual payments for a client — registrations, verified payments, certificates and reminders, run by background workers.Read the case study