A·MIST Résumé ↗
← All work

Case study 01Real-time · Multi-tenant · Infra2024 – 2026

A multi-tenant platform for large, live, proctored quiz contests — payments, WebSockets, queues and auto-scaling infrastructure, built solo from the first diagram to the load test.

Role
Solo — architecture to load test
Target
10,000 concurrent participants
Infra
AWS · ALB · Auto Scaling · ElastiCache · Terraform
7,500
concurrent sockets, load-tested
$37
a month at idle
< 100 ms
API at peak
Field notes

7,500concurrent WebSocket users, load-tested against a 10K architecture.rest
1 → 8instances behind the load balancer, pre-warmed before the peak.rest
0answers scored twice. Every submission takes an idempotent Redis lock.rest
6BullMQ queues keep heavy work off the request path — the API stays under 100 ms.rest
The architecture

One participant, one answer — followed from payment to certificate through every service it touches. Scroll to move the camera.

QuizBuzz01 / 07

Every contest begins as a single entity. Everything else grows out of it.

Decisions & challenges · 06

  1. 01Shipped

    630× faster test-data seeding

    42 min → 4.1 s

    Problem
    Seeding 10,000 participants for a load test took 42 minutes — sequential Prisma upserts over an SSH tunnel at ~100 ms a round-trip.
    Approach
    Bulk INSERT … ON CONFLICT DO NOTHING through a raw pg.Pool, 500-row batches, four in flight, with deterministic keys.
    Outcome
    For bulk work, round-trip count dominates query complexity.
  2. 02Shipped

    Zero job loss across a dual-Redis switch

    0 jobs lost

    Problem
    Jobs scheduled before go-live sat in the local Redis; after switching to ElastiCache, workers listened elsewhere and contests never started.
    Approach
    A DUMP/RESTORE migration tool that copies every key — types and TTLs intact — both ways, run inside the live container by the go-live and go-idle scripts.
    Outcome
    “Always remember to do X” is not a control. Move the execution so the system guarantees it.
  3. 03Shipped

    Closing the auto-submit coverage gap

    full coverage

    Problem
    Auto-submit saved 140 of 1,323 live participants: it only read the active set, and an OOM crash had dropped the rest into the disconnected set.
    Approach
    Union both sets on time expiry, then submit in batches of 50 with Promise.allSettled so one failure can’t abort the rest.
    Outcome
    Once you find one instance of a failure class, predict its symmetric twin.
  4. 04Shipped

    Debugging the Socket.IO protocol layer

    50% → ~100%

    Problem
    The k6 client sent plain JSON over raw WebSockets; Socket.IO speaks Engine.IO v4 on top, so the server ignored everything. Half the sockets “worked”.
    Approach
    Verified the transport with curl, then rewrote the load client as a state machine speaking the real EIO4 handshake, namespace auth, framing and ping/pong.
    Outcome
    A clean transport connection says nothing about the protocol above it.
  5. 05Designed

    Rethinking autoscaling for WebSocket load

    CPU ≠ load

    Problem
    Scaling triggered on CPU at 60%, but WebSocket load is memory- and IO-bound: CPU sat at 20% while the heap hit 95%. New instances also took ~7 minutes.
    Approach
    Three options weighed: a memory alarm, a custom connection-count metric, or pre-warming from the registration count before the contest starts.
    Outcome
    Pre-warming chosen as primary — the peak is known in advance — with CloudWatch as defence in depth.
  6. 06Shipped

    Auditing every rate limiter, not just the broken one

    600 ms ≠ 600 s

    Problem
    An OTP limiter guarded a login route with no OTP; every k6 user shared one IP and exhausted it. The audit found a second limiter missing a ×1000.
    Approach
    Removed the misplaced limiter, then audited every limiter in the codebase and fixed the unit mismatch in both places.
    Outcome
    Finding one bug of a class is the moment to look for its siblings.
Next case studyPR SentinelAn AI pull-request reviewer, running as a GitHub App. A verified webhook lands on a BullMQ queue; an orchestrator plans the review, runs specialised agents in two parallel waves with 12 permission-scoped tools, filters false positives in a reflection pass, and posts line comments back to the PR.Read the case study