A multi-tenant platform for large, live, proctored quiz contests — payments, WebSockets, queues and auto-scaling infrastructure, built solo from the first diagram to the load test.
- 7,500
- concurrent sockets, load-tested
- $37
- a month at idle
- < 100 ms
- API at peak
One participant, one answer — followed from payment to certificate through every service it touches. Scroll to move the camera.
QuizBuzz01 / 07
Every contest begins as a single entity. Everything else grows out of it.
630× faster test-data seeding
42 min → 4.1 s
- Problem
- Seeding 10,000 participants for a load test took 42 minutes — sequential Prisma upserts over an SSH tunnel at ~100 ms a round-trip.
- Approach
- Bulk INSERT … ON CONFLICT DO NOTHING through a raw pg.Pool, 500-row batches, four in flight, with deterministic keys.
- Outcome
- For bulk work, round-trip count dominates query complexity.
Zero job loss across a dual-Redis switch
0 jobs lost
- Problem
- Jobs scheduled before go-live sat in the local Redis; after switching to ElastiCache, workers listened elsewhere and contests never started.
- Approach
- A DUMP/RESTORE migration tool that copies every key — types and TTLs intact — both ways, run inside the live container by the go-live and go-idle scripts.
- Outcome
- “Always remember to do X” is not a control. Move the execution so the system guarantees it.
Closing the auto-submit coverage gap
full coverage
- Problem
- Auto-submit saved 140 of 1,323 live participants: it only read the active set, and an OOM crash had dropped the rest into the disconnected set.
- Approach
- Union both sets on time expiry, then submit in batches of 50 with Promise.allSettled so one failure can’t abort the rest.
- Outcome
- Once you find one instance of a failure class, predict its symmetric twin.
Debugging the Socket.IO protocol layer
50% → ~100%
- Problem
- The k6 client sent plain JSON over raw WebSockets; Socket.IO speaks Engine.IO v4 on top, so the server ignored everything. Half the sockets “worked”.
- Approach
- Verified the transport with curl, then rewrote the load client as a state machine speaking the real EIO4 handshake, namespace auth, framing and ping/pong.
- Outcome
- A clean transport connection says nothing about the protocol above it.
Rethinking autoscaling for WebSocket load
CPU ≠ load
- Problem
- Scaling triggered on CPU at 60%, but WebSocket load is memory- and IO-bound: CPU sat at 20% while the heap hit 95%. New instances also took ~7 minutes.
- Approach
- Three options weighed: a memory alarm, a custom connection-count metric, or pre-warming from the registration count before the contest starts.
- Outcome
- Pre-warming chosen as primary — the peak is known in advance — with CloudWatch as defence in depth.
Auditing every rate limiter, not just the broken one
600 ms ≠ 600 s
- Problem
- An OTP limiter guarded a login route with no OTP; every k6 user shared one IP and exhausted it. The audit found a second limiter missing a ×1000.
- Approach
- Removed the misplaced limiter, then audited every limiter in the codebase and fixed the unit mismatch in both places.
- Outcome
- Finding one bug of a class is the moment to look for its siblings.