Post-mortems, benchmarks,
and the occasional opinion.
Why we picked EPYC 9754 over Sapphire Rapids for our Enterprise tier
128 Zen 4c cores at 2.25 GHz base. On our workload mix (heavy Redis, Postgres, and Nginx), the 9754 landed 22% ahead of a similarly-priced Xeon Platinum 8592+. Here's the exact benchmark methodology and the raw numbers.
Warsaw: what the launch actually looked like from the NOC
The first 72 hours after a new PoP goes live are always a little tense. Here's the timeline of what we caught, what we missed, and the one thing we're changing for the next launch.
Gen5 NVMe in practice: 3× the sequential, 1.4× the random
Vendor spec sheets promise 14 GB/s. In a real multi-tenant fleet with QoS and encryption at rest, you get closer to 3.4 GB/s. That's still a huge jump, but let's talk about where the gap actually goes.
Against per-hour rounding: an argument for billing to the second
The main reason to round up to the hour is that your billing pipeline can't handle finer granularity. That's a solvable engineering problem, not a defensible business model. Here's how we do it.
Post-mortem: the 11-minute SYD API latency incident (Jul 4)
A control-plane migration that should have been transparent doubled API latency in a single region for 11 minutes. Full timeline, what our monitoring missed, and the two changes we made afterward.
Why our whole internal fleet runs NixOS
We host on our own platform, of course. But we run our internal fleet — control plane, monitoring, build agents — on NixOS. Here's the honest tradeoff analysis three years in.