
Event Sourcing in Production: Three Years of Append-Only Truth
Event sourcing solves some auditability problems cleanly and creates new ones around schema evolution and event replay performance. An honest retrospective.
Clean code practices, system design, architecture patterns, and development methodologies

Head-based sampling at 1% misses the rare slow requests you actually care about. Tail-based sampling solves this and introduces a different set of operational challenges.

Event sourcing solves some auditability problems cleanly and creates new ones around schema evolution and event replay performance. An honest retrospective.

We migrated to OpenTelemetry from three different proprietary SDKs. The cardinality explosion from naively porting existing metrics nearly took down our Prometheus cluster.

Deep dive into microservices architecture with practical design patterns, implementation strategies, and real-world examples for building scalable distributed systems.

The Istio / Envoy sidecar adds measurable latency. On p99 for our mix of RPCs, the overhead was 4 ms — acceptable for some services and not for others.

Feature flags accumulate. After eighteen months and three product releases, our flag system had 140 active flags and three engineers who understood any given flag's interaction graph.

We migrated 14 repositories into a Nx monorepo. The codemod tools handled 70% of the work. The remaining 30% took most of the calendar time.

We didn't notice egress was 18% of our cloud bill until a board review forced a line-item audit. The architecture changes that followed were uncomfortable but necessary.

Consumer lag is a lagging indicator. By the time it appears on a dashboard, the ordering guarantees you depended on may already be violated. How we set up lag alerting.

Five-person team, two time zones, one on-call rotation that doesn't burn people out. Here is what we landed on and what we changed after six months.

Most postmortems get filed and forgotten. The difference between a useful one and a formality is specificity of contributing factors and follow-up ownership.

The worst incidents I've been part of were slow not because the fix was hard but because decision authority was unclear. A runbook that helps with that.