insights
Notes from production
Longer write-ups on the problems I get called about most: Kubernetes stuck in a bad state, cloud bills nobody can explain, and RAG/LLM systems that worked in a demo and fell over in production. Written the way I'd explain it to another engineer, with the commands I actually ran.
Self-Hosted RAG in Production: The Issues That Never Show Up in the Demo
Why RAG systems that demo beautifully fall apart in production, and the concrete fixes for retrieval quality, hallucination, cost, and staleness that a demo never forces you to solve.
Kubernetes Cost Optimization: Where the Money Actually Goes and How to Cut It
A practical breakdown of where Kubernetes clusters bleed money and the concrete order of operations to cut the bill 20-50% without hurting reliability.
Kubernetes CrashLoopBackOff in Production: A Debugging Playbook That Actually Works
A field playbook for debugging CrashLoopBackOff in production without panic-restarting pods, built from real incidents where the "obvious" fix was wrong.
Living through one of these right now?
Book a free 30-minute call. We diagnose it together, and you walk away with a plan you can act on. You’ll get a straight read either way.