Blog

AI products in production: agents, evals, and the decisions behind them.

10 essays on building AI in production, written while shipping it.

Model safety ends where the customer's WhatsApp begins
Model-layer safety is the lab's job. Deployment-layer safety is mine: pointing a safe model at a real customer's WhatsApp without getting burned.
The system behind 230K interactions a day
The two numbers on my CV, unpacked: what counts as an interaction, which layer the 99.65% is measured at, and the eval, guardrail, and cost subsystems that keep both honest.
Creating a tenant is an INSERT
Voltade's third swing at the same problem. Studio asked SMEs to build, Vobase had us building for them, and Volty bets they'll configure. What changed each time, and why.
How I keep production agents on the rails
Production agent safety is not about jailbreaks. It is about an agent confidently doing the wrong thing on a customer's WhatsApp for six hours before anyone notices. Here are the failure modes that actually happen, the guardrails that work, and how I prove the guardrails are working.
How I Evaluate AI Agents (and Why Most Teams Get It Wrong)
Most agent evals measure the wrong things. After running two agents in production for six months, here's the framework I actually use, with real metrics, LLM-as-judge calibration data, and the $300 lesson that started it all.
The Trust Budget
Autonomy isn't a property of the agent. It's a budget you allocate per action, priced by how reversible the action is and how big the blast radius gets. Here's the framework I use across every agent I run.
Why Studio didn't work
We built a no-code agent builder so SMEs could build their own agents. Almost nobody did. Why Studio failed, and what we built instead.
Seven agents on a Mac Mini: four months of breaking my OpenClaw harness
I run seven personal AI agents on a Mac Mini in my flat. They book my gym classes, manage my inbox, handle outreach for two academies, and watch Voltade customer groups. The harness took four months to stabilise. The two things that fixed it were scope discipline and deterministic flows. Here is what each agent does, what broke, and the patterns that finally stuck.
The 20% That Is the Business
Templates handle the common 80%. The remaining 20% is every customer's actual business. Here's what that 20% looked like for one bakery, one day, five bugs.
I Gave Claude $100 and Told It to Trade Crypto
Building an autonomous crypto trading bot with Claude Code, $100 in USDC, and a lot of guardrails. From first trades to losing everything, then rebuilding the strategy from scratch based on actual research.

Navigate with j/k, open with Enter