What actually breaks when you ship AI agents to production
Every agent demo looks the same: a clean happy path, a confident tone, a task completed in one shot. Production traffic doesn't look like that. The gap between the two is where most of my last year went.
The first failure mode is silent context loss. Long conversations quietly drop earlier constraints once they fall outside the window a retrieval step actually pulls back in, and the agent starts contradicting itself with total confidence. The fix wasn't a bigger context window — it was forcing the agent to restate its working assumptions before every tool call, so drift becomes visible instead of silent.
The second is cost. A flow that costs a fraction of a cent in testing can spike hard once real users start asking follow-up questions that trigger extra retrieval and retries. Budgeting per conversation, not per request, is what actually caught this before it hit a bill.
The third is the hardest to explain to non-engineers: agents fail differently than APIs. An API times out or throws. An agent can return a perfectly formatted, entirely wrong answer, and nothing in the response tells you that happened. Most of the reliability work ends up being building the signals that used to be free.
Next post
Local-first is harder than it looks