AI Agent Store Logo - Find Right AI Agent For The Job
AI Agent Store
find AI Agent for your use case

What Actually Breaks After the Demo And What We Changed Because of It

September 24, 2026 · 3 min read

The demo went perfectly. An agent reading support tickets, classifying them, drafting replies, and escalating the hard ones to a human. The client loved it. We shipped it to production two weeks later, and by day four it had refunded a customer twice for the same order because it couldn't tell that its own first refund had already gone through. Nobody had told it to check. Nobody had thought to.

That gap between demo and production is where most agent projects actually live, and it rarely shows up in the places people expect. The model didn't fail. The prompt was fine. What was missing was everything around the model: what it was allowed to touch, what happened when it got something wrong, and who was supposed to catch it before it reached a customer.

Here's what we changed in how we build agent systems after that, and why each change exists.

Permissions became a first-class design step, not a config afterward. Early on, we gave agents the same access a competent employee would have, because that felt reasonable. It isn't. An employee has judgment about when to use access. An agent has whatever guardrails you wrote down, and nothing else. Now every agent gets a permission model scoped to the smallest set of actions it needs for its actual job, written before a line of orchestration code, not bolted on when something goes wrong. A refund agent gets read access to order history and write access to one specific refund action, with a hard cap on amount and frequency per order. It cannot see or touch anything else, because there was never a reason for it to.

Recovery paths got built for the agent's mistakes, not just the user's. Most agent architectures we saw early on assumed the happy path and treated failure as an exception to log. That's backwards for anything touching money, data, or customer communication. We now build explicit recovery flows: if an agent's action needs to be undone, there has to be an undo, not a support ticket that says "figure it out." For the refund case, that meant a reconciliation check that runs before any refund action, not after. The agent asks "has this already happened" before it acts, every time, as a hard gate rather than a step it can skip under load.

Evals stopped being a launch-day checklist and became a running system. A one-time eval before launch tells you the agent worked on the day you tested it. It tells you nothing about the agent three weeks later, after the underlying data shifted, after an edge case nobody scripted for showed up in production, after a model update changed behavior slightly. We run evals continuously against a growing set of real production cases, not just the synthetic ones from before launch. When an agent's output drifts from what the eval expects, that's a signal to look, not something averaged away in a dashboard.

Human gates got placed based on cost of error, not based on how confident the agent claimed to be. Confidence scores are cheap to generate and easy to trust, and that combination makes them dangerous. We stopped gating human review on the agent's self-reported confidence and started gating it on what happens if the agent is wrong. Low-cost, reversible actions run autonomously even at moderate confidence. High-cost or irreversible ones, like a refund over a threshold or anything that touches a customer's account status, get a human in the loop regardless of how sure the agent claims to be. The agent's confidence is a poor judge of its own blind spots.

Someone owns the architecture, not just the prompt. The teams that get hurt worst are the ones where an AI tool wrote most of the orchestration and nobody senior actually understood the failure modes of what got built. AI tools are genuinely good at generating the boilerplate, the API glue, the standard classification logic. They are not good at knowing that a refund agent needs a reconciliation check before it acts, because that came from a production incident, not from a training pattern. Someone has to own that judgment. At Groovy Web, that's the model our AI agent development team builds around: AI handles the volume work, a senior engineer owns the parts that break in production.

If you're an agency starting agent work next month, the advice that would have saved us the most time is this: build the permission model and the recovery path before you build the happy path, because the happy path is the part that was never going to fail. Budget real time for continuous eval infrastructure, not a pre-launch checklist you run once and forget. And don't let confidence scores decide what needs a human, let the cost of being wrong decide that instead.

The demo will always look good. Production is where you find out what you actually built.

Try it on real work

Turn this idea into an agent that runs after your browser closes.

Start with one task and clear approval rules. We handle hosting, saved memory, restarts, and messaging connections.

Runs without your laptopBrowser + messaging appsCredits, keys, or subscriptionsMemory survives restarts

Plans start at $29/month. Cancel anytime.

Hosted agent

OpenClaw or Hermes

saved state
Browser
WhatsApp
Telegram
Slack
“I checked the inbox, handled the routine messages, and sent you the one question that needs a decision.”