An agent running the return process for a mid-sized retailer approved a refund on an order that fell weeks outside the return window, then approved it again eleven more times before finance caught the pattern in a reconciliation report.
That refund was the obvious cost, but what it revealed was much more important: a system handed genuine authority before anyone confirmed it deserved it.
In this blog, we break down where agent deployments actually go wrong, why the failures cluster around integration, hallucination, and oversight rather than the model's capability, and what it takes to fix it before it costs you something bigger than a refund.
What "Deploying An Agent" Actually Means
Deploying an agent is not the same as running a good demo. A demo proves a model can follow instructions inside a clean environment built to show it off. Deployment means handing that same model real permissions inside a business that was never built with an agent in mind: messy data, systems that predate the model by a decade, and employees who assumed a human would always make the call. This means that the issue surfaces the moment you give software real authority inside a business that was never set up to manage it.
The Promise Is Real, The Failure Rate Realer
The scale of the opportunity is not exaggerated. McKinsey estimates that agentic AI could unlock between $2.6 and $4.4 trillion a year in value across more than sixty business use cases, and Deloitte found that 74% of organizations expect to be using agents at least moderately by 2027.
But what happens once they arrive is a different story. McKinsey reports that only 1% of organizations describe their AI adoption as mature, while 80% say they have already seen risky agent behavior, including exposed data and unauthorized system access. A widely cited MIT study reported by Forbes, found that 95% of enterprise generative AI pilots deliver no measurable return. A generic tool bolted onto a workflow it was never built for rarely survives contact with production while something built around how the business runs tends to hold up.
The reasons for the rare survival cluster into three things nobody puts on the demo slide, and none of them are about the model getting smarter next quarter.
Want to know which side of that statistic your use case is likely to land on before you commit a budget? Book a call with Calda and we'll look at where it currently stands.
Where The Real Costs Actually Show Up
Your Process Isn't What You Think It Is
When a company tries to automate a process, it assumes that process is written down in a document. In practice, the written procedure that actually exists is the polished version, while the actual work needed for the process to run, goes through judgment calls, informal messages, and workarounds that everyone knows about but nobody wrote down. Feed an agent the clean version and you get a system that does the wrong thing confidently, quickly, and with excellent logs to prove it.
The wiring underneath is just as unforgiving. Every point where an agent has to bridge two systems that were never built to talk to each other, whether through an export, a manual entry step, or an old workaround nobody remembers, is a place the automation can quietly break without telling anyone.
So what does that mean?
The exceptions are the job, not the noise around it. A process that looks 90% routine on paper often needs judgment far more often than that once someone actually watches it happen, and an agent that hits one of those cases either stops cold or sails through and gets it wrong silently.
That same silent failure shows up again once you look past integration and into how reliable the model's own behavior actually is.
If you are still working out what actually separates an agent from a chatbot, our explanation on what AI agents really are is worth a read before you go further.
Fast, Confident, And Frequently Wrong
Even with clean data and solid connections, agents carry a reliability problem ordinary software does not. Fiddler review of agent reliability puts agent failure rates in production at 70% to 95% depending on the task, and found that roughly 88% of agents that perform well in a controlled demo fail once they hit a real workflow. Run the same request twice and you can get two different answers, which is a behaviour called non-determinism, and it's exactly what makes these failures so hard to reproduce and debug.
The problem compounds the moment you chain steps together. If each step in a three-step agent workflow succeeds 70% of the time on its own, the odds of the whole chain finishing correctly fall to roughly 34%.
That kind of unpredictability is one thing when an agent is only generating text. It's a different problem entirely once that same agent has the authority to act.
The Security Problem Nobody's Watching
Give an agent the power to act and you have effectively onboarded an employee with broad access and no instinct for when something feels off. That is why they are often called "digital insiders," systems operating inside your infrastructure with real privileges that can cause damage by accident or because someone turned them against you. McKinsey catalogs failure patterns that barely existed in older software: one agent's mistake cascading into others it works alongside, a compromised agent talking a trusted one into handing over access it should never receive, and data quietly moving between agents with no record anyone can later audit.
None of this is hypothetical anymore. Security researchers at Unit 42, built two ordinary agent applications and ran nine attacks against them, succeeding at stealing credentials, reaching into the internal network, and pulling records out of connected databases. Their most uncomfortable finding is that prompt injection, hiding malicious instructions inside text an agent reads, is not even always required, since a loosely scoped agent can be talked out of its lane with an ordinary request. Even Anthropic's own published testing found an attacker could hijack an agent's behavior 78.6% of the time without safeguards, and still 57.1% of the time after hardening, a rate that would trigger an immediate freeze in most security reviews.
But Who Owns The Outcome?
The piece that gets missed most often is the simplest to state. When an agent makes a costly decision, a human still has to own what happens next. Only a small share of enterprises have mature governance for their agents, which means most are running systems with real authority and no clear line for which decisions it can make alone.That line has to exist before the agent goes live, not after the mistake forces the question.
How The Three Failures Feed Each Other
None of these three sit in isolation, and that is what makes a bad deployment so much more expensive than it first appears. A process automated from a simplified description rather than the fully accurate one produces confident, wrong outputs. Those outputs go unnoticed because reliability failures get treated as one-off glitches rather than a pattern worth investigating. And because nobody built the oversight to catch either problem early, the agent keeps its broad access until the mistake is expensive enough that someone finally asks who approved it.
Each failure is easy to write off on its own. A wrong refund gets treated as a fluke, a flaky output gets blamed on the model having an off day, and a permissions gap gets patched and forgotten. Nobody connects the dots because the dots live in different reviews, reported to different people.
What The Few Who Succeed Actually Do
The companies getting this right look almost boring from the outside, and the pattern has nothing to do with picking the smartest model available. In practice, it looks like this:
- Start with one tedious, low-stakes task, something like invoice matching or ticket triage, and learn to run an agent well before trusting one with real financial or legal weight.
- Map the real process by watching people work, rather than automating whatever the surface level documentation claims it is.
- Design the human handoff as a deliberate feature, building agents that escalate cleanly the moment they hit something uncertain.
- Decide what "good" looks like and build a way to measure it before writing a single instruction for the agent.
- Give each agent the narrowest access it needs and its own distinct identity, so a mistake stays contained rather than spreading.
One point on that list matters more than the rest, so it's worth pulling out rather than leaving it as one line among five. Mapping the real process and cleaning up the data behind it makes up the bulk of the actual work. The organizations that pull this off spend months on exactly that, clarifying the process and aligning the people involved before any agent goes live. Skip that process, and your pilot quietly disappears a few months in. Do it properly, and the same pilot is what ends up running in production a year later.
Prefer to have certified partners of OpenAI handle your agent integration? Book a FREE call with Calda
What It Actually Comes Down To
The demo was never going to reveal any of this, because none of these costs exist yet on the day an agent is switched on. They arrive gradually, in an exception the agent handled with total confidence and zero accuracy, in a chain of steps whose failure rate nobody multiplied out ahead of time, and in an access grant that sat unquestioned until it became a headline. Each one looks small enough to explain away alone, but together, they are the actual cost of deploying an agent without professional help, and they keep accruing long after the launch announcement goes out.
The companies that avoid this outcome are not the fastest movers or the biggest spenders. They either treated the messy process, the unreliable output, and the missing oversight as one connected problem from the start, or they brought in a team that already knew how to.
If you are looking for that team, reach out to Calda and we will look at what you are dealing with.
FAQ:
How long should we realistically budget for a first agent deployment?
Plan in months, not weeks. Timelines vary by company size, but mid-market companies typically move in roughly ninety days, while larger enterprises usually need closer to nine months once integration, data cleanup, and testing are accounted for. Any promise of a working production deployment in a few weeks is either scoped very narrowly or skipping steps you will pay for later.
How do we choose which process to automate first?
Pick something tedious, frequent, and low-stakes rather than your highest-value workflow. A process where a mistake costs little and gets noticed quickly is the right place to learn how your organization's data, systems, and people actually behave around an agent, before you hand one anything that touches money, legal exposure, or a customer relationship you cannot afford to damage.
What does a minimum acceptable oversight setup actually look like?
At minimum, it means a written line between decisions the agent can make on its own and decisions that require a person, a way to see every action the agent took after the fact, and a defined limit on what systems and data it can reach. Think of it as layered controls rather than one safeguard, and that framing holds even for a small first deployment.
Who is actually liable when an agent makes a costly mistake?
There is no single answer yet, and that is part of what has stalled deployments inside companies where nobody wants to put their name on the outcome. The practical fix is deciding this before the agent goes live, not after: define which actions require sign-off, keep a full record of what the agent did and why, and treat that record as the thing that will matter if a decision is ever questioned.
How do we know if our organization is even ready to start?
You are ready enough to begin a small pilot if you can name one process precisely, watch someone do it end to end, and identify who would need to approve an exception. You are not ready for a broader rollout if your data is inconsistent across systems, your process only exists as a surface level document nobody follows exactly, or nobody in the room can say who owns the outcome when the agent gets something wrong.
