What happens when an AI agent makes a mistake with a customer

By Precipitate · 6 September 2026

Close-up of a woven fabric with one thread pulled loose and tied off with a small red knot where it frays, set against a dark background

Yes, and the honest answer is that any agent making decisions and taking real actions without a person checking every step will eventually get something wrong: a wrong answer or a refund it should not have approved. The real question for an owner-operator is not whether that happens. It is whether the system catches its own mistake, tells someone, and leaves a record you can check afterward.

The mistake is real, not hypothetical

Air Canada found this out in front of a tribunal. Its chatbot told a grieving customer he could claim a bereavement fare after he had already bought a full-price ticket. That information was false. When the airline refused the refund and argued the chatbot was a separate entity responsible for its own words, the tribunal disagreed and ruled Air Canada liable, as CMSWire reported. The precedent held: the business owns what its agent says, not the software underneath it.

Cursor ran into a faster version of the same problem. Its AI support agent, nicknamed Sam, invented a policy limiting customers to one device per subscription and described it as a security measure, according to CMSWire. No such policy existed. The fabricated rule spread through developer forums within hours, and subscription cancellations followed before a person at Cursor could correct the record.

The legal ground under agents is still moving

Agentic payments sharpen the liability question further. A white paper from the Consumer Bankers Association and the law firm Davis Wright Tremaine, covered by The Financial Brand, warns that the consumer protections built into the Electronic Fund Transfer Act may not apply cleanly once an agent initiates the transaction instead of a person. Customers could end up liable for their own agent's mistakes, and banks, the paper notes, will be the first place people call when a payment goes wrong.

None of this is settled law yet, but the direction is clear enough to plan around now. If your agent books, orders, refunds, or replies on your behalf, the customer on the other end will treat its mistake as yours, not as a bug in someone else's software. That is not an argument against using agents. It is an argument for knowing, before you deploy one, exactly what it does the moment it is unsure.

Automation without a backstop makes it worse

Money spent has not guaranteed good outcomes. Organizations put $47 billion into AI initiatives in the first half of 2025, and 89% of that spending produced minimal returns, CMSWire reported, pointing to compliance complexity and organizational chaos as the more common cause than the underlying model. Customer service came up as one of the riskiest places to apply AI without enough care, because a bad answer lands on someone who is often already frustrated.

Results were better, per the same reporting, where AI supported a human agent rather than replacing one: faster responses and more empathy, with a person still accountable for the interaction. For a small operation, that argues for a specific split: hand the agent the repetitive volume, and keep a person as the one who steps in for anything upset or outside the normal pattern. Our breakdown of workflow automation versus agentic decision-making is a useful gut check for where that line sits in your own business.

Some mistakes cost more than a wrong answer

The Air Canada case was not really about a wrong fare policy. It was about a customer grieving a death, being told something false at the exact moment he needed the answer to be right, part of what CMSWire described as the emotional damage an AI mistake can cause. A pricing error is annoying. A false answer delivered to someone at a hard moment is the kind of mistake that ends up in front of a tribunal, or on social media, before it ends up on your desk.

This is the honest limit worth naming: an agent can be built to recognize that a situation is emotionally loaded or unusual, and to hand it to a person instead of answering on its own. It cannot replace the judgment that situation actually needs. Any vendor or system that claims otherwise is selling you the part that is easy to automate and hiding the part that is not.

What changes when the agent can catch itself

An agent built to operate, rather than just respond, does two things a simple chatbot does not: it checks the outcome of its own action, and it knows which situations sit outside its authority and stops to ask a person before they turn into a promise it cannot keep. That has to be a design decision from the start, not something added later. A refund that matches a hundred prior refunds can be approved automatically. A refund tied to the kind of situation that put Air Canada in front of a tribunal should not be.

Most agent failures actually happen at the seam between the agent's reasoning and the real system it acts on: a calendar, a payment processor, a CRM. We wrote about that seam separately in where AI agents fail when they touch real systems. If you are evaluating a vendor rather than building this yourself, our notes on how to evaluate an AI automation vendor cover the specific questions worth asking about that seam before you sign anything.

What to check this week

Pull up whatever agent already touches your customers, whether it answers the phone or processes refunds for you, and ask one question: what does it do the moment it is not sure? A good answer names a specific action, such as flagging the message or holding the transaction, and a specific place you can go look afterward.

Then go look at that log, if one exists. If it does not, that absence is the real finding, ahead of anything the agent might say next week.

Sources

Want this answered for your own business?

Get a straight answer

More from the blog

Straight answers