One in 40 AI-Handled Support Emails Needs a Human to Come Back and Fix It
Who Gives A Crap sells toilet paper by subscription and gives away half its profits to sanitation charities. Decent company, decent people. You'd expect them to get the boring admin right.
But in July this year, one of its customers got an email saying their regular delivery of bamboo toilet paper was about to jump from $66 for 48 rolls to $69.50 for 24 rolls. Read that slowly. It's roughly the same money for half the rolls. That's not a price rise, mate. That's the price per roll more than doubling, from $1.38 to $2.90.
The customer asked if that was actually right. The company's "Customer Happiness Team" wrote back and confirmed it. Yes, your subscription's changing from 48 rolls at $66 to 24 rolls at $69.50, starting the following month.
Except it wasn't right. It was wrong twice, in two separate emails, from two separate replies. According to SmartCompany, which saw the correspondence, the actual change was a straightforward price bump on the same 48 rolls. Not a quantity halving dressed up as a small increase. A company spokesperson told SmartCompany there'd been a typo in the original email, and the new $69.50 was meant to cover 48 rolls, not 24. Both the original email and the reply that "confirmed" the wrong version were written by the company's AI email agent. Once the mix-up surfaced, Who Gives A Crap shut the tool down and sent a correcting email. "AI agents can be prone to error and we are working on the quality," the spokesperson said. That's corporate for "sorted, but that one got through."
What makes this one worth a proper look isn't the toilet roll maths. It's the second email. The first mistake was a slip, the kind any system, human or machine, can produce on a busy day. The second mistake was an AI agent looking at a customer who was directly questioning a number, and confidently telling them the wrong number was correct. No pause. No "let me check that". No escalation to a person. Just a calm, well-punctuated confirmation of a mistake, sent straight back to the person who'd caught it. CXM World picked the story up a week later and framed it as part of a wider pattern: AI agents doubling down when questioned, instead of flagging that they might be wrong.
And that pattern isn't rare. It isn't just a toilet paper problem either. A support platform called Robylon published a failure-mode breakdown of roughly 9.4 million inbound support emails over an 18-month window, across a dozen industries. Of the emails their AI resolved without a human touching them, 2.9% needed a person to come back and correct something afterwards. That ranged from 1.8% up to 4.6% depending on the account. Sounds small, sort of a rounding error, until you do the maths on volume. If your business handles 2,000 support emails a month, and half go through AI without a human eyeballing them, you're looking at roughly 29 wrong answers a month sitting in customer inboxes before anyone catches them. Multiply that by however many clients your agency runs support inboxes for, and it stops being a rounding error.
So what actually goes wrong? The same report breaks it down, and it's rarely the dramatic, invented-a-policy-from-nowhere failure everyone worries about. Genuinely fabricated details, what the report calls "hallucinated specifics", only made up 5.4% of the failure cases. The bigger culprits were more boring, and more dangerous precisely because they're boring: stale information (21.6% of failures), a job only half finished while the email says it's done (17.4%), and an AI confidently making a policy call it had no authority to make (11.8%). None of that needs the AI to invent anything. It just needs to be working off an out-of-date price list, or answering with total confidence when the honest answer is "I'm not sure, let me get someone."
Here's the bit that should worry you more than the toilet roll story itself. The failure that took the tool offline wasn't the drafting mistake. It was the follow-up. Most businesses testing an AI email agent check whether it can write a decent first response. Fewer check what it does when a customer pushes back and says "that doesn't look right". That's the exact moment an AI agent needs to hand off to a human. And it's the exact moment most setups never bothered to write a rule for.
I run this kind of automation for my own agency and for GHL partners, so this one landed close to home. If you've got an AI tool writing or answering customer emails, or you're weighing one up, here are three checks before you trust it with anything price or policy related:
Test the pushback, not just the pitch. Send your AI agent a message where the customer politely disagrees with a number it just gave. Does it check, or does it confirm its own mistake with the same confidence as the first email? If you've never actually tested this, you don't know the answer, you're guessing.
Anything with a number in it should trace back to a system, not a sentence. A price, a quantity, a date, none of that should come from the AI reasoning its way to a plausible figure. It should be pulled from your actual product data or CRM record every single time, with the email blocked from sending if that lookup fails or comes back empty.
Build the "I'm not sure" exit before you need it, not after. A hard rule that any question involving a price change, a policy exception, or a direct customer challenge gets routed to a person, not just "answered as best it can". That's not slowing your AI down. That's the bit that keeps you out of a SmartCompany headline.
None of this is an argument for ditching AI email agents. They're still faster and cheaper than a human writing every reply from scratch, and most of them get the routine stuff right, most of the time. So the actual risk was never the AI writing a wrong sentence once. It's the AI being asked "are you sure?" and saying yes anyway, with nobody watching for that exact moment.
Want to see how this could work in your business? Book a call and let's talk about where you're at and what's possible.
Brewed by Steven, poured by Viktor
About Steven Tann: Steven helps business owners build systems that run themselves using AI. After 10+ years helping 7,000+ businesses and building his own autonomous operations, he's the bloke who actually does it, not just talks about it. Find out more at steventann.com.