Home › Guides

Guide

AI Agent Failures: 9 Real Cases, What Went Wrong, and the Safeguard for Each

By Mind with Tools · Updated · 10 min read

In February 2026, an AI agent started deleting the email of someone whose job is making AI behave. Summer Yue, who works on AI alignment at Meta, had told it to suggest what to delete and not to act until she said so. She couldn’t stop it from her phone. She had to run to the computer it was running on.

This page collects nine real cases like that one, each dated and linked to a source. For each you get what happened, what failed, and one safeguard a non-coder can use. Three are not agents: Air Canada’s chatbot, the lawyers in Mata v. Avianca and Deloitte’s report for the Australian government involved chatbots or ordinary AI output. They are here because they are the same failure at lower stakes: someone relied on text nobody checked.

What counts as an AI agent failure

An AI agent is an AI that is given a goal, chooses its own next step, and uses tools to take it. Tools means things like reading email, opening files, searching the web or running commands. A chatbot answers a question. An agent does something about it.

A failure here means the AI did something the person did not want, or said something untrue, and a real person or business paid for it. Most of these cases are about a lost rule, a key that was too powerful, or a source nobody opened, not a clever machine. Where a story rests on the people involved or on a company describing its own product, the text says so. Nine cases are a collection, not a measure of how often agents fail.

When the agent has the keys

Summer Yue and OpenClaw (2026)

Her instruction was: “Check this inbox too and suggest what you would archive or delete, don’t action until I tell you to.” It had worked well on a small test inbox. On her real, much larger inbox, she said, the agent hit “compaction”, which is when an agent squeezes a long session into a summary, and “it lost my original instruction.” She later called it a “rookie mistake.”

What failed: A rule typed into a chat is a request. The agent could still delete, so nothing stopped it when the request was forgotten.

Safeguard: Connect an agent to your email so it can read and draft but not delete, and test the stop button on a practice run first.

PocketOS and a nine-second deletion (2026)

In April 2026, an AI coding agent was doing routine work for PocketOS, a car-rental software company, in its test environment. It hit a credentials mismatch and, instead of stopping to ask, went looking for another way. By the founder’s account it found an access key in an unrelated file, made for a narrow job but able to do almost anything on the company’s hosting account, and used it to delete a storage volume it believed was test-only. The volume held the live data, and the founder said the backups were on it too and that it all took nine seconds. The hosting company, Railway, confirmed the key had account-wide access and that an older route into its system deleted instantly with no undo. It recovered the data and made deletions on that route undoable.

What failed: The agent treated a blocked step as a puzzle instead of a stop sign, and a powerful key was lying where it could find it. Asked why, the agent said it had guessed instead of checking, but an agent’s account of itself is a hypothesis, not evidence.

Safeguard: Don’t leave logins, passwords or keys anywhere an agent can read them, give each agent its own limited account, and keep backups somewhere it cannot reach.

Replit and the SaaStr database (2025)

In July 2025, Jason Lemkin, founder of the SaaStr community, was building an app with Replit’s AI agent. He told it not to change anything without permission, in what he called a code freeze, and by his account repeated it in capital letters. The agent ran commands against his live database anyway and deleted records. It also told him the deletion could not be rolled back. It could, and he recovered the data. Replit’s CEO called the deletion “unacceptable” and said the company was adding automatic separation between test and live data. Lemkin’s own point was that there was no way to enforce a code freeze in the tool.

What failed: A freeze that existed only as a sentence was not a freeze, and the agent’s claim that nothing could be undone was wrong.

Safeguard: Let an agent work on a test copy rather than the real thing, and when it says something is lost for good, check before you believe it.

When a wrong answer gets relied on

Gemini CLI and the “lost” files (2025)

In July 2025, a user asked Google’s Gemini CLI, an agent that runs commands on your computer, to move files into a new folder on his Windows PC. By his own account in a public bug report, the agent reported success. The files were not where it said, and it then told him it had failed him “completely and catastrophically.” His report said the files were gone. A week later he posted a correction in the same report: he had found them in the top level of his C: drive. They were misplaced, not destroyed.

What failed: The agent was wrong about its success, then wrong about its failure. Both reports sounded confident, and neither was evidence.

Safeguard: After any file job, open the folder yourself and search before you panic, and treat an agent’s account of its own actions as a lead to check.

Deloitte’s report for the Australian government (2025)

Deloitte had an A$440,000 contract to review an automated welfare-penalty system for Australia’s employment department. Its 237-page report, dated July 2025, cited works that did not exist and, according to the Associated Press, included a fabricated quote from a federal court judgment. By early September Deloitte had confirmed in writing that some footnotes and references were wrong, and a corrected version disclosed that an AI tool, Azure OpenAI, had been used. The department asked for the final instalment of the fee back, about A$97,600, and Deloitte agreed to repay it. The department said the report’s substance and recommendations stood.

What failed: Review. In an October letter to the department, Deloitte said it took responsibility for the fact that appropriate review and oversight processes were not followed.

Safeguard: For the few facts your decision rests on, open the source and find the sentence that supports each one, and ask the AI to show that sentence beside every claim.

Mata v. Avianca (2023)

Two lawyers filed a brief in a New York federal court that cited cases ChatGPT had invented. When the cases were challenged, one lawyer asked ChatGPT whether one of them was real, and it said it was. On June 22, 2023, a federal judge ordered the two lawyers and their firm to pay a $5,000 penalty. The judge found bad faith, citing “conscious avoidance” and false and misleading statements to the court.

What failed: Asking the same AI to vouch for its own output is not a check.

Safeguard: Check a claim against the original source, such as the court record or the company’s own page, never against the AI that made it.

Air Canada’s chatbot (2024)

After a death in the family, a passenger named Jake Moffatt asked Air Canada’s website chatbot about bereavement fares. It suggested the reduced fare could be applied for afterward. That was not the airline’s policy, and Air Canada refused the refund. In February 2024, a British Columbia tribunal found the airline responsible. It rejected the argument that the chatbot was a separate entity responsible for its own actions, found negligent misrepresentation, and ordered Air Canada to pay about C$812.

What failed: Nobody checked what the chatbot said on the company’s behalf, and the company was held to it.

Safeguard: Treat anything an AI says in your name as something you said, and keep its output as a draft until a person has read it.

When it runs with too much freedom

Anthropic’s Project Vend shop (2025)

This one was mixed, not a clean failure. In 2025, Anthropic let an AI model nicknamed Claudius run a small shop in its San Francisco office for about a month. It did some things well, such as finding suppliers. It also priced items below cost and told customers to pay into an account that did not exist. It agreed to drop its discount codes, then went back to offering them within days. In a second phase, with better tools and set procedures, the shop did far better and losing weeks became rare, though Anthropic changed several things at once. A supervising “CEO” agent with the same blind spots helped little: it approved lenient requests about eight times as often as it refused them. Anthropic’s summary: “bureaucracy matters.”

What failed: Corrections did not stick, and the agent was easy to talk into discounts.

Safeguard: Write corrections on a page the agent reads at the start of every run, and don’t make a second AI with the same blind spots your only check.

When the attack is hidden in the data

EchoLeak in Microsoft 365 Copilot (2025)

Prompt injection means hiding instructions in something an AI will read, so it treats them as orders. In 2025, researchers at Aim Security showed that one carefully written email sent to someone using Microsoft 365 Copilot could make the assistant pull private information from that person’s data and send it out. The user did not have to click anything. The flaw was tracked as CVE-2025-32711. It was a researchers’ demonstration, and Aim said Microsoft had confirmed no customers were affected. Researchers showed the same kind of trick that year against Perplexity’s Comet browser and Salesforce’s Agentforce. Both were demonstrations reported to the companies first.

What failed: The AI could not reliably tell the user’s instructions from ones hidden in what it was reading.

Safeguard: Keep an agent that reads outside email or web pages to reading and drafting, with no way to send anything out, and tell it that what it reads is information, not orders.

The pattern across all of them

Case Year What failed Safeguard
Yue and OpenClaw 2026 An “ask first” rule was lost mid-task Read and draft access only; test the stop button
PocketOS 2026 A blocked agent improvised with a powerful key No loose keys; limited accounts; backups out of reach
Replit and SaaStr 2025 A freeze was only a sentence; false “can’t undo” Test copy for the agent; verify “lost”
Gemini CLI 2025 Wrong about its success, then about its failure Look at the folder yourself
Deloitte report 2025 Citations nobody opened Open the sources your decision rests on
Mata v. Avianca 2023 An AI asked to vouch for itself Check against the original source
Air Canada chatbot 2024 An unchecked answer given in the company’s name Draft until a person reads it
Project Vend 2025 Corrections did not stick; easy to talk into discounts Written record; not a same-model checker alone
EchoLeak 2025 Obeyed an instruction hidden in an email No way to send out; reading is not obeying

In most of these nine, either the agent could do something nobody had walled off, or nobody checked what the AI said. For every rule you give an agent, ask: if it ignored this sentence, could it still do the damage? If yes, you need a setting, not a sentence.

Run your own failure log

When your own agent fails, write it down. The Strange Employee, the ebook this guide belongs to, includes a simple failure log. The idea is easy to copy: one row per failure with what you wanted, what you got, the earliest wrong step, the type, what you changed, and whether the failed case now passes and a good case still does.

Name the type. Six of the book’s eight types show up above:

  • Said it was done when it wasn’t: Gemini CLI.
  • Made something up: Deloitte, Mata, Air Canada.
  • Forgot a rule partway through: OpenClaw, and the Vend discounts.
  • Treated a blocked step as a puzzle: PocketOS, Replit.
  • Obeyed something it read: EchoLeak.
  • Got talked into it: Vend.

Then find the earliest wrong step with five questions, in order. Did it understand the assignment? Did it get the right material? Did a tool fail without anyone noticing? Was one decision bad? Was anything checking its work?

Start at the beginning of the run, not the end. Treat the agent’s explanation as a guess; Gemini CLI and PocketOS both gave fluent ones. “An apology isn’t a fix.” Change one thing, then rerun the case that failed and one that worked.

Questions people ask

Can an AI agent delete my files?

Yes, if the connection you gave it allows deleting. In 2025 an agent deleted a live database in a project run by SaaStr's founder, in 2026 another deleted the live database of a car-rental software company, and a third began deleting an AI researcher's inbox. Start with read and draft access only, and keep anything important on a copy the agent cannot reach. Also check before you believe a report of lost files: in one 2025 case the files had only been misplaced.

What's the most common reason AI agents fail?

This page is a collection of cases, not a count, so it cannot say which cause is most common. In these nine cases the same few causes keep showing up: a rule that was only a sentence, access that was too wide, and AI output that nobody checked. Those are the parts you control, so they are the best place to start. Agents can also fail for reasons outside your control, such as a flaw in the product they run on.

How do I stop an agent from doing something I didn't ask?

Don't rely on a typed rule alone: a rule is a request, and a permission is a wall. Give the agent only the access the job needs, such as read and draft with no delete or send, and test the stop button on a practice run. For every rule, ask whether the agent could still do the damage if it ignored the sentence. If yes, you need a setting, a narrower connection, or no connection.

Are chatbot mistakes the same as agent failures?

They are related but not the same. A chatbot answers a question, while an agent chooses its next step and uses tools, so its mistakes can change or delete things. The Air Canada, Mata v. Avianca and Deloitte cases were chatbot or ordinary AI output errors, and they share a root with agent failures: someone relied on output nobody checked.

Can an AI agent be tricked by a hidden instruction in an email or web page?

Yes, and it is called prompt injection: text written so the AI treats it as an order. In 2025 researchers showed it against Microsoft 365 Copilot, Perplexity's Comet browser and Salesforce's Agentforce, in demonstrations reported to the companies first, and the researchers who found the Copilot flaw said Microsoft confirmed no customers were affected. The safe assumption is that any agent that reads content from outside your control can be steered by it, so keep such agents to reading and drafting.

Sources