Silent Failure Taxonomy for Production AI Agents

A framework for finding AI agent failures that don't trigger any alerts.

Staff Writer · · 10 min read
Cover illustration for “Silent Failure Taxonomy for Production AI Agents”
Agent Failure Modes · October 8, 2026 · 10 min read · 2,303 words

Production AI agents fail most dangerously when nothing fails at all, at least nothing that an infrastructure dashboard can see. The agent returns a successful status, the latency graph stays flat, the error monitor shows green, and the agent has already issued a refund that violates policy, deleted a table it was never authorized to touch, or cited a regulation that does not exist. This piece builds a taxonomy for that class of failure, organized by the mechanism that produces it rather than the symptom a user happens to notice, so engineering teams can find each pattern in their own traces and know which part of the system needs fixing.

Why production AI agents fail without triggering any alert

A crash is a conversation between a system and its operator. The system says something broke, and the operator believes it, because the signal and the failure are the same event. Agent failure breaks that arrangement. An agent can complete its assigned task, return a well-formed output, and still have done the wrong thing entirely, with no part of the infrastructure stack positioned to notice the difference.

The gap is structural. Infrastructure metrics measure what happened at the level of the system: did the process return, did the call complete, did the response arrive inside the latency budget. None of those measurements ask whether what happened was correct relative to what the agent was actually supposed to do. A 500 error is detectable because the infrastructure agrees something went wrong. A hallucinated refund policy, a deleted database, or a fabricated regulatory citation all return success at the infrastructure level, because success and correctness are different properties and only one of them is instrumented. Dexing Liu's analysis of more than 100,000 agent interactions found this same pattern recurring across different architectures, different models, and different task domains: the task completes, the output is degraded, and no error fires. That recurrence across unrelated systems is what makes the failure a structural property of the agent category, not a bug in one company's implementation.

This taxonomy is scoped deliberately to non-adversarial failure. Microsoft's AI Red Team updated its failure mode taxonomy in June 2026 to cover categories like agentic supply chain compromise, goal hijacking, inter-agent trust escalation, computer-use agent visual attacks, session context contamination, MCP and plugin abuse, and capability or architecture disclosure. Those are real and growing risks, but they all assume an attacker: injected input, a malicious actor, a compromised dependency. The failures this piece catalogs need none of that. They emerge from an agent doing its job under normal conditions, with no adversary in the loop and no red team exercise positioned to surface them, because there is nothing to attack and nothing to defend against. The agent just gets it wrong, quietly, while every system around it says everything is fine.

Priyanka Bajaj's work on production AI failure gives this phenomenon a formal name: evaluation blindness, defined as a measurement function producing a value indistinguishable from a non-failing state while the system is actually failing, with no auxiliary signal around to flag the gap. In her six-class taxonomy of production failures, the Operational class is silent by definition: there is no version of an operational failure that trips an alarm on its own, because the entire category is defined by the absence of a distinguishing signal.

Why existing monitoring frameworks cannot see these failures

Standard observability was built for deterministic systems. Error rates, latency histograms, and pass/fail health checks all assume something: a wrong output and a correct output will show up as different signals somewhere in the stack. That assumption holds for a web server returning a malformed response. It does not hold for an agent, because agents are allowed to reach the same instruction through different paths, and a wrong answer can look, structurally, exactly like a right one.

The same prompt can trigger a different sequence of tool calls on two separate runs, and both sequences can return outputs that pass every syntactic check even when they are semantically wrong in ways no schema validator would catch. Binary monitoring built on top of that variability doesn't produce safety, it produces false confidence, which is arguably worse than no monitoring at all, because it tells a team to stop looking. Bajaj's review of 50 real-world production incidents found that 53% of verifiable public incidents were silent: the measurement function that should have caught the problem returned a value indistinguishable from normal operation, and nobody downstream had a reason to look twice.

Existing failure taxonomies make the problem harder to fix because most of them are tied to a single benchmark. They catalog fine-grained failure modes inside one evaluation setup, so they help there, but none of the categories carry over once you move to a different production stack. Scale AI's interaction-centric taxonomy, published in July 2026, names the consequence directly: the same visible failure might call for model post-training, harness engineering, environment redesign, or benchmark repair, depending entirely on where it actually originated, and without a shared structure for naming where failures originate, teams end up fixing the wrong component.

That absence of shared vocabulary is where the damage compounds. A write-up describing a production incident as "the agent made a bad decision" gives an engineering team nothing to act on. The WOWHOW taxonomy's framing, described in a DEV Community post, makes the opposite case well: a name like Scope Creep Execution tells a team what pattern to search their logs for and which instrumentation point would have caught it. If description stays vague, so do post-mortems, mitigations turn generic, and the same failure keeps recurring. The rest of this piece is an attempt to build the vocabulary that prevents that.

Organizing a silent failure taxonomy by mechanism, not symptom

Sorting agent failures by symptom, meaning wrong output, slow response, angry user, groups together failures that have nothing in common except how they were noticed. A wrong output caused by a bad input needs a fix different from one caused by a flawed plan, but a symptom-based log entry makes them look the same. Sorting by mechanism instead, by the layer of the system where disorder actually accumulated, produces categories a team can act on.

Liu's Entropy Principle paper gives you the layered structure you need to sort failures this way. It identifies five layers across an agent's lifecycle: Transmission (L1), Memory (L2), Execution (L3), Coordination (L4), and a fifth layer, each with its own failure dynamics and its own measurable signals of accumulating entropy. Crucially, failures do not stay put in the layer where they start. A memory-layer failure can show up later as a broken execution-layer output. A transmission-layer failure corrupts every agent downstream of it in a multi-agent chain. That displacement is why a taxonomy built around surface symptoms sends teams to the wrong place: the symptom appears three layers away from the mechanism that produced it.

Scale AI's interaction-centric taxonomy adds a second axis that makes the categories even more useful. It assigns each of 41 failure modes to the interaction edge where it originates (model-context, model-tool, model-memory, model-local environment, model-external environment) and to a fault side, model or harness, that tells a team where the actual repair belongs. A model-side failure points toward post-training. A harness-side failure points toward scaffolding and tool-integration work. An environment-side failure points toward redesigning the conditions the agent operates in. In the paper's Claude Code example, an agent that ignores an instruction it was given earlier might be failing because the harness's context compaction silently dropped that instruction (a harness-side problem), or because the instruction was still present and the model simply failed to follow it (a model-side problem). The visible behavior is identical in both cases. Only the fault-side assignment tells an engineering team which one they are actually looking at, and which fix would do anything.

The categories that follow use mechanism as the primary axis and fault-side as the secondary one. Each failure mode is named by what broke, which layer it broke in, and which component is responsible for fixing it, so that finding the pattern in a trace comes with an answer about what to do next.

Diagram: Where Silent Failures Actually Originate: Five Agent Layers. Visualizes: Show the five layers of an agent's lifecycle from Liu's Entropy Principle framework — L1 Transmission, L2 Memory, L3 Execution, L4 Coordination, and the fifth layer —…

Perception failures: the agent acts on a wrong model of its inputs

Almost every failure that gets blamed on bad reasoning actually starts upstream, as a perception failure. An agent that acts on a wrong model of its inputs is reasoning correctly from false premises, and the output looks like a reasoning failure only because nobody checked the premises first. If you fix the planning or execution layer downstream of a perception failure, nothing changes, because the plan was built on bad ground truth from the start.

The WOWHOW taxonomy names four perception failure modes. Ambiguity Collapse (Mode 1) is rated High severity: the agent encounters a spec with more than one reasonable reading, picks one silently, and never surfaces the alternatives it discarded. Its signature in a trace is an inference step with no recorded alternatives, no explicit statement of what the agent assumed the spec meant, just a single resolution that looks deliberate but was never checked against the other options. Log every branch where the agent resolves an ambiguous term to a specific referent and flag any resolution that happened in one step with nothing else considered. The mitigation is a pre-task clarification gate: a structured interpretation block the agent has to produce before it is allowed to call a single tool.

Salience Inversion (Mode 3), also rated High severity, is the mirror image: the agent fixes its attention on a minor detail while missing the constraint the whole task actually depends on. Scale AI's interaction-centric taxonomy frames both of these as failures on the model-context edge. The fault sits on the model side when the agent had the option to ask for clarification and didn't. It sits on the harness side when the context compaction or retrieval mechanism never delivered the right material. Retraining a model does nothing for a context window that was poisoned by a harness-side retrieval bug. The fault-side assignment has to happen before anyone commits engineering time to a fix.

Planning failures: the agent builds a strategy that is wrong before execution begins

Planning failures are dangerous in a specific way that perception failures are not: they are internally coherent. An agent with a flawed plan commits to a multi-step strategy with high apparent confidence, and nothing about the plan's own structure signals that anything is wrong with it. It reads like a good plan. It just isn't one.

The WOWHOW taxonomy names four planning failure modes. Phantom Dependency Assumption (Mode 5), rated High severity, is the one that shows up most often in production incidents involving hallucinated tool calls: the agent builds its plan around a library, API, or helper function that does not actually exist in the environment it's running in. Its signature is a plan that references a resource the agent never verified. Validate every dependency named in a plan against the live environment before letting execution begin.

Horizon Truncation (Mode 6), also High severity, is a plan that solves the immediate sub-goal but quietly invalidates a step further down the sequence, or blocks a downstream agent that depends on the state the plan leaves behind. Its signature is a plan that achieves its stated goal but creates conditions that break something scheduled to happen next. Simulate a plan through to its end state before you commit to it, and check specifically for invariant violations in later steps.

Confidence-Evidence Mismatch (Mode 7) carries the highest severity rating in the planning category, Critical. The agent can produce a detailed, well-formed, multi-step plan while it holds close to zero verified evidence that any of its key assumptions actually hold in the environment. The plan looks finished. Nothing about its form distinguishes it from a plan built on solid verification. Require a confidence-evidence register alongside every plan, one piece of verified evidence per key assumption, so a plan with zero entries in that register is visibly different from one that has done the work.

You can see its signature directly in tool call logs, as repetition without progress.

Liu's entropy framework places planning failures at the task execution layer, L3, where errors do not announce themselves immediately but accumulate across every step that follows. The Entropy Principle formula, S(t) = S₀·eᵅᵗ, captures why a planning failure gets worse the longer it runs undetected: each additional execution step adds its own disorder on top of the error the plan started with, so the cost of catching a planning failure late is categorically higher than the cost of catching it before execution starts.

Execution failures: the agent does something other than what it planned

Execution failures are the most consequential category in this taxonomy, because by the time one occurs, every upstream check has already passed. The inputs were read correctly. The plan was sound, verified, and free of phantom dependencies. And the agent still does something that falls outside the scope of what it was actually asked to do, or leaves the system sitting in a broken intermediate state partway through a multi-step operation.

This is the layer where the gap between infrastructure health and task correctness becomes impossible to ignore. An agent that deletes a production table instead of a staging table has executed a tool call successfully: the call returned, the database responded, no process crashed. Nothing in that sequence resembles a system error, because nothing about it was one. The deviation lives between the plan and the action taken, in the layer where intent is supposed to translate into a specific, bounded operation, and it is precisely that layer where a sound plan and a correct set of inputs still leave room for the agent to do something else.

Sources

  1. Silent Failure in LLM Agent Systems: The Entropy Principle and the Inevitable Disorder of Autonomous Agents
  2. Microsoft Updates Taxonomy of AI System Failure Modes
  3. Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment
  4. Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

More in Agent Failure Modes