Failure Modes

Browse the canonical taxonomy, then drill into a mode page for definitions, examples, detection approaches, mitigations, and related patterns.

104 modes · 13 categories

Fabrication

8 modes

The model invents facts, citations, or details that have no support in its sources or available evidence.

Faithfulness

6 modes

The response misrepresents the input, source material, or the model's own earlier statements in ways that distort their meaning.

Freshness

4 modes

The model presents outdated, time-sensitive, or version-specific information as if it were current.

Retrieval

9 modes

The system fails to fetch, rank, filter, or apply the right external evidence.

Context

6 modes

The model loses track of information in long inputs, missing, diluting, or overwriting details that matter.

Memory

7 modes

State carried across turns or sessions is missing, corrupted, out of date, or applied where it doesn't belong.

Control

12 modes

The system fails to follow instructions, respect constraints, stay in role, produce the required output format, or behave consistently across phrasings, runs, and model versions.

Instruction Noncompliance

Fails to follow an explicit, applicable instruction from the governing prompt, user request, or task procedure.

Control

Constraint Violation

Breaks a stated limit, requirement, policy, boundary, allowed action set, or output constraint that should govern the task, including dropping a constraint partway through multi-step reasoning or execution.

Control

Format Failure

Produces an answer in the wrong shape, organization, medium, style, or presentation format for the requested output.

Control

JSON/Schema Failure

Emits invalid JSON, malformed structured data, or output that does not satisfy the required schema.

Control

Refusal Overreach

Refuses, blocks, or safety-wraps a request more broadly than policy, risk, or context requires.

Control

Refusal Underreach

Fails to refuse, limit, redirect, or safety-constrain a request that requires stronger boundaries.

Control

Role Confusion

Misunderstands or drifts from its assigned role, persona, authority boundary, operating mode, or relationship to the user and other agents.

Control

Priority Confusion

Applies the wrong hierarchy among system, developer, user, tool, policy, memory, or task-level instructions.

Control

Clarification Underuse

Proceeds without asking when missing or ambiguous information materially affects correctness, safety, or user intent, committing to an interpretation that should have been confirmed first.

Control

Clarification Overuse

Asks the user for clarification when the task is already sufficiently specified, stalling on details the system could reasonably infer or safely proceed without.

Control

Prompt Brittleness

Produces materially different answers when the prompt is reworded, reformatted, or rerun, even though nothing meaningful about the request changed.

Control

Model Update Regression

Behavior, quality, or compliance shifts when the underlying model is updated, swapped, or silently revised, breaking prompts and pipelines tuned against the previous version.

Control

Reasoning

9 modes

The model errs while interpreting goals, weighing constraints, planning steps, or checking its own work.

Reasoning Error

Draws the wrong conclusion through invalid inference, faulty assumptions, mistaken causal reasoning, unsupported logical steps, or framing the problem with the wrong representation or abstraction.

Reasoning

Arithmetic Error

Computes or transforms numeric inputs incorrectly, including arithmetic, aggregation, unit conversion, comparison, or formula application.

Reasoning

Goal Misinterpretation

Solves the wrong problem because it misunderstood the user's objective, success condition, scope, or intended outcome.

Reasoning

Planning Failure

Builds an ineffective, unsafe, incomplete, or poorly ordered plan for achieving the user's goal.

Reasoning

Step Omission

Leaves out a necessary reasoning, verification, retrieval, tool, communication, or execution step needed for the task to succeed.

Reasoning

Compositional Failure

Fails to combine multiple facts, constraints, operations, sources, or subproblem results into a coherent answer.

Reasoning

Error Accumulation

Allows small mistakes, approximations, stale assumptions, or unverified intermediate results to compound across a multi-step task until the final output fails.

Reasoning

Verification Failure

Does not adequately check whether intermediate steps, tool results, cited evidence, assumptions, or the final answer are correct before relying on them.

Reasoning

Overthinking

Spends far more reasoning than the problem warrants — long deliberation on trivial questions, redundant re-derivations, or second-guessing that talks the model out of a correct answer.

Reasoning

Tools

9 modes

The system skips a needed tool, misuses one, invokes it unsafely, or mishandles its results.

Agency

8 modes

The agent miscalibrates initiative, stopping short of completing the task or acting well beyond its scope.

Security

9 modes

Adversarial inputs manipulate the system into leaking protected information or behaving unsafely.

Prompt Injection

Lets untrusted input attempt to override, weaken, or redirect the system's intended instructions, policies, tool-use rules, or data boundaries.

Security

Jailbreak

Manipulates the model into bypassing safety, policy, or behavioral controls that should remain enforced.

Security

Indirect Prompt Injection

Lets retrieved, browsed, uploaded, tool-supplied, or otherwise external content carry malicious instructions into the model's context.

Security

System Prompt Leakage

Reveals hidden system, developer, policy, tool, chain-of-thought, or other protected prompt content that should not be exposed.

Security

Sensitive Information Disclosure

Exposes secrets, credentials, personal data, confidential business information, private user content, or other protected information.

Security

Data Exfiltration

Enables unauthorized extraction, transfer, or reconstruction of protected data from tools, files, memory, retrieval systems, databases, or context.

Security

Insecure Output Handling

Produces output that is unsafe for downstream rendering, execution, storage, parsing, logging, or human trust without sanitization or validation.

Security

Unbounded Consumption

Consumes or triggers excessive tokens, compute, time, bandwidth, money, API quota, storage, or external resources without adequate limits or stopping conditions.

Security

Supply Chain Vulnerability

Introduces or recommends risk through compromised, malicious, abandoned, typosquatted, untrusted, or poorly pinned dependencies, tools, plugins, models, datasets, or upstream content.

Security

Alignment

7 modes

The model prioritizes pleasing, persuading, or mirroring the user over truthfulness and safety.

Response Integrity

10 modes

The final answer misses the mark on task fit, audience, locale, or actionability, even when the underlying content is sound.

Verbosity Failure

Provides more detail, repetition, caveats, background, or explanation than the task, user, medium, or decision requires.

Response Integrity

Incompleteness

Leaves out information, constraints, caveats, steps, options, or outputs needed to satisfy the user's task.

Response Integrity

Irrelevance

Includes content that does not materially help answer the user's question, solve the task, or support the needed decision.

Response Integrity

Genericism

Gives vague, boilerplate, or template-like guidance that is too nonspecific or abstract for the user to act on, instead of concrete help grounded in their task.

Response Integrity

Audience Mismatch

Uses terminology, assumptions, depth, examples, tone, or framing that does not fit the intended reader's expertise, role, goals, or context.

Response Integrity

Concision Failure

Compresses the answer so aggressively that necessary context, reasoning, caveats, instructions, or operational detail is lost.

Response Integrity

Poor Structure

Organizes information in a way that makes the answer hard to scan, compare, execute, or verify.

Response Integrity

Calibration Failure

Misstates confidence, uncertainty, evidence strength, risk, tradeoffs, or likelihood in the final answer.

Response Integrity

Localization Failure

Ignores or misapplies locale-specific language, spelling, units, currencies, laws, formats, idioms, accessibility expectations, or cultural conventions.

Response Integrity

Output Truncation

Delivers a response cut off mid-thought by a token limit, stop sequence, or timeout — often without the system or the model registering that the output is incomplete.

Response Integrity