Abstract
Multi-agent large language model (LLM) workflows are increasingly deployed for complex, multi-step computational tasks. Yet the fault tolerance of these systems (how individual agents respond to malformed inputs, tool failures, and subtle numerical errors) remains poorly characterized. We present the Agent Error Simulator (AES), a benchmark framework that evaluates LLM agent fault tolerance through controlled, declarative fault injection. AES intercepts agent-to-model HTTP traffic via a transparent proxy, injecting three categories of faults (format, logic, and tool call) at configurable points in a hierarchical planner-worker-aggregator workflow. We evaluate four models spanning 2B–32B parameters (Qwen2.5-Coder 32B, Gemma4 27B, Mistral 7B, and Granite4 2B) across 194 jobs organized into 10 test groups covering baseline behavior, multi-error injection, workflow scale, context pressure, compaction, and detection timing. Our results reveal a model-size-dependent logic error recovery threshold between 7B and 27B parameters; universal catastrophic failure under aggregator-level logic injection; and model-specific aggregator arithmetic deficits that interact with injected faults in non-obvious ways. We release the benchmark suite, injection tooling, and result data to support reproducible multi-agent fault tolerance research.