Skip to main content

Agent Error Simulator: Fault Injection for Testing in Agentic Workflows

Authors: K. Bateman, J. Cernuda, L. Logan, B. Nicolae, F. Cappello, X.-H. Sun, A. Kougkas

Date: September, 2026

Venue: The 22nd IEEE International Conference on eScience (eScience'26)

Type: Conference

Abstract

Multi-agent large language model workflows are increasingly deployed for complex, multi-step computational tasks, yet their fault tolerance remains poorly characterized. We present the Agent Error Simulator (AES), a benchmark framework that evaluates LLM agent fault tolerance through controlled, declarative fault injection. AES intercepts agent-to-model HTTP traffic via a transparent proxy and injects format, logic, and tool-call faults at configurable points in a hierarchical planner-worker-aggregator workflow. We evaluate four models spanning 2B to 32B parameters across 194 jobs in ten test groups covering baseline behavior, multi-error injection, workflow scale, context pressure, compaction, and detection timing. The results reveal a model-size-dependent logic-error recovery threshold between 7B and 27B parameters, universal catastrophic failure under aggregator-level logic injection, and model-specific aggregator arithmetic deficits that interact with injected faults. We release the benchmark suite, injection tooling, and result data to support reproducible multi-agent fault-tolerance research.

Tags

AILLM AgentsAgentic WorkflowsFault InjectionFault ToleranceBenchmarking