Skip to main content

StageOnce: Speculative Cross-Job Data Staging for LLM Agent Fleets

Authors: I. Yildirim, X.-H. Sun, A. Kougkas

Date: November, 2026

Venue: The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC'26)

Type: Poster

Abstract

Scientific workflows increasingly run as fleets of LLM agents analysing immutable datasets. On a shared compute node, independent staging repeats source reads and creates storage contention, while each job discovers its inputs during reasoning. StageOnce coordinates these speculative path predictions through a node-local daemon, coalescing concurrent misses into one immutable copy while preserving dataset versions and private outputs. Tools read staged inputs directly from the filesystem, without per-read coordination. Recorded-session replays in genomics and astronomy on SSD and HDD backends reduce fleet makespan by 31-51% at 12 concurrent jobs against independent staging. Logical staging remains one copy across every tested fleet size. An ablation isolates speculation, which reduces makespan by 8-49% against reactive shared caching. StageOnce beats independent staging when fleet demand exceeds what reasoning windows can hide and reuse amortises coordination overhead, while smaller working sets can make direct reads or prologue bulk copies preferable.

Tags

LLM AgentsData StagingSpeculative ExecutionHPC I/O