MPI4AI
MPI4AI brings recurring AI communication patterns into community-governed Open MPI (opens in a new tab). It gives scientific MPI codes a portable path to GPU communication, AI-oriented collectives, compute-stream integration, and recovery without binding the research to one vendor library.
Research Program
MPI4AI is developing four connected capabilities:
- Native GPU communication for GPU-to-GPU data movement within Open MPI.
- AI-oriented collectives for communication patterns that recur in model training and inference.
- Compute-stream integration for coordinating communication with GPU computation.
- Fault tolerance and malleability for recovery and adaptive resource use.
The research targets neural architecture search with transfer learning, key-value prefix caching for large language model inference, and distributed data-parallel training. These workloads place different demands on synchronization, data movement, and recovery, making them useful tests for the proposed abstractions.
This communication layer complements the AI data paths studied in DeepIO and the task and provenance mechanisms developed in DTIO.
Open Standards
The project plans to contribute its extensions to Open MPI and seek their inclusion in the MPI-5 and MPI-6 standards. Standardization would make the abstractions available beyond a single implementation.
The collaboration brings together Illinois Tech, Tennessee Technological University, the University of Tennessee Knoxville, and SUNY at Stony Brook.