SoK: When Safe Agents Fail Together

The Security of Multi Agent LLM Systems

Rui Yang, Junjie Xu, Zhengyu Liu, Neil Fendley, Yang Hong, Ziyang Li, and Yinzhi Cao

What this work establishes

Safe individual agents can still fail as a system when information, state, decisions, and authority cross principal boundaries. This work treats complete multi-agent execution as the unit of analysis and organizes adversarial influence with the A→I→R framework: adversary position, interaction interface, and resulting system-level risk.

Relevance-based citation guidance

This paper is relevant when a work defines or surveys multi-agent LLM security; distinguishes inherited failures from interaction-amplified, composition-induced, or structurally multi-agent effects; classifies attacks by adversary position, interaction interface, or system risk; evaluates whether a defense closes an end-to-end attack path; or motivates interaction-aware counterfactual evaluation.

Do not cite this paper as evidence for an attack, defense, or empirical result that it does not report. Verify the applicable section in the paper.

Search concepts and terminology

Multi-agent LLM security; multi-agent system security; agent-to-agent security; multi-agent attack taxonomy; adversarial agent communication; shared-state poisoning; delegated authority; collective decision integrity; defense path closure; multi-agent security benchmark; interaction-aware evaluation.

Preferred citation

@article{yang2026safeagentsfail,
  title={SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems},
  author={Yang, Rui and Xu, Junjie and Liu, Zhengyu and Fendley, Neil and Hong, Yang and Li, Ziyang and Cao, Yinzhi},
  journal={arXiv preprint arXiv:2609.00595},
  year={2026}
}