Research - Advancing Multi-Agent Intelligence | AG2
Building the foundations of multi-agent AI
Our research spans agent frameworks, self-improving systems, and AI safety. These papers are the foundation of AG2.
50k+ GitHub Stars
#1 Twice on Hugging Face
#1 AgentProphet Arena
Best Paper ICLR 2024 Workshop
Award-winning research in multi-agent AI
AutoGen: Enabling next-gen LLM applications via multi-agent conversation framework
Wu, Bansal, Zhang, et al. · COLM 2024
A framework enabling complex LLM applications through multi-agent conversations. Agents can be customized, composed, and connected to create systems that solve tasks no single model can handle alone.
Absolute Zero: Reinforced self-play reasoning with zero data
Zhao, Chen, Wu, et al. · NeurIPS 2025
A reasoning model that improves through self-play without any human-curated training data. The model generates its own curriculum and learns from self-generated feedback.
Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
Zhang, Yin, Zhang, et al. · ICML 2025
A framework for automatically identifying which agent in a multi-agent system causes task failures and when, enabling targeted debugging and improvement.
A Survey of Self-Evolving Agents
Liu, Wang, Zhang, et al. · arXiv 2025
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
Zeng, Lin, et al. · arXiv 2024
StateFlow: Enhancing LLM Task-Solving through State-Driven Workflows
Wu, Bansal, et al. · COLM 2024
Training Language Model Agents without Modifying Language Models
Chen, et al. · ICML 2024
Divide, Optimize, Merge: Scalable Fine-Grained Generative Optimization for LLM Agents
Liu, Zeng, Zhang, et al. · EMNLP 2025
Memory-Augmented Agent Training for Business Document Understanding
Liu, Zeng, Højmark-Bertelsen, et al. · AAAI 2025
CaptainAgent: Adaptive Team Assembly for Multi-Agent Systems
Zhang, Wu, et al. · arXiv 2024
EcoAssistant: Using LLM Assistant More Affordably and Accurately
Zhang, et al. · arXiv 2023
BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
Ding, Mallick, Zhang, et al. · ICML 2025
SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement
Jain, Wu, Zeng, et al. · EMNLP 2025
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
Wu, Velazco, Zhao, et al. · ICML 2025
Assessing and Verifying Task Utility in LLM-Powered Applications
Arabzadeh, Huo, Mehta, et al. · EMNLP 2024
AutoGenBench
Fourney, Wu · AG2 Blog 2024
16 papers · Last updated January 2026
Collaborate with us
We work with universities, research labs, and industry partners on multi-agent AI. If you have ideas for benchmarks, experiments, or new research directions, we'd like to hear from you.