Research - Advancing Multi-Agent Intelligence | AG2

Building the foundations of multi-agent AI

Our research spans agent frameworks, self-improving systems, and AI safety. These papers are the foundation of AG2.

50k+ GitHub Stars

#1 Twice on Hugging Face

#1 AgentProphet Arena

Best Paper ICLR 2024 Workshop

Award-winning research in multi-agent AI

AutoGen: Enabling next-gen LLM applications via multi-agent conversation framework

Wu, Bansal, Zhang, et al. · COLM 2024
A framework enabling complex LLM applications through multi-agent conversations. Agents can be customized, composed, and connected to create systems that solve tasks no single model can handle alone.

Absolute Zero: Reinforced self-play reasoning with zero data

Zhao, Chen, Wu, et al. · NeurIPS 2025
A reasoning model that improves through self-play without any human-curated training data. The model generates its own curriculum and learns from self-generated feedback.

Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems

Zhang, Yin, Zhang, et al. · ICML 2025
A framework for automatically identifying which agent in a multi-agent system causes task failures and when, enabling targeted debugging and improvement.

A Survey of Self-Evolving Agents

Liu, Wang, Zhang, et al. · arXiv 2025

AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks

Zeng, Lin, et al. · arXiv 2024

StateFlow: Enhancing LLM Task-Solving through State-Driven Workflows

Wu, Bansal, et al. · COLM 2024

Training Language Model Agents without Modifying Language Models

Chen, et al. · ICML 2024

Divide, Optimize, Merge: Scalable Fine-Grained Generative Optimization for LLM Agents

Liu, Zeng, Zhang, et al. · EMNLP 2025

Memory-Augmented Agent Training for Business Document Understanding

Liu, Zeng, Højmark-Bertelsen, et al. · AAAI 2025

CaptainAgent: Adaptive Team Assembly for Multi-Agent Systems

Zhang, Wu, et al. · arXiv 2024

EcoAssistant: Using LLM Assistant More Affordably and Accurately

Zhang, et al. · arXiv 2023

BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute

Ding, Mallick, Zhang, et al. · ICML 2025

SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement

Jain, Wu, Zeng, et al. · EMNLP 2025

ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation

Wu, Velazco, Zhao, et al. · ICML 2025

Assessing and Verifying Task Utility in LLM-Powered Applications

Arabzadeh, Huo, Mehta, et al. · EMNLP 2024

AutoGenBench

Fourney, Wu · AG2 Blog 2024

16 papers · Last updated January 2026

Collaborate with us

We work with universities, research labs, and industry partners on multi-agent AI. If you have ideas for benchmarks, experiments, or new research directions, we'd like to hear from you.

Get in touch →