GitHub - AIDC-AI/complex-mcp: [ICML 2026] ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox

ComplexMCP is a benchmark for evaluating model performance in complex software workflows and large API tool ecosystems.

1) Build Environment Via Docker

docker build -t complexmcp:latest .

docker run -d --name complexmcp \
  -p 8000-8007:8000-8007 \
  -p 9000-9006:9000-9006 \
  complexmcp:latest

2) Create `.env`

Create a .env file in the project root, following .env.example format.

cp .env.example .env

Then fill values in .env as needed.

3) Run Benchmark

python run_benchmark.py --tool-config config/general.yaml \
  --model [model_name]

If you find this work helpful, please cite our paper:

@misc{li2026complexmcpevaluationllmagents,
      title={ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox}, 
      author={Yuanyang Li and Xue Yang and Longyue Wang and Weihua Luo and Hongyang Chen},
      year={2026},
      eprint={2605.10787},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2605.10787}, 
}

Name		Name	Last commit message	Last commit date
Latest commit History 68 Commits
assets		assets
benchmark		benchmark
client		client
config		config
docker		docker
servers		servers
software		software
.dockerignore		.dockerignore
.env.example		.env.example
.gitignore		.gitignore
Dockerfile		Dockerfile
LICENSE		LICENSE
README.md		README.md
requirements.txt		requirements.txt
run_benchmark.py		run_benchmark.py
start_servers.sh		start_servers.sh
start_softwares.sh		start_softwares.sh

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox

1) Build Environment Via Docker

2) Create `.env`

3) Run Benchmark

About

Uh oh!

Releases

Packages

Uh oh!

Contributors

Uh oh!

Languages

Folders and files

Latest commit

History

Repository files navigation

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox

1) Build Environment Via Docker

2) Create .env

3) Run Benchmark

About

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Uh oh!

Contributors

Uh oh!

Languages

2) Create `.env`

Packages