MCP-AtlasAniGG-EthAlicense-Not gradedqualityBmaintenanceA large-scale benchmark that evaluates AI agents' tool-use competency across 36 real MCP servers using a reproducible Docker sandbox and LLM-as-judge scoring. Updated 5 days ago (2026-08-18 09:55 UTC)MIT