Skip to main content
Glama
anandapurva55

multi-source-mcp-benchmark

Related Servers

Alternatives to multi-source-mcp-benchmark

No user-submitted related servers found.

    Related Servers

    • A
      license
      B
      quality
      C
      maintenance
      Enables deterministic security testing of AI agents that use tools by serving synthetic MCP environments with poisoned data, fake secrets, and privileged actions. Records agent tool calls and evaluates security invariants (e.g., canary leaks, forbidden access, approval binding) without an LLM judge or real systems.
      8
      MIT
    • A
      license
      Not graded
      quality
      B
      maintenance
      A large-scale benchmark that evaluates AI agents' tool-use competency across 36 real MCP servers using a reproducible Docker sandbox and LLM-as-judge scoring.
      MIT
    • A
      license
      A
      quality
      B
      maintenance
      Enables AI agents to perform software engineering tasks inside an isolated, deterministic sandbox—exploring repositories, reproducing failures, applying patches, running tests, and verifying solutions against hidden suites through MCP tools.
      8
      MIT