Skip to main content
Glama
88,810 servers. Updated

Matching MCP tools:

Matching MCP Connectors:

"OpenAI Gym" matching MCP servers:

GET /v1/servers – MCP directory API reference
  • A
    license
    A
    quality
    A
    maintenance
    MCP server that lets coding agents test AI agents. Create YAML test cases, snapshot golden baselines, check for regressions, and generate visual reports all from inside Claude Code or any MCP-compatible tool. Works with LangGraph, CrewAI, OpenAI, Claude, Mistral, and any HTTP API.
    10
    58 npm
    584 PyPI
    134
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    An MCP server that allows agents to test and compare LLM prompts across OpenAI and Anthropic models, supporting single tests, side-by-side comparisons, and multi-turn conversations.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to fetch daily deterministic constraint challenges and submit JSON answers for immediate, stateless evaluation via MCP.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to drive SECS/GEM host testing by connecting to equipment, sending SECS messages, running E30 startup conformance reports, and monitoring events and status.
    484 npm
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables ChatGPT to control Mobile Safari on a real USB-connected iPhone through an OpenAI Secure MCP Tunnel and Appium MCP, including device selection, WDA preparation, and Safari session automation.
    1
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    An MCP server that performs independent verification of artifacts against criteria using a different AI model lineage (codex, OpenAI, or Gemini) to catch defects that same-family checks might miss.
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Panel Review is the guardian agent for AI's highest-stakes coding decisions. Before an agent's riskiest designs, diffs, or commits ship, four frontier models, from OpenAI, Anthropic, Google, and xAI, argue them through and return severity-tagged findings, with review gates the agent cannot silently skip.
    117 npm
    2
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables comparison of responses from multiple LLMs (OpenAI, Anthropic, Gemini) to the same prompt, returning a validated divergence score based on sentence embeddings.
    26 npm
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Analyzes app screenshots to identify UI/UX issues, compare designs with implementations, and provide actionable fixes using GPT-4o/GPT-5.2 vision capabilities. Supports single/batch analysis, design comparison, and automated report generation for iOS, Android, web, and desktop platforms.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Autonomously analyzes projects, generates Playwright tests via OpenAI, runs them, classifies failures, and recommends fixes, iterating up to 3 times.
    11 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables running the same prompt(s) across roughly 300 models from Anthropic, OpenAI, Google, Meta, Mistral, DeepSeek, Qwen and others, then comparing per-model outputs, latency, token usage, error rates and billed cost side by side. Supports free cost estimates before launching, asynchronous run creation with status polling and cancellation, prepaid x402 credit top-ups, and an optional AI-written comparison summary of the results.
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Provides browser automation, AI-powered analysis, visual processing, web scraping, automated test generation, and DevTools analysis capabilities. Supports multiple AI providers (OpenAI, Anthropic, Google, Ollama) for intelligent web interaction and data extraction.
    -
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI-powered document analysis and querying for project documentation using vector embeddings stored in Redis. Supports document upload, context-aware Q\&A, automatic test case generation, and requirements traceability through OpenAI integration.
    205 npm
    -