Skip to main content
Glama
alphaparkinc

genpark-multi-model-prompt-regression-benchmark-evaluator-skill

Related Servers

Alternatives to genpark-multi-model-prompt-regression-benchmark-evaluator-skill

No user-submitted related servers found.

    Related Servers

    • A
      license
      Not graded
      quality
      B
      maintenance
      Provides deployable, stateless MCP services for rigorous inference evaluation and benchmarking, including a profiled lm-evaluation-harness controller/worker and a bounded vLLM forward-pass benchmark adapter.
      Apache 2.0
    • A
      license
      Not graded
      quality
      C
      maintenance
      Enables benchmarking of local LLM models (performance and quality) and sharing results to a public leaderboard via MCP tools.
      15 npm
      8
      Apache 2.0
    • A
      license
      Not graded
      quality
      A
      maintenance
      This MCP server provides a stateful, resettable, verifiable API runtime that gates every tool call, enabling agents to run long workflows against provider-shaped environments without live provider write access. It records decisions, side effects, and outcome evidence for replayable, verifiable benchmark runs.
      Apache 2.0