Skip to main content
Glama
heygabedev

MCP Server Sorter

by heygabedev

MCP Server Sorter

Find, compare, and keep MCP server selections with evidence you can inspect.

A local integration workbench with deterministic search, optional validated reranking, versioned catalogs, reproducible evaluations, and recovery tools. The default showcase runs entirely offline after installation: 30 synthetic server records, repeatable mock providers, no credentials, and no paid calls.

Catalog browser

Run the showcase

Build prerequisites: Python 3.13, Node.js 24, and uv. Dependency installation requires internet access; the running demo does not. Native operation requires no Docker.

git clone https://github.com/heygabedev/mcp-server-sorter.git
cd mcp-server-sorter
uv sync --frozen --python 3.13
npm ci --prefix web
uv run python scripts/build_release.py
uv tool run --python 3.13 --from ./dist/mcp_server_sorter-0.1.0-py3-none-any.whl mcp-sorter serve

Open http://127.0.0.1:8000. The wheel contains the interface; Node is needed only to build it. Data stays in .data beneath the directory where the service starts. Set SORTER_DATA_DIR to use a different location.

An already-built wheel can be launched with the last command alone. The release assets include artifact checksums. Fully offline installations use a previously downloaded wheelhouse; see recovery and artifact management.

Alternatively, download the Linux amd64 container archive from the same release, verify its checksum, and load it locally:

docker load --input mcp-server-sorter-0.1.0-linux-amd64.tar.gz
docker compose -p mcp-sorter up -d

Run Compose from this repository. It binds to 127.0.0.1:8000, stores data in a named volume, and uses a read-only container filesystem. The archive needs no registry credentials. Only Linux amd64 images are verified; native wheel checks cover Windows and Linux. See container recovery for backup and pinned-image rollback.

Related MCP server: @curatedmcp/mcp

Walk through the demo

  1. Discover: search for github pull requests or local database, apply filters, and open a record to inspect its evidence. Deprecated records are excluded by default; unknown metadata stays unknown.

  2. Compare: select up to four records. Compare transport, authentication, and license, then save a named collection. Export and import selections with their pinned versions and evidence.

  3. Evaluate: run the evidence baseline, compare it with a saved run, then try the mock fast or malformed-output profiles. Inspect each case, constraint failure, unsupported explanation, and fallback reason.

  4. Operate: submit a catalog refresh, inspect durable jobs and request traces, and see local alert conditions.

  5. Recover: compare catalog versions, activate an earlier catalog/configuration pair, and create a verified backup. Restoration uses a new data directory and preserves collections added after the backup.

The server names are familiar integration examples. Their capabilities, licenses, versions, and evidence are illustrative fixtures, not verified claims about vendor products. Model presets simulate success, timeouts, rate limits, malformed output, and unavailable providers. The displayed timings measure this local implementation, not provider latency.

Evaluation results

Architecture

flowchart LR
    Web[React interface] --> API[FastAPI /api/v1]
    CLI[Typer CLI] --> Services[Shared application services]
    MCP[Read-only MCP stdio tools] --> Services
    API --> Services
    Services --> Search[SQLite FTS5 / BM25]
    Services --> State[SQLite state / Alembic]
    Services --> Jobs[Durable local jobs]
    Services --> Eval[Versioned evaluation runner]
    Services --> Gateway[Mock or explicitly enabled gateway]
    Search --> Catalog[Immutable catalog snapshots]
    Gateway --> Policy[Allowlist / bounds / output validation]

Python 3.13, FastAPI, Pydantic, SQLAlchemy/Alembic, SQLite FTS5, HTTPX, and Typer serve the same application logic through REST, CLI, and MCP. React/TypeScript/Vite builds into the Python wheel.

Each catalog snapshot has its own immutable SQLite index. Historical snapshots cannot change the active snapshot's BM25 corpus statistics. Filters run before result limits; evidence count, timestamp capped to the snapshot date, and server ID resolve ranking ties. Optional reranking is restricted to the first 20 known candidates and their evidence IDs. Invalid output falls back to deterministic retrieval.

Jobs persist their inputs, configuration and catalog pins, idempotency keys, attempts, and renewable leases. Owner fencing prevents an expired worker from committing a result after another worker takes over. The local SQLite state store uses short transactions, a busy timeout, and rollback journaling. This is a single-machine demo, not a distributed task queue.

Version boundaries are independent:

Artifact

Identity and compatibility

Application

Package version, pinned wheel SHA-256, container image ID

State schema

Alembic revision plus an application compatibility manifest

Catalog

Canonical-content SHA-256 and independent FTS index

Ranking configuration

Immutable configuration SHA-256, policy version and weights

Prompts and model profiles

Explicit versions and prompt/schema hashes when reranking runs

Evaluation dataset

Dataset version, content hash, intent split and review status

Evaluation run

Immutable report ID, source hash and all relevant artifact pins

Catalog/configuration activation changes both pointers in one transaction. Recovery instructions cover data restoration and application rollback.

Evaluation results

The bundled golden mock dataset contains 90 scenarios: 30 development and 60 held-out, separated by intent, plus six adversarial query fixtures. Judgments include relevance grades, supporting evidence, expected abstention, rationale, and fixture-authored review status.

Fixture profile

NDCG@5

Recall@10

Reciprocal rank

Fallback rate

Evidence baseline

1.000

1.000

1.000

0%

Mock balanced

1.000

1.000

1.000

0%

Mock fast

0.969

1.000

0.958

0%

Mock timeout / malformed / rate-limited / unavailable

1.000

1.000

1.000

66.7%

These are synthetic regression results, not live model quality measurements. Thirty cases require abstention and never call a reranker. The mock fast profile deliberately swaps candidates; its NDCG decrease of 0.0308 fails the 0.02 regression threshold. Failure presets keep the baseline results through fallback. All profiles produce zero constraint violations, invalid evidence references, and unsupported generated explanation templates.

Reports contain JSON and escaped HTML, per-case failures, latency, abstention, schema failures, fallback rate, available usage metadata, and paired comparisons with a deterministic bootstrap over intents. The explanation check verifies generated templates against pinned catalog facts; it does not independently verify vendor claims. See raw results and test evidence.

uv run mcp-sorter eval run
uv run mcp-sorter eval run --profile demo-malformed
uv run mcp-sorter eval compare BASELINE_ID CANDIDATE_ID
uv run python scripts/record_fixture_results.py

Tests and performance

The v0.1.1 regression gate requires correct abstention on every case and checks NDCG changes separately for development and held-out cases, as well as overall. Comparison output identifies the failed scope and rule; simulated provider failures remain visible without failing a correct fallback ranking.

Evaluation JSON is immutable. If a worker exits after saving JSON, its next attempt regenerates the HTML export from that saved result without rerunning the evaluation.

Unit, property/metamorphic, SQLite integration, API schema, external contract, security, reliability, UI, accessibility, packaging, and recovery checks are implemented. Network access is denied in ordinary Python tests except loopback and local sockets. No external credentials are needed.

uv run ruff check .
uv run ruff format --check .
uv run mypy
uv run pytest --cov=mcp_sorter --cov-report=json
uv run python scripts/check_coverage.py
uv run python scripts/export_openapi.py --check
npm --prefix web run lint
npm --prefix web run format:check
npm --prefix web run build
npm --prefix web test

For browser tests, run npx playwright install chromium from web, then npm run test:e2e. Playwright starts the local API and web dev server. The journeys include axe checks, keyboard focus, mobile layout, comparisons, collections, evaluations, mock failures, version activation, and backups.

Run uv run python scripts/mutation_check.py and uv run python scripts/benchmark.py for the bounded mutation and load suites. The Linux reference run used 10,000 synthetic records: 74.9 ms p95 at 10 RPS, 97.9 ms at 50 RPS, and 70.0 ms during a three-minute 10 RPS soak, with zero errors across 2,600 requests. This is a short synthetic workload; it does not establish long-term or production capacity.

Local verification reached 94.74% overall branch coverage, with at least 95% in ranking, evaluation, and recovery. All 13 targeted mutations were detected. Full scope, machine details, raw reports, and platform limits are in verification.

GitHub Actions workflows remain in the repository but are currently disabled at the owner's request. Local checks are the recorded release evidence; historical failed hosted runs are not represented as passes.

Operations and integrations

/health/live, /health/ready, /metrics, and the Operations screen expose health, job state, controlled fallback events, and local telemetry. Structured logs contain request IDs, trace IDs, and statuses, without query strings, credentials, or raw provider content. OpenTelemetry traces retain the latest 100 spans in memory; SQLite retains 1,000 controlled events. Prometheus configuration, alert rules, and a Grafana dashboard are in monitoring/. External monitoring deployment is optional and has not been live-verified.

The default listener binds to loopback. There is no multi-user authentication or authorization layer; the application is intended for a trusted local machine. Keep the listener local.

The MCP server exposes only search_servers and compare_servers over stdio:

uv run mcp-sorter mcp

Registry, GitHub metadata, OpenRouter, LiteLLM Proxy, and approved remote-probe adapters are contract-tested with controlled responses. Credentials alone never enable networking. Integration setup describes the explicit opt-in controls and live-verification limits. .env.example contains placeholders and is not loaded automatically.

MIT license. See LICENSE.

Available Tools

2 tools
compare_serversC
Read-onlyIdempotent

Compare two to four distinct local server records in a pinned catalog.

ParametersJSON Schema
NameRequiredDescriptionDefault
idsYes
snapshotNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds the constraint that records must be 'distinct' and that they are in a 'pinned catalog', which is useful context beyond the annotations. No behavioral contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no fluff. It communicates the core purpose efficiently, though it is terse and lacks necessary parameter details. It earns its place but could be expanded without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has two parameters (one required) and an output schema, but the description provides no guidance on how to populate the required 'ids' array or the optional 'snapshot'. An agent cannot reliably call this tool correctly without additional knowledge, despite the annotations covering safety. The description is incomplete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the 'ids' and 'snapshot' parameters. It does not explain what 'ids' should contain (server identifiers), the expected array length, or what 'snapshot' refers to. The phrase 'two to four distinct local server records' only hints at the ids count but does not clarify the parameter itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('compare') and a clear resource ('local server records in a pinned catalog'), and it bounds the operation to two to four distinct records. While it does not explicitly name the sibling tool, the verb 'compare' clearly distinguishes it from 'search_servers', so purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The usage context is implied by the verb and object — this tool is for comparing existing server records, not searching. However, there is no explicit guidance on when to use this versus the sibling 'search_servers', nor any exclusions (e.g., 'use search_servers to find, then compare here').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_serversB
Read-onlyIdempotent

Search local catalog records and return evidence-linked deterministic rankings.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNo
filtersNo
snapshotNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeNo
as_ofYes
queryYes
filtersYes
resultsYes
snapshotYes
configurationNo
model_metadataNo
policy_versionNo
prompt_versionNo
schema_versionNo
sqlite_versionNo
fallback_reasonNo
application_versionNo
model_profiles_versionNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive behavior, so the description only needs to add context beyond that. It adds 'evidence-linked deterministic rankings,' which communicates ranking stability and evidence sourcing, but it does not disclose ordering, pagination, filtering behavior, or snapshot semantics. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly packed sentence with no filler. It front-loads the action and output, and every word adds meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having strong annotations and an output schema, the description omits when to use the tool, parameter semantics, and how it differs from compare_servers. For a search tool with three optional parameters, more context is needed for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carried the burden of explaining the three parameters (query, filters, snapshot) but does not mention any of them. The schema provides names, defaults, and enums, but the meaning of 'snapshot' and how filters interact with rankings remains unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Search'), a clear resource ('local catalog records'), and a concrete output ('evidence-linked deterministic rankings'). It is easily distinguished from the sibling 'compare_servers' because searching and ranking are different operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus compare_servers or any other alternative. It implies a local-catalog scope but does not state exclusions, prerequisites, or when a sibling would be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.1
    • First observedcompare_servers
    • First observedsearch_servers

TDQS

B3.4/5.0

Scored across 2 tools

Disambiguation5/5

search_servers and compare_servers have distinct roles: one finds and ranks matching catalog records, the other directly compares a small fixed set of records. There is no real overlap because compare explicitly requires two to four pinned records rather than returning ranked search results.

Naming Consistency5/5

Both tools follow the same verb_noun pattern with plural object nouns: search_servers and compare_servers. This is perfectly consistent and predictable.

Tool Count3/5

With only two tools the server feels minimal for a catalog sorter, though the narrow scope of search-and-compare keeps it from being unreasonable. It is borderline rather than clearly well-scoped.

Completeness4/5

The core workflows of discovering ranked servers and comparing selected candidates are covered. A direct way to list the full catalog or fetch a single record would be a minor gap, but search can likely cover discovery.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    Not graded
    maintenance
    Stores metadata for MCP servers and provides smart search capabilities, allowing users to find appropriate MCP servers for their queries and route requests to the most suitable server.
    12
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables discovering and managing MCP servers through a registry, supporting listing, searching, and configuration.
    -