Skip to main content
Glama

QMCP - Model Context Protocol Server

An AI assistant can only use a tool somebody has exposed to it. qmcp is the server that does the exposing: it publishes your tools, runs them when an assistant asks, records every invocation, and stops for a human when a step needs one.

It speaks the Model Context Protocol, so any client that speaks it can use these tools without being told about them in advance. Built with FastAPI.

Features

  • Tool Discovery - List available tools via /v1/tools

  • Tool Invocation - Execute tools via /v1/tools/{name}

  • Invocation History - Audit trail via /v1/invocations

  • Human-in-the-Loop - Request human input via /v1/human/*

  • Persistence - SQLite with SQLModel/aiosqlite

  • Python Client - qmcp.client.MCPClient for workflows

  • Metaflow Examples - Ready-to-use flow templates

  • Agent Framework - SQLModel schemas + mixins for agent types/topologies

  • PydanticAI Integration - Create agents from QMCP models with full audit trail

  • Structured Logging - JSON logs with structlog

  • Request Tracing - Correlation IDs across requests

  • Metrics - Prometheus-compatible /metrics endpoint

  • CLI Interface - Manage via qmcp command

Related MCP server: AI Agent MCP Server

Quick Start

# Install dependencies
uv sync

# Start the server
uv run qmcp serve

# Or with development reload
uv run qmcp serve --reload

See quickstart.md for a copy-paste walkthrough.

Three commands worth knowing before the rest:

uv run qmcp topology show governed --level 2   # the seam a model is called through
uv run qmcp orchestration plane                # what every shape would do, declared
uv run qmcp human list                         # what is waiting on a person

The first is the one to read. Model output reaches the human queue through one door, and docs/human_in_loop.md says what that door does and does not enforce.

Adoption and Onboarding

Adoption checklist:

  • Decide how the server is hosted (local, container, or VM) and who can reach it.

  • Set QMCP_HOST, QMCP_PORT, and QMCP_DATABASE_URL for your environment.

  • Standardize X-Correlation-ID values for audit trails across clients.

  • Decide how humans submit HITL responses (UI or API).

  • Wire /metrics into your monitoring stack.

Onboarding path:

  1. uv sync --all-extras

  2. Run the end-to-end tutorial below.

  3. uv run qmcp serve for local exploration.

End-to-End Tutorial (HITL approval workflow)

This tutorial mirrors the end-to-end test tests/test_hitl.py::TestHITLWorkflow::test_complete_approval_workflow.

Copy and paste:

uv sync --all-extras
uv run pytest tests/test_hitl.py::TestHITLWorkflow::test_complete_approval_workflow -v

Client Library

from qmcp.client import MCPClient

with MCPClient(base_url="http://localhost:3141") as client:
    # List tools
    tools = client.list_tools()

    # Invoke a tool
    result = client.invoke_tool("echo", {"message": "Hello!"})
    print(result.result)

    # Human-in-the-loop
    request = client.create_human_request(
        request_id="approval-001",
        request_type="approval",
        prompt="Approve deployment?",
        options=["approve", "reject"]
    )
    response = client.wait_for_response("approval-001", timeout=3600)

See docs/client.md for full API documentation.

Validating itself, and the human loop

qmcp can be pointed at its own repository. Each check is a real subprocess, recorded as the same invocation row the server writes when a tool is called over HTTP — so qmcp dashboard reads a self-check back like anything else.

# run this repository's own gates, recording each one
qmcp selfcheck --database run.db

# a failing check raises a question. What is waiting on a person:
qmcp human list --database run.db

# answer it. This is a person acting, and it is recorded as one:
qmcp human respond selfcheck-tag-claims defer --database run.db --by "your name"

A failing check becomes a unit of work; a passing one becomes nothing, because a green gate is not work. It opens at brainstorm — noticing is not deciding — and moves to planning once somebody answers. It goes no further: re-running a check cannot establish that anybody acted on it.

walkthrough/02-a-run-that-found-something.md executes the whole thing, with real subprocesses.

The pair. qmcp is the harness; dossier is the control panel. Neither imports the other — what crosses is a schema, and an address names the same row on both sides:

qmcp selfcheck --database run.db --deltas > deltas.json   # units of work
qmcp dashboard --database run.db --json   > harness.json  # what has run
# then, in dossier: dossier deltas ingest / dossier harness ingest

CLI Commands

# Start the server
qmcp serve [--host HOST] [--port PORT] [--reload]

# Start the server for Docker-based flows
qmcp cookbook serve [--host 0.0.0.0] [--port PORT] [--reload]

# List registered tools
qmcp tools list

# Show configuration
qmcp info

# Run a cookbook flow in Docker
qmcp cookbook simple-plan --goal "Deploy a web service"

# Start the server + run a cookbook flow (unified dev)
qmcp cookbook dev simple-plan --goal "Deploy a web service"

# Run a cookbook flow via the generic runner
qmcp cookbook run simple-plan --goal "Deploy a web service"

# Run other cookbook recipes (flow args are passed through)
qmcp cookbook run approved-deploy --service "api-gateway" --environment "staging"
qmcp cookbook dev local-qc-gauntlet --change-summary "Add audit fields" --target-area "metrics, logging"

# Run a cookbook flow in Docker explicitly
qmcp cookbook docker simple-plan --goal "Deploy a web service"

# If the qmcp shim cannot be installed (Windows)
uv run --no-sync python -m qmcp cookbook run simple-plan --goal "Deploy a web service"

# Run tests with auto setup/teardown
qmcp test [-v] [--coverage] [TEST_PATH]

Cookbook flows run in Docker and require Docker Desktop (Linux engine). Add --no-sync to skip syncing flow dependencies if the image is already built.

API Endpoints

Endpoint

Method

Description

/health

GET

Health check

/v1/tools

GET

List available tools

/v1/tools/{name}

POST

Invoke a tool

/v1/invocations

GET

List invocation history

/v1/invocations/{id}

GET

Get single invocation

/v1/human/requests

POST

Create human request

/v1/human/requests

GET

List human requests

/v1/human/requests/{id}

GET

Get request with response

/v1/human/responses

POST

Submit human response

/v1/topology

GET

Every topology shape, and the encoding a window must honour

/v1/topology/encoding

GET

Which visual channel carries which data axis

/v1/topology/shape/{kind}

GET

One shape as boxes and arrows, at a level

/v1/topology/schema/{kind}

GET

The JSON schema of one kind's configuration class

/v1/orchestration/plane

GET

What every shape would do, declared: status, spends/writes/decides, needs, refusals, drift

/v1/orchestration/runnable

GET

What a hand (workers, budget, model, built) could run, and what each shape is short of

/v1/topologies

POST

Save a topology design, validated through its kind's configuration class

/v1/topologies

GET

Every saved design, addressed, with the plane's verdict

/v1/topologies/{ref}

GET

One design by id or by name; ?act= judges a pairing

/v1/topologies/{ref}

PUT

Change a design's description, config or version

/metrics

GET

Prometheus metrics

/metrics/json

GET

Metrics as JSON

The topology, orchestration and design routes name nobody and are served wherever the server is bound. walkthrough/07-saving-a-shape-is-not-running-it.md exercises them, and the reason a refused shape can be saved is in qmcp/topology_designs.py. The archive-derived readings under /v1/topology/relations/ and /v1/threads are loopback-only and are not in this table.

Built-in Tools

  • echo - Echo input back (for testing)

  • planner - Create execution plans

  • executor - Execute approved plans

  • reviewer - Review and assess results

Development

# Install dev dependencies
uv sync --all-extras

# Run tests (with auto cleanup)
uv run qmcp test -v

# Run tests with coverage
uv run qmcp test --coverage

# Run linter
uv run ruff check .

Architecture

See docs/architecture.md for the full architectural overview.

The system follows a three-plane architecture:

  1. Client/Orchestration - Metaflow workflows (MCP client)

  2. MCP Server - FastAPI service (this project)

  3. Execution/Storage - Tools and database

Documentation

Example Flows

See examples/flows/ for Metaflow integration examples:

  • simple_plan.py - Basic tool invocation

  • approved_deploy.py - HITL approval workflow

  • local_agent_chain.py - Local LLM plan -> review -> refine with SQLModel artifacts

  • local_qc_gauntlet.py - Local LLM QC checklist/task/gate builder

  • local_release_notes.py - Local LLM release notes and doc update suggestions

For local LLM flows, install extras with uv sync --extra flows. Start uv run qmcp serve --host 0.0.0.0 when --use-mcp True to enable MCP calls from Docker-based flows. On Windows, prefer running flows in a Linux container to avoid platform-specific Metaflow dependencies.

Docker runner (recommended on Windows):

docker compose -f docker-compose.flows.yml build
docker compose -f docker-compose.flows.yml run --rm flow-runner \
  examples/flows/local_agent_chain.py run --use-mcp True --goal "..."

Set MCP_URL and LLM_BASE_URL (or pass --mcp-url / --llm-base-url) when running in Docker, e.g. http://host.docker.internal:3141.

License

MIT

Related MCP Connectors

Related MCP Servers