Skip to main content
Glama

๐Ÿ”Œ MCP Agent Toolkit

A production-shaped Model Context Protocol server + autonomous agent client โ€” an LLM discovers four guarded tools over the open standard at runtime, recovers from live tool errors, and produces a cited report, with every invocation logged, rate-limited, and correlated by request ID.

Node 20+ TypeScript MCP CI Coverage Lint Type checked: tsc License: MIT

Recorded demo โ€” agent โ†” MCP server over stdio

Real run, live OpenAI tool-calling loop โ€” watch it discover the tools, hit a real SQL error (no such column: category), retry with the correct schema, search the live web, and write a cited report into the sandbox. Full transcript: docs/demo-transcript.txt โ€” recorded via npm run demo:record (asciinema cast in docs/demo.cast).


๐Ÿ“Œ What this is

A standalone MCP server + agent client in strict TypeScript. The server exposes four tools over the official @modelcontextprotocol/sdk with stdio transport: web_search, db_query (read-only SQLite), read_file and write_file (sandboxed). The client spawns the server as a child process and speaks plain MCP โ€” tools/list, tools/call, params._meta correlation โ€” exactly like Claude Desktop or any other MCP-capable host would. Nothing is imported in-process; the protocol is the boundary.

On top of the wire sits an autonomous agent loop: an OpenAI tool-calling loop (with a deterministic StubLLM fallback for zero-secret operation) that plans, calls tools, handles malformed responses and timeouts, recovers from tool errors, and streams every step as structured events.

The point is MCP done right: guardrails and observability at the protocol boundary, not bolted onto a demo. Every tool call is schema-validated, rate-limited, sanitized on the way back (the web is untrusted โ€” OWASP LLM01), logged as JSON with a correlated request ID, and counted in Prometheus metrics.


Related MCP server: MCP Server

๐Ÿ’ก Why this matters โ€” for your project, or for a technical reviewer

  • ๐Ÿ”Œ The protocol is the real boundary. The agent discovers tools via tools/list and invokes them via tools/call over JSON-RPC stdio โ€” no shared imports, no framework glue. Swap the client for any MCP host and the server just works.

  • ๐Ÿ›ก๏ธ Defense in depth, not a single regex. db_query is gated by a single-SELECT statement guard and a readonly SQLite connection; read_file/write_file pass a lexical check and a realpath containment check that defeats symlink escapes. Each layer fails closed on its own.

  • ๐Ÿงผ Untrusted output never becomes instructions. Every tool result is sanitized server-side (control/bidi/zero-width chars stripped, injection patterns neutralized, MAX_TOOL_OUTPUT_CHARS cap) before it reaches the model context โ€” the OWASP LLM01 control for indirect prompt injection.

  • ๐Ÿ”‘ The server never sees the LLM key. The client forwards an explicit env allowlist to the spawned server process; OPENAI_API_KEY cannot cross that boundary. Empty secrets count as absent, so CI runs with zero credentials.

  • ๐Ÿงพ Request-ID correlation end-to-end. The client stamps a UUID into params._meta.requestId per call; the server logs it on every invocation line โ€” client event stream joins server logs trivially.

  • ๐Ÿšฆ Every failure is a first-class signal. Timeouts, malformed responses, unknown tools and guardrail rejections normalize to typed error envelopes โ€” the agent sees a clean failure reason, not a stack trace (the demo transcript shows a live SQL error recovered).

  • ๐Ÿงช The transport is real in tests too. The integration suite spawns the actual server over stdio and asserts request-ID correlation and rate-limit enforcement across the wire โ€” not a mock of MCP.

  • ๐Ÿ“‰ CI green with zero secrets. Fork, clone, npm ci && npm test โ€” 110 tests pass offline; the deterministic StubLLM and SEARCH_PROVIDER=none need no keys.


๐Ÿ—๏ธ Architecture

graph LR
    subgraph client["Agent client"]
        CLI[CLI entry] --> LOOP[Agent loop<br/>OpenAI tool calling]
        LOOP --> MC[MCP client<br/>StdioClientTransport]
        LOOP -.->|no OPENAI_API_KEY| STUB[StubLLM<br/>deterministic]
        LOOP -->|chat.completions| OAI[(OpenAI API)]
    end

    MC ==>|"JSON-RPC over stdio<br/>tools/list ยท tools/call<br/>requestId via params._meta"| MS

    subgraph server["MCP server (child process)"]
        MS[McpServer<br/>StdioServerTransport] --> REG[Tool registry<br/>+ guardrail wrapper]
        REG --> RL[Rate limiter<br/>token bucket]
        RL --> T1[web_search]
        RL --> T2[db_query]
        RL --> T3[read_file]
        RL --> T4[write_file]
        T1 --> SANE[Output sanitizer]
        T2 --> SQLG[SELECT-only guard<br/>+ readonly conn]
        T3 --> FS[Path sandbox]
        T4 --> FS
        REG --> LOG[pino JSON โ†’ stderr]
        REG --> MET[Prometheus metrics<br/>optional :METRICS_PORT]
    end

    T1 -->|HTTPS| EXT1[(Tavily /<br/>DuckDuckGo)]
    T2 --> EXT2[(SQLite<br/>readonly)]
    T3 --> EXT3[(data/sandbox/)]
    T4 --> EXT3

Full rationale and the tool/guardrail matrix: docs/architecture.md.


๐Ÿ“Š Evaluation-Driven Development (EDD)

The quality gates are binary and enforced in CI โ€” a regression fails the build rather than being shrugged off:

Dimension

Metric

Threshold

Latest run

Test suite

unit + security + real-stdio integration

100% green

110/110 โœ…

Coverage

lines / branches / functions / statements

โ‰ฅ 85% each

98.5 / 85.1 / 100 / 97.5 โœ…

Type safety

tsc --strict

0 errors

โœ…

Lint

eslint strict (typescript-eslint)

0 warnings

โœ…

Security suite

traversal ยท write SQL ยท injection ยท rate-limit rejections

all must fail loudly

โœ…

Dependency audit

npm audit --omit=dev

0 high/critical

โœ…

Reproduce: npm run test:cov (coverage gate), npm run verify-equivalent sequence in CONTRIBUTING.md.


๐Ÿ”Œ API

The server speaks standard MCP (JSON-RPC 2.0 over stdio): initialize, tools/list, tools/call. Tools:

Tool

Input

Guardrails

Returns

web_search

query (1โ€“400 chars), max_results (default 5)

provider allowlist, timeout, output sanitizer + size cap

{ provider, results: [{title, url, snippet}] }

db_query

query โ€” a single SELECT/WITH

statement guard + readonly connection, 500-row cap

{ rows, row_count, truncated }

read_file

path relative to sandbox

lexical + realpath containment, READ_FILE_MAX_BYTES cap

{ path, content }

write_file

path, content

same containment, WRITE_FILE_MAX_BYTES cap

{ path, bytes_written }

Example tools/call payload (real shape โ€” the client adds _meta.requestId automatically):

{
  "jsonrpc": "2.0", "id": 7, "method": "tools/call",
  "params": {
    "name": "db_query",
    "arguments": { "query": "SELECT title, author FROM reports ORDER BY published_at DESC LIMIT 5" },
    "_meta": { "requestId": "de30734e-8f77-4418-8a27-65557d8ade40" }
  }
}

Agent CLI:

npm run agent -- "question"          # live OpenAI loop (needs OPENAI_API_KEY)
npm run agent -- "question" --stub   # deterministic, zero secrets
npm run agent -- "question" --json   # machine-readable event stream
npm run demo                         # seeded E2E run โ†’ docs/demo-transcript.txt
npm run demo:record                  # record docs/demo.cast โ†’ GIF via asciinema/agg

๐Ÿ“ก Observability & cost

  • pino JSON logs on stderr (stdout is reserved for the MCP transport): every tool invocation logs tool name, sanitized input args, output preview, duration, ok/error, session_id and request_id.

  • Request-ID correlation: the client stamps params._meta.requestId; the server echoes it on every log line โ€” join client events to server logs by UUID.

  • Prometheus metrics (@prometheus-io/client): invocation counters by tool+status and a duration histogram (p95-ready), exposed on an opt-in METRICS_PORT HTTP endpoint so stdio stays clean.

  • Cost: the only metered dependency is the LLM. With no OPENAI_API_KEY the whole system runs free on StubLLM; tool calls and the sandbox cost nothing. MAX_TOOL_ITERATIONS bounds worst-case token spend per run.

Real server log line from the recorded demo:

{"level":30,"service":"mcp-server","event":"tool_invocation","tool":"db_query",
 "request_id":"de30734e-8f77-4418-8a27-65557d8ade40","session_id":"3bb996d6-โ€ฆ",
 "duration_ms":0,"ok":true,"input":"{\"query\":\"SELECT title, author FROM reports\"}",
 "output":"[{\"type\":\"text\",\"text\":\"{\\\"status\\\": \\\"ok\\\"โ€ฆ"}

๐Ÿ”’ Security posture

Sandboxed filesystem (lexical + realpath containment), read-only SQL (statement guard and readonly connection), output sanitization against indirect prompt injection (OWASP LLM01), per-tool/session rate limiting, env allowlist so the server never sees OPENAI_API_KEY, secret-redacted logs, gitleaks + npm audit in CI on pull_request with contents: read. Full threat model and mechanism table: SECURITY.md.


๐Ÿš€ Quick Start

npm ci
cp .env.example .env        # optional โ€” paste OPENAI_API_KEY for the live loop
npm run seed:db             # fixture DB for db_query (8 reports)
npm run demo                # recorded E2E: seeds + runs, saves transcript
npm run agent -- "your question"

Degrades cleanly with zero secrets: no OPENAI_API_KEY โ†’ deterministic StubLLM; no SEARCH_API_KEY โ†’ keyless DuckDuckGo; SEARCH_PROVIDER=none โ†’ fully offline. The entire CI suite runs with no credentials.

docker compose up                    # runs the default demo question once
docker compose run --rm agent "โ€ฆ"    # custom question
docker compose run --rm agent --stub "offline run, no keys"

The image is multi-stage (non-root node user), seeds the fixture DB at build time, and mounts only data/sandbox/ as a volume โ€” the single writable surface. .env is read at runtime via env_file; nothing is baked in.

โš™๏ธ Environment Variables

Variable

Required

Secret?

Purpose

OPENAI_API_KEY

no

โœ…

Live LLM for the agent loop; StubLLM when absent

CHAT_MODEL_NAME

no

โ€”

Model override (default gpt-4o-mini)

LLM_PROVIDER

no

โ€”

auto | openai | stub

SEARCH_PROVIDER

no

โ€”

auto | tavily | duckduckgo | none

SEARCH_API_KEY

no

โœ…

Tavily key (optional paid provider)

MCP_SERVER_COMMAND / MCP_SERVER_ARGS

no

โ€”

Override how the client spawns the server

SANDBOX_ROOT

no

โ€”

The only filesystem root tools may touch

DB_PATH

no

โ€”

SQLite fixture path (npm run seed:db)

RATE_LIMIT_PER_MINUTE

no

โ€”

Per tool, per session (default 60)

MAX_TOOL_ITERATIONS

no

โ€”

Agent loop bound โ€” must terminate

MAX_TOOL_OUTPUT_CHARS

no

โ€”

Tool output truncation cap (untrusted data)

READ_FILE_MAX_BYTES / WRITE_FILE_MAX_BYTES

no

โ€”

FS size caps

LOG_LEVEL

no

โ€”

pino level; logs go to stderr

METRICS_PORT

no

โ€”

Empty = disabled; e.g. 9108 โ†’ /metrics HTTP

OUTPUT_FORMAT

no

โ€”

pretty | json agent event stream

Never commit secrets. See .env.example for the full, commented template.

๐Ÿ“ Project Structure

src/
โ”œโ”€โ”€ server/            # MCP server (stdio) โ€” separate process
โ”‚   โ”œโ”€โ”€ server.ts      #   tool registration + guardrail wrapper + logging
โ”‚   โ”œโ”€โ”€ tools/         #   web_search, db_query, read_file, write_file
โ”‚   โ”œโ”€โ”€ guardrails/    #   path sandbox, SQL guard, rate limiter, sanitizer
โ”‚   โ””โ”€โ”€ observability/ #   pino logger, Prometheus metrics, request ctx
โ”œโ”€โ”€ client/            # agent CLI โ€” spawns the server, speaks MCP
โ”‚   โ”œโ”€โ”€ mcp-client.ts  #   tools/list + tools/call + env allowlist
โ”‚   โ”œโ”€โ”€ agent-loop.ts  #   tool-calling loop (bounded, must terminate)
โ”‚   โ”œโ”€โ”€ events.ts      #   structured event stream (pretty|json)
โ”‚   โ””โ”€โ”€ llm/           #   OpenAI impl + deterministic StubLLM
โ””โ”€โ”€ shared/            # config (env > .env, empty secrets = absent),
                       # ToolError + result envelopes
scripts/               # seed-db.ts, demo.ts (transcript), record-demo.ts (cast)
tests/                 # 110 tests โ€” unit, security, real-stdio integration
docs/                  # architecture.md, demo.gif + demo.cast (real recording)
data/sandbox/          # the ONLY writable surface for the fs tools
.github/workflows/     # CI: lint โ†’ typecheck โ†’ tests+coverage โ†’ build โ†’ audit

๐Ÿงช Verification

npm run lint           # eslint โ€” 0 warnings
npm run typecheck      # tsc --noEmit strict โ€” 0 errors
npm run test:cov       # 110 tests offline; โ‰ฅ85% all metrics (actual ~98.5% lines)
npm run build          # compile to dist/
npm audit --omit=dev   # 0 vulnerabilities

CI (.github/workflows/ci.yml): lint โ†’ typecheck โ†’ tests+coverage โ†’ build โ†’ npm audit + gitleaks, on pull_request, Node 20/22/24 matrix, permissions: contents: read. A fork passes with zero secrets.

๐Ÿ› ๏ธ Stack

Node 20+ ยท TypeScript 5 strict ESM ยท @modelcontextprotocol/sdk (stdio transport, tools/list/tools/call, _meta correlation) ยท OpenAI tool-calling (chat.completions) ยท better-sqlite3 (readonly fixture DB) ยท zod (input schemas) ยท pino (JSON logs) ยท @prometheus-io/client (metrics) ยท Vitest + @vitest/coverage-v8 ยท ESLint 10 + typescript-eslint ยท Docker multi-stage.

Known limitations & next steps

  • stdio only. The transport trusts the spawning client โ€” the correct MCP model for a local server. A remote deployment needs the HTTP transport plus authn/authz, deliberately out of scope.

  • In-process rate limiting. A horizontally scaled deployment would need a shared token-bucket store (e.g. Redis).

  • Keyless search quality. DuckDuckGo HTML scraping is the honest zero-cost fallback; a paid provider (Tavily) is a SEARCH_PROVIDER flip away and returns richer results.

  • No streaming token UX. The event stream is structured JSON/pretty lines; an SSE surface for a browser console is a natural next step (AgentEvent union is already shaped for it).

  • SQLite-centric SQL guard. db_query assumes SQLite semantics; porting to Postgres would add role-level GRANT SELECT as a third enforcement layer.


๐Ÿ’ผ Need this for your own project?

I build production-shaped agentic AI systems โ€” MCP servers and clients, guardrail-first tool surfaces, cost-aware agent loops, and the observability discipline to actually trust them in production.

If you need an internal capability exposed safely to agents over an open protocol โ€” with sandboxing, injection defenses, audit-grade logging and a test suite that proves the guardrails fail loudly โ€” this codebase is the working template.

This project is the protocol layer underneath the Agentic AI portfolio ladder (agentic-api โ†’ agentic-rag-system โ†’ agentic-web-researcher โ†’ langgraph-multiagent-orchestrator): the open standard those agents consume tools through, implemented end-to-end.

๐Ÿ“„ License

MIT

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    A security-focused Model Context Protocol server that enables controlled local tool execution through strict network firewalls, filesystem protections, and rate-limiting policies. It features a plugin-based architecture for progressive tool discovery and includes reference implementations for web searching and bug tracking.
    15
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    A secure Model Context Protocol server providing HTTP endpoints for AI agent tool execution, including file system operations, shell commands, and LLM-based code generation.
    1
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    A Model Context Protocol server that advertises tools with JSON schemas and executes tool calls safely, enabling AI agents to perform actions on real systems.
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Exposes database and external API tools via the Model Context Protocol, enabling LLMs to query SQLite and call public APIs with validated inputs and error handling.
    MIT