MCP Agent Toolkit
README.md
<div align="center">
# ๐ MCP Agent Toolkit
**A production-shaped Model Context Protocol server + autonomous agent client โ an LLM discovers four guarded tools over the open standard at runtime, recovers from live tool errors, and produces a cited report, with every invocation logged, rate-limited, and correlated by request ID.**
[](https://nodejs.org/)
[](https://www.typescriptlang.org/)
[](https://modelcontextprotocol.io/)
[](https://github.com/kanderson-ai-dev/mcp-agent-toolkit/actions/workflows/ci.yml)
[](https://github.com/kanderson-ai-dev/mcp-agent-toolkit/actions/workflows/ci.yml)
[](https://eslint.org/)
[](https://www.typescriptlang.org/tsconfig#strict)
[](LICENSE)

*Real run, live OpenAI tool-calling loop โ watch it discover the tools, hit
a real SQL error (`no such column: category`), retry with the correct
schema, search the live web, and write a cited report into the sandbox.
Full transcript: [`docs/demo-transcript.txt`](docs/demo-transcript.txt) โ
recorded via `npm run demo:record` (asciinema cast in
[`docs/demo.cast`](docs/demo.cast)).*
</div>
---
## ๐ What this is
A **standalone MCP server + agent client** in strict TypeScript. The server
exposes four tools over the official [`@modelcontextprotocol/sdk`](https://github.com/modelcontextprotocol/typescript-sdk)
with stdio transport: `web_search`, `db_query` (read-only SQLite),
`read_file` and `write_file` (sandboxed). The client spawns the server as a
child process and speaks plain MCP โ `tools/list`, `tools/call`,
`params._meta` correlation โ exactly like Claude Desktop or any other
MCP-capable host would. Nothing is imported in-process; the protocol is the
boundary.
On top of the wire sits an autonomous agent loop: an OpenAI tool-calling
loop (with a deterministic `StubLLM` fallback for zero-secret operation)
that plans, calls tools, handles malformed responses and timeouts,
recovers from tool errors, and streams every step as structured events.
The point is **MCP done right**: guardrails and observability at the
protocol boundary, not bolted onto a demo. Every tool call is
schema-validated, rate-limited, sanitized on the way back (the web is
untrusted โ OWASP LLM01), logged as JSON with a correlated request ID, and
counted in Prometheus metrics.
---
## ๐ก Why this matters โ for your project, or for a technical reviewer
- ๐ **The protocol is the real boundary.** The agent discovers tools via
`tools/list` and invokes them via `tools/call` over JSON-RPC stdio โ no
shared imports, no framework glue. Swap the client for any MCP host and
the server just works.
- ๐ก๏ธ **Defense in depth, not a single regex.** `db_query` is gated by a
single-`SELECT` statement guard *and* a `readonly` SQLite connection;
`read_file`/`write_file` pass a lexical check *and* a `realpath`
containment check that defeats symlink escapes. Each layer fails closed
on its own.
- ๐งผ **Untrusted output never becomes instructions.** Every tool result is
sanitized server-side (control/bidi/zero-width chars stripped, injection
patterns neutralized, `MAX_TOOL_OUTPUT_CHARS` cap) before it reaches the
model context โ the OWASP LLM01 control for indirect prompt injection.
- ๐ **The server never sees the LLM key.** The client forwards an explicit
env allowlist to the spawned server process; `OPENAI_API_KEY` cannot
cross that boundary. Empty secrets count as absent, so CI runs with zero
credentials.
- ๐งพ **Request-ID correlation end-to-end.** The client stamps a UUID into
`params._meta.requestId` per call; the server logs it on every
invocation line โ client event stream joins server logs trivially.
- ๐ฆ **Every failure is a first-class signal.** Timeouts, malformed
responses, unknown tools and guardrail rejections normalize to typed
error envelopes โ the agent sees a clean `failure` reason, not a stack
trace (the demo transcript shows a live SQL error recovered).
- ๐งช **The transport is real in tests too.** The integration suite spawns
the actual server over stdio and asserts request-ID correlation and
rate-limit enforcement across the wire โ not a mock of MCP.
- ๐ **CI green with zero secrets.** Fork, clone, `npm ci && npm test` โ
110 tests pass offline; the deterministic `StubLLM` and
`SEARCH_PROVIDER=none` need no keys.
---
## ๐๏ธ Architecture
```mermaid
graph LR
subgraph client["Agent client"]
CLI[CLI entry] --> LOOP[Agent loop<br/>OpenAI tool calling]
LOOP --> MC[MCP client<br/>StdioClientTransport]
LOOP -.->|no OPENAI_API_KEY| STUB[StubLLM<br/>deterministic]
LOOP -->|chat.completions| OAI[(OpenAI API)]
end
MC ==>|"JSON-RPC over stdio<br/>tools/list ยท tools/call<br/>requestId via params._meta"| MS
subgraph server["MCP server (child process)"]
MS[McpServer<br/>StdioServerTransport] --> REG[Tool registry<br/>+ guardrail wrapper]
REG --> RL[Rate limiter<br/>token bucket]
RL --> T1[web_search]
RL --> T2[db_query]
RL --> T3[read_file]
RL --> T4[write_file]
T1 --> SANE[Output sanitizer]
T2 --> SQLG[SELECT-only guard<br/>+ readonly conn]
T3 --> FS[Path sandbox]
T4 --> FS
REG --> LOG[pino JSON โ stderr]
REG --> MET[Prometheus metrics<br/>optional :METRICS_PORT]
end
T1 -->|HTTPS| EXT1[(Tavily /<br/>DuckDuckGo)]
T2 --> EXT2[(SQLite<br/>readonly)]
T3 --> EXT3[(data/sandbox/)]
T4 --> EXT3
```
Full rationale and the tool/guardrail matrix:
[`docs/architecture.md`](docs/architecture.md).
---
## ๐ Evaluation-Driven Development (EDD)
The quality gates are binary and enforced in CI โ a regression fails the
build rather than being shrugged off:
| Dimension | Metric | Threshold | Latest run |
|---|---|---|---|
| Test suite | unit + security + real-stdio integration | 100% green | 110/110 โ
|
| Coverage | lines / branches / functions / statements | โฅ 85% each | 98.5 / 85.1 / 100 / 97.5 โ
|
| Type safety | `tsc --strict` | 0 errors | โ
|
| Lint | `eslint` strict (typescript-eslint) | 0 warnings | โ
|
| Security suite | traversal ยท write SQL ยท injection ยท rate-limit rejections | all must **fail loudly** | โ
|
| Dependency audit | `npm audit --omit=dev` | 0 high/critical | โ
|
Reproduce: `npm run test:cov` (coverage gate), `npm run verify`-equivalent
sequence in [`CONTRIBUTING.md`](CONTRIBUTING.md).
---
## ๐ API
The server speaks standard MCP (JSON-RPC 2.0 over stdio): `initialize`,
`tools/list`, `tools/call`. Tools:
| Tool | Input | Guardrails | Returns |
|---|---|---|---|
| `web_search` | `query` (1โ400 chars), `max_results` (default 5) | provider allowlist, timeout, output sanitizer + size cap | `{ provider, results: [{title, url, snippet}] }` |
| `db_query` | `query` โ a single `SELECT`/`WITH` | statement guard **+** `readonly` connection, 500-row cap | `{ rows, row_count, truncated }` |
| `read_file` | `path` relative to sandbox | lexical + `realpath` containment, `READ_FILE_MAX_BYTES` cap | `{ path, content }` |
| `write_file` | `path`, `content` | same containment, `WRITE_FILE_MAX_BYTES` cap | `{ path, bytes_written }` |
Example `tools/call` payload (real shape โ the client adds
`_meta.requestId` automatically):
```json
{
"jsonrpc": "2.0", "id": 7, "method": "tools/call",
"params": {
"name": "db_query",
"arguments": { "query": "SELECT title, author FROM reports ORDER BY published_at DESC LIMIT 5" },
"_meta": { "requestId": "de30734e-8f77-4418-8a27-65557d8ade40" }
}
}
```
Agent CLI:
```bash
npm run agent -- "question" # live OpenAI loop (needs OPENAI_API_KEY)
npm run agent -- "question" --stub # deterministic, zero secrets
npm run agent -- "question" --json # machine-readable event stream
npm run demo # seeded E2E run โ docs/demo-transcript.txt
npm run demo:record # record docs/demo.cast โ GIF via asciinema/agg
```
---
## ๐ก Observability & cost
- **pino JSON logs on stderr** (stdout is reserved for the MCP transport):
every tool invocation logs tool name, sanitized input args, output
preview, duration, ok/error, `session_id` and `request_id`.
- **Request-ID correlation**: the client stamps `params._meta.requestId`;
the server echoes it on every log line โ join client events to server
logs by UUID.
- **Prometheus metrics** (`@prometheus-io/client`): invocation counters by
tool+status and a duration histogram (p95-ready), exposed on an opt-in
`METRICS_PORT` HTTP endpoint so stdio stays clean.
- **Cost**: the only metered dependency is the LLM. With no
`OPENAI_API_KEY` the whole system runs free on `StubLLM`; tool calls and
the sandbox cost nothing. `MAX_TOOL_ITERATIONS` bounds worst-case token
spend per run.
Real server log line from the recorded demo:
```json
{"level":30,"service":"mcp-server","event":"tool_invocation","tool":"db_query",
"request_id":"de30734e-8f77-4418-8a27-65557d8ade40","session_id":"3bb996d6-โฆ",
"duration_ms":0,"ok":true,"input":"{\"query\":\"SELECT title, author FROM reports\"}",
"output":"[{\"type\":\"text\",\"text\":\"{\\\"status\\\": \\\"ok\\\"โฆ"}
```
---
## ๐ Security posture
Sandboxed filesystem (lexical + `realpath` containment), read-only SQL
(statement guard **and** `readonly` connection), output sanitization
against indirect prompt injection (OWASP LLM01), per-tool/session rate
limiting, env allowlist so the server never sees `OPENAI_API_KEY`,
secret-redacted logs, `gitleaks` + `npm audit` in CI on `pull_request`
with `contents: read`. Full threat model and mechanism table:
[`SECURITY.md`](SECURITY.md).
---
## ๐ Quick Start
```bash
npm ci
cp .env.example .env # optional โ paste OPENAI_API_KEY for the live loop
npm run seed:db # fixture DB for db_query (8 reports)
npm run demo # recorded E2E: seeds + runs, saves transcript
npm run agent -- "your question"
```
Degrades cleanly with zero secrets: no `OPENAI_API_KEY` โ deterministic
`StubLLM`; no `SEARCH_API_KEY` โ keyless DuckDuckGo; `SEARCH_PROVIDER=none`
โ fully offline. The entire CI suite runs with no credentials.
```bash
docker compose up # runs the default demo question once
docker compose run --rm agent "โฆ" # custom question
docker compose run --rm agent --stub "offline run, no keys"
```
The image is multi-stage (non-root `node` user), seeds the fixture DB at
build time, and mounts only `data/sandbox/` as a volume โ the single
writable surface. `.env` is read at runtime via `env_file`; nothing is
baked in.
## โ๏ธ Environment Variables
| Variable | Required | Secret? | Purpose |
|---|---|---|---|
| `OPENAI_API_KEY` | no | โ
| Live LLM for the agent loop; `StubLLM` when absent |
| `CHAT_MODEL_NAME` | no | โ | Model override (default `gpt-4o-mini`) |
| `LLM_PROVIDER` | no | โ | `auto` \| `openai` \| `stub` |
| `SEARCH_PROVIDER` | no | โ | `auto` \| `tavily` \| `duckduckgo` \| `none` |
| `SEARCH_API_KEY` | no | โ
| Tavily key (optional paid provider) |
| `MCP_SERVER_COMMAND` / `MCP_SERVER_ARGS` | no | โ | Override how the client spawns the server |
| `SANDBOX_ROOT` | no | โ | The only filesystem root tools may touch |
| `DB_PATH` | no | โ | SQLite fixture path (`npm run seed:db`) |
| `RATE_LIMIT_PER_MINUTE` | no | โ | Per tool, per session (default 60) |
| `MAX_TOOL_ITERATIONS` | no | โ | Agent loop bound โ must terminate |
| `MAX_TOOL_OUTPUT_CHARS` | no | โ | Tool output truncation cap (untrusted data) |
| `READ_FILE_MAX_BYTES` / `WRITE_FILE_MAX_BYTES` | no | โ | FS size caps |
| `LOG_LEVEL` | no | โ | pino level; logs go to stderr |
| `METRICS_PORT` | no | โ | Empty = disabled; e.g. `9108` โ `/metrics` HTTP |
| `OUTPUT_FORMAT` | no | โ | `pretty` \| `json` agent event stream |
Never commit secrets. See [`.env.example`](.env.example) for the full,
commented template.
## ๐ Project Structure
```
src/
โโโ server/ # MCP server (stdio) โ separate process
โ โโโ server.ts # tool registration + guardrail wrapper + logging
โ โโโ tools/ # web_search, db_query, read_file, write_file
โ โโโ guardrails/ # path sandbox, SQL guard, rate limiter, sanitizer
โ โโโ observability/ # pino logger, Prometheus metrics, request ctx
โโโ client/ # agent CLI โ spawns the server, speaks MCP
โ โโโ mcp-client.ts # tools/list + tools/call + env allowlist
โ โโโ agent-loop.ts # tool-calling loop (bounded, must terminate)
โ โโโ events.ts # structured event stream (pretty|json)
โ โโโ llm/ # OpenAI impl + deterministic StubLLM
โโโ shared/ # config (env > .env, empty secrets = absent),
# ToolError + result envelopes
scripts/ # seed-db.ts, demo.ts (transcript), record-demo.ts (cast)
tests/ # 110 tests โ unit, security, real-stdio integration
docs/ # architecture.md, demo.gif + demo.cast (real recording)
data/sandbox/ # the ONLY writable surface for the fs tools
.github/workflows/ # CI: lint โ typecheck โ tests+coverage โ build โ audit
```
## ๐งช Verification
```bash
npm run lint # eslint โ 0 warnings
npm run typecheck # tsc --noEmit strict โ 0 errors
npm run test:cov # 110 tests offline; โฅ85% all metrics (actual ~98.5% lines)
npm run build # compile to dist/
npm audit --omit=dev # 0 vulnerabilities
```
CI (`.github/workflows/ci.yml`): lint โ typecheck โ tests+coverage โ build
โ `npm audit` + `gitleaks`, on `pull_request`, Node 20/22/24 matrix,
`permissions: contents: read`. A fork passes with zero secrets.
## ๐ ๏ธ Stack
Node 20+ ยท TypeScript 5 strict ESM ยท `@modelcontextprotocol/sdk` (stdio
transport, `tools/list`/`tools/call`, `_meta` correlation) ยท OpenAI
tool-calling (`chat.completions`) ยท `better-sqlite3` (readonly fixture DB)
ยท `zod` (input schemas) ยท `pino` (JSON logs) ยท `@prometheus-io/client`
(metrics) ยท Vitest + `@vitest/coverage-v8` ยท ESLint 10 +
typescript-eslint ยท Docker multi-stage.
## Known limitations & next steps
- **stdio only.** The transport trusts the spawning client โ the correct
MCP model for a local server. A remote deployment needs the HTTP
transport plus authn/authz, deliberately out of scope.
- **In-process rate limiting.** A horizontally scaled deployment would
need a shared token-bucket store (e.g. Redis).
- **Keyless search quality.** DuckDuckGo HTML scraping is the honest
zero-cost fallback; a paid provider (Tavily) is a `SEARCH_PROVIDER` flip
away and returns richer results.
- **No streaming token UX.** The event stream is structured JSON/pretty
lines; an SSE surface for a browser console is a natural next step
(`AgentEvent` union is already shaped for it).
- **SQLite-centric SQL guard.** `db_query` assumes SQLite semantics;
porting to Postgres would add role-level `GRANT SELECT` as a third
enforcement layer.
---
## ๐ผ Need this for your own project?
**I build production-shaped agentic AI systems โ MCP servers and clients,
guardrail-first tool surfaces, cost-aware agent loops, and the
observability discipline to actually trust them in production.**
If you need an internal capability exposed safely to agents over an open
protocol โ with sandboxing, injection defenses, audit-grade logging and a
test suite that proves the guardrails fail loudly โ this codebase is the
working template.
This project is the protocol layer underneath the Agentic AI portfolio
ladder (`agentic-api` โ `agentic-rag-system` โ `agentic-web-researcher` โ
`langgraph-multiagent-orchestrator`): the open standard those agents
consume tools through, implemented end-to-end.
## ๐ License
[MIT](LICENSE)
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues