Skip to main content
Glama
dkreme514

customer-mcp

by dkreme514
README.md
# AI Solutions Engineer Assessment Submission

Python 3.11 solution for the four implementation tasks supplied in the assessment. The project uses the official MCP Python SDK, FastAPI/httpx, and an on-disk SQLite token ledger.

## What is implemented

- **Task 1:** `get_customer_record` and `trigger_refund` MCP tools with strict Pydantic schemas. The same server supports stdio and Streamable HTTP. Logging stays on stderr so stdio JSON-RPC is not corrupted.
- **Task 2:** `/mcp` proxy parses JSON-RPC and bearer-token roles, forwards `tools/list` and ordinary calls unchanged, and intercepts unauthorized `admin_*` calls locally with JSON-RPC error `-32001`.
- **Task 3:** `/v1/chat/completions/stream` parses OpenAI-compatible SSE JSON deltas and incrementally redacts email addresses, SSNs, and payment-card-like numbers—even across delta and network-chunk boundaries. A bounded overlap controls memory and latency.
- **Task 4:** `/v1/chat/completions` uses a transactional SQLite sliding-window token limiter (50,000/minute by default), a 3-second primary timeout, failover on timeout or HTTP 429, and sanitized error envelopes.

## Run

```bash
python -m venv .venv
. .venv/bin/activate
pip install -e '.[dev]'
cp .env.example .env                 # export values with your preferred env loader

# Task 1, local MCP transport
customer-mcp

# Task 1, remote MCP transport
customer-mcp --transport streamable-http --port 8001

# Tasks 2-4 gateway
fde-gateway
```

Inspect the MCP server with `uv run mcp dev src/fde_assessment/mcp_server.py`. Run the suite with `pytest`.

## Security and production notes

The static environment-token mapping is intentionally compact for an assessment. In production, terminate TLS at a trusted proxy, validate JWT signature/issuer/audience, derive roles from signed claims, store secrets in a secret manager, and add audit logging without request bodies. The SQLite limiter is safe for concurrent workers on one host through WAL plus `BEGIN IMMEDIATE`; a multi-host deployment should use a shared transactional counter such as Redis.

Token reservations use an input estimate plus requested `max_tokens`, which is conservative and race-safe. A production implementation can reconcile the reservation with actual usage after a successful completion. The streaming guardrail intentionally uses bounded PII patterns: unbounded pattern matching and guaranteed low-latency emission are incompatible without a maximum match length.

TDQS

B3.3/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have completely distinct purposes: retrieving a customer record versus triggering a refund. There is no overlap or ambiguity in their responsibilities.

Naming Consistency5/5

Both tool names follow a consistent lowercase snake_case verb_noun pattern: get_customer_record and trigger_refund. This makes the naming predictable and easy to follow.

Tool Count3/5

At only two tools, the server feels minimal but not absurdly sparse. The operations are focused, yet the count is on the thin side for a customer-related server.

Completeness2/5

The tool surface lacks basic customer lifecycle operations such as create, update, delete, or list. It only supports retrieval and a single financial action, which leaves significant gaps for typical customer management workflows.

Maintenance

ActivityMaintained
ResponsivenessNo issues