Skip to main content
Glama
dkreme514

customer-mcp

by dkreme514

AI Solutions Engineer Assessment Submission

Python 3.11 solution for the four implementation tasks supplied in the assessment. The project uses the official MCP Python SDK, FastAPI/httpx, and an on-disk SQLite token ledger.

What is implemented

  • Task 1: get_customer_record and trigger_refund MCP tools with strict Pydantic schemas. The same server supports stdio and Streamable HTTP. Logging stays on stderr so stdio JSON-RPC is not corrupted.

  • Task 2: /mcp proxy parses JSON-RPC and bearer-token roles, forwards tools/list and ordinary calls unchanged, and intercepts unauthorized admin_* calls locally with JSON-RPC error -32001.

  • Task 3: /v1/chat/completions/stream parses OpenAI-compatible SSE JSON deltas and incrementally redacts email addresses, SSNs, and payment-card-like numbers—even across delta and network-chunk boundaries. A bounded overlap controls memory and latency.

  • Task 4: /v1/chat/completions uses a transactional SQLite sliding-window token limiter (50,000/minute by default), a 3-second primary timeout, failover on timeout or HTTP 429, and sanitized error envelopes.

Related MCP server: MCP Customer Support Demo

Run

python -m venv .venv
. .venv/bin/activate
pip install -e '.[dev]'
cp .env.example .env                 # export values with your preferred env loader

# Task 1, local MCP transport
customer-mcp

# Task 1, remote MCP transport
customer-mcp --transport streamable-http --port 8001

# Tasks 2-4 gateway
fde-gateway

Inspect the MCP server with uv run mcp dev src/fde_assessment/mcp_server.py. Run the suite with pytest.

Security and production notes

The static environment-token mapping is intentionally compact for an assessment. In production, terminate TLS at a trusted proxy, validate JWT signature/issuer/audience, derive roles from signed claims, store secrets in a secret manager, and add audit logging without request bodies. The SQLite limiter is safe for concurrent workers on one host through WAL plus BEGIN IMMEDIATE; a multi-host deployment should use a shared transactional counter such as Redis.

Token reservations use an input estimate plus requested max_tokens, which is conservative and race-safe. A production implementation can reconcile the reservation with actual usage after a successful completion. The streaming guardrail intentionally uses bounded PII patterns: unbounded pattern matching and guaranteed low-latency emission are incompatible without a maximum match length.

Available Tools

2 tools
get_customer_recordA

Return a customer record for an ID formatted exactly as CUST-XXXXX.

ParametersJSON Schema
NameRequiredDescriptionDefault
customer_idYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It states the core behavior ('Return a customer record') but does not disclose error handling, not-found behavior, authentication needs, or whether the operation is strictly read-only. The read intent is implied but not explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. Every word contributes to the tool's purpose and input requirement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter get operation, the description covers the action and input format. However, with no output schema and no annotations, it leaves the return structure and error behavior unspecified, so an agent has incomplete information about what to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds a human-readable format ('CUST-XXXXX') that reinforces the schema's pattern, but it does not explain the semantic meaning of customer_id beyond what the tool name and schema title already imply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and resource ('customer record') and clearly scopes it to an ID with a strict format. This makes the tool's purpose unambiguous and easily distinguished from the sibling trigger_refund.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus trigger_refund or any other alternative. The description only implies that it is for retrieving a customer record by ID, but does not state contexts, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trigger_refundC

Trigger a refund after strictly validating the customer, amount, and reason.

ParametersJSON Schema
NameRequiredDescriptionDefault
amountYes
reasonYes
customer_idYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It reveals a 'strictly validating' step, but does not explain the side effects of triggering a refund (e.g., whether it is immediately executed, irreversible, requires permissions, or what happens on validation failure). The term 'trigger' is vague about the actual mutation and its consequences.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that states the action first and then the validation condition. It contains no filler or redundant information, making it appropriately concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with three required parameters, no annotations, and no output schema, the description is too thin. It does not explain the expected outcome, error conditions, or any post-validation behavior, leaving an agent without enough context to predict the tool's full effect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it merely lists 'customer, amount, and reason' without adding any meaning beyond the parameter names. It does not explain the format, purpose, or relationships between parameters, nor does it clarify constraints like the customer_id pattern or reason length.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Trigger a refund') and a specific resource ('refund'), with a clear scope ('after strictly validating the customer, amount, and reason'). It is easily distinguishable from the sibling tool get_customer_record, which is a read operation, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide explicit guidance on when to use this tool versus the sibling get_customer_record, nor does it mention any exclusions or alternative tools. The only implied usage is that it is for triggering refunds, but no contextual conditions or prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 2 tool updatesv0.1.0
    • First observedget_customer_record
    • First observedtrigger_refund

TDQS

B3.3/5.0
Disambiguation5/5

The two tools have completely distinct purposes: retrieving a customer record versus triggering a refund. There is no overlap or ambiguity in their responsibilities.

Naming Consistency5/5

Both tool names follow a consistent lowercase snake_case verb_noun pattern: get_customer_record and trigger_refund. This makes the naming predictable and easy to follow.

Tool Count3/5

At only two tools, the server feels minimal but not absurdly sparse. The operations are focused, yet the count is on the thin side for a customer-related server.

Completeness2/5

The tool surface lacks basic customer lifecycle operations such as create, update, delete, or list. It only supports retrieval and a single financial action, which leaves significant gaps for typical customer management workflows.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    C
    maintenance
    Provides order lookup, customer lookup, and refund issuance tools with categorized errors to ensure accurate routing and distinguish access failures from valid empty results.
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server for regulated enterprises, providing per-tool RBAC, redacted audit logging, and structured error handling. Exposes bank tools for customer lookup, statement search, and dispute resolution over stdio and HTTPS transports.
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables MCP-compatible clients to retrieve invoice and purchase order data live over stdio through tools for listing invoices, getting invoice details, and fetching purchase orders.
    -

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/dkreme514/quilr-ai-solutions-engineer-assessment'

If you have feedback or need assistance with the MCP directory API, please join our Discord server