Skip to main content
Glama
BhavikMoradiya

MCP Enterprise Tool Gateway

README.md
# MCP Enterprise Tool Gateway

[![CI](https://github.com/bmoradiya/mcp-enterprise-tool-gateway/actions/workflows/ci.yml/badge.svg)](https://github.com/bmoradiya/mcp-enterprise-tool-gateway/actions/workflows/ci.yml)

## Executive Summary

MCP Enterprise Tool Gateway is a production-oriented portfolio repository for enterprise AI architecture and implementation. It demonstrates MCP servers, tool schemas, RBAC, audit logs, rate limits, prompt-injection boundaries.

Current status: **architecture scaffold plus working service skeleton**. Benchmarks are intentionally not fabricated. This repository includes evaluation/benchmark methodology and executable hooks so measured results can be added only after real runs.

## Real Business Problem

LLM applications need governed access to internal systems without giving models unrestricted credentials or ambiguous tool permissions.

## Why AI Is Appropriate

MCP is appropriate because it standardizes tool discovery and execution boundaries between AI clients and enterprise tools.

## Architecture

```mermaid
flowchart LR
    Client["MCP Client"]
    Gateway["MCP Tool Gateway"]
    Policy["RBAC + Policy Engine"]
    Audit["Audit Log"]
    Tools["Enterprise Tools"]
    Approval["Dangerous Action Approval"]
    Client --> Gateway
    Gateway --> Policy
    Policy --> Tools
    Gateway --> Audit
    Policy --> Approval
```

## Request/Data Flow

1. A client calls the FastAPI boundary with a correlation ID.
2. Input is validated with typed Pydantic models.
3. The service layer applies policy, routing, retrieval, orchestration, or evaluation logic depending on the project.
4. Provider and infrastructure dependencies are accessed through interfaces so local development can use deterministic mocks.
5. Structured logs, traces, and metrics capture latency, errors, and AI-specific operational signals.

## Technology Decisions

Primary stack: Python MCP server/client skeleton, FastAPI control plane, Pydantic schemas, OpenTelemetry.

The repository favors typed Python, small modules, explicit interfaces, deterministic local tests, and optional cloud/provider integrations. AWS is the primary production architecture target where infrastructure is relevant, but local development must not require paid services.

## Repository Structure

```text
.
├── src/
├── tests/
├── docs/
│   ├── adr/
│   ├── architecture/
│   └── security/
├── examples/
├── infrastructure/
├── .github/workflows/
├── .env.example
├── pyproject.toml
└── README.md
```

## Prerequisites

Required for local development:

- Python 3.11 or newer
- Git
- `make`
- Internet access for the first dependency installation

Optional, depending on the implementation phase:

- Docker Desktop or another Docker-compatible runtime
- AWS CLI v2 configured with a non-production profile
- Terraform 1.6 or newer
- Ollama or vLLM for local model experiments
- Provider API keys for OpenAI, Anthropic, Google, or Amazon Bedrock

No provider key is required for the current scaffold. The default local provider mode is `mock`.

## Local Quick Start

```bash
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
pytest
uvicorn mcp_enterprise_tool_gateway.api:app --reload
```

Recommended one-command validation:

```bash
make validate
```

Health check:

```bash
curl http://127.0.0.1:8000/healthz
```

For a detailed local runbook, see `docs/local-development.md`.

## Configuration

Copy `.env.example` to `.env` for local development. Provider credentials are optional and should never be committed.

## Example Requests

```bash
curl -s http://127.0.0.1:8000/healthz
curl -s http://127.0.0.1:8000/readiness
```

## Example Outputs

```json
{"status":"ok","service":"mcp-enterprise-tool-gateway"}
```

## Evaluation Approach

The evaluation plan is documented in `docs/evaluation.md`. The repo includes a benchmark harness placeholder and testable metrics schema. Results must be generated from real runs before publication.

## Performance Considerations

Track p50/p95 latency, provider latency, tool latency, queue time where applicable, token usage, cost per request, cache hit rate, and error rate. Optimize only after measuring bottlenecks.

## Security Considerations

See `docs/security/threat-model.md`. The design assumes model inputs, retrieved documents, user uploads, and tool outputs are untrusted.

## Reliability Considerations

Production deployment should include timeouts, retries with backoff, circuit breakers, idempotency where applicable, health checks, readiness checks, structured logging, and alertable SLOs.

## Cost Considerations

Major cost drivers are model tokens, embeddings, vector/graph/search infrastructure, compute, storage, network transfer, and observability volume. This repository avoids publishing precise cost numbers until measured in a specific environment.

## Observability

The service skeleton exposes correlation-friendly boundaries. Production implementation should emit OpenTelemetry traces, structured logs, and AI metrics such as tokens/request, latency, tool success rate, groundedness, and evaluation score trends.

## Testing Strategy

Tests should cover deterministic business logic, provider contract boundaries, security/adversarial cases, and evaluation regressions. The current CI runs the scaffold tests.

## Deployment

Infrastructure examples live under `infrastructure/`. They are intentionally not auto-applied because cloud deployments may create paid resources.

## Architectural Tradeoffs

Important decisions are captured as ADRs in `docs/adr/`. Each ADR states context, options, decision, tradeoffs, and operational consequences.

## Limitations

- This initial version is a scaffold and vertical-slice foundation.
- Benchmarks are not published until measured.
- Cloud deployment modules are architecture-ready examples, not automatically deployed infrastructure.
- Provider adapters default to local/mock behavior until credentials are configured.

## Future Improvements

- Implement the complete domain workflow.
- Add provider-specific integrations.
- Add realistic sample datasets.
- Run measured benchmarks and publish reproducible reports.
- Add deeper security and adversarial test coverage.