Skip to main content
Glama
eyasu11321238a

Secure Clinical LLM Access via MCP

README.md
# Secure Clinical LLM Access via MCP

**A research-oriented framework for secure access to structured clinical data using open-weight large language models and the Model Context Protocol (MCP).**

[![Python](https://img.shields.io/badge/Python-3.10%2B-blue.svg)](https://www.python.org/)
[![MCP](https://img.shields.io/badge/MCP-Model%20Context%20Protocol-purple.svg)](https://modelcontextprotocol.io/)
[![Ollama](https://img.shields.io/badge/LLM-Ollama-black.svg)](https://ollama.com/)
[![SQLite](https://img.shields.io/badge/Database-SQLite-lightgrey.svg)](https://www.sqlite.org/)
[![Synthea](https://img.shields.io/badge/Data-Synthetic%20Clinical%20Data-green.svg)](https://synthetichealth.github.io/synthea/)

## Overview

Large language models can provide an intuitive interface for exploring complex data, but directly connecting an LLM to a clinical database introduces important challenges around **data access, security, reliability, and control**.

This project investigates a controlled alternative: exposing structured clinical data to an LLM through a set of **narrowly scoped, validated, and auditable tools using the Model Context Protocol (MCP)**.

Instead of allowing the model to generate and execute arbitrary SQL, the system provides purpose-specific tools for accessing clinical information. This creates an explicit security boundary between the language model and the underlying database.

The project combines:

* Open-weight LLMs
* Model Context Protocol (MCP)
* Structured relational clinical data
* Python-based data transformation pipelines
* Tool-scoped database access
* Input validation and result-size limits
* Audit logging
* Security testing
* LLM-based clinical question answering

**All clinical data used in this project is synthetic. No real patient data is stored or processed.**

---

## Research Motivation

The project explores the following question:

> **How can open-weight language models access and analyze structured clinical data while maintaining a controlled and auditable data-access boundary?**

This is particularly relevant to clinical research environments where medical professionals may need natural-language access to complex datasets without giving an LLM unrestricted access to the underlying database.

The system therefore focuses on the interface between:

```text
Human
  │
  │ Natural-language question
  ▼
Open-weight LLM
  │
  │ Tool selection
  ▼
MCP Interface
  │
  │ Validated, scoped request
  ▼
Clinical Database
  │
  ▼
Synthetic Clinical Data
```

---

## System Architecture

```text
                  ┌──────────────────────────┐
                  │       User / Clinician   │
                  └────────────┬─────────────┘
                               │
                       Natural-language
                            question
                               │
                               ▼
                  ┌──────────────────────────┐
                  │    Open-weight LLM       │
                  │      (Ollama)             │
                  └────────────┬─────────────┘
                               │
                         Tool selection
                               │
                               ▼
                  ┌──────────────────────────┐
                  │       MCP Client          │
                  │     Clinical Agent        │
                  └────────────┬─────────────┘
                               │
                         MCP protocol
                               │
                               ▼
                  ┌──────────────────────────┐
                  │       MCP Server          │
                  │                          │
                  │  • Input validation      │
                  │  • Tool scoping          │
                  │  • Result limits         │
                  │  • PII exclusion         │
                  │  • Audit logging          │
                  └────────────┬─────────────┘
                               │
                       Controlled tools
                               │
                               ▼
                  ┌──────────────────────────┐
                  │    Relational Database   │
                  │        SQLite            │
                  └────────────┬─────────────┘
                               │
                               ▼
                  ┌──────────────────────────┐
                  │   Synthetic Clinical     │
                  │        Data              │
                  │        Synthea            │
                  └──────────────────────────┘
```

---

## Data Pipeline

Clinical data is generated using **Synthea**, a synthetic patient generator designed to produce realistic but non-identifiable medical records.

The generated heterogeneous CSV files are transformed into a linked relational database:

```text
Synthea
   │
   ├── Patients
   ├── Encounters
   ├── Conditions
   ├── Observations
   └── Medications
          │
          ▼
     Ingestion Pipeline
          │
          ▼
   Relational SQLite DB
```

The ingestion pipeline is implemented in Python and establishes relationships between patients, encounters, diagnoses, observations, and medications.

---

## Why MCP Instead of Direct SQL?

A central design decision is **not to expose a generic SQL execution tool to the language model**.

A naive architecture might look like:

```text
LLM → Generate SQL → Execute SQL → Database
```

This approach gives the model significant control over the database and creates unnecessary security risks.

This project instead uses:

```text
LLM
 │
 ▼
MCP Tool
 │
 ├── Validate parameters
 ├── Restrict operation
 ├── Limit result size
 ├── Exclude identifying fields
 └── Record audit event
 │
 ▼
Database
```

Each tool has a narrowly defined purpose.

For example:

```text
get_patient_summary()
get_conditions()
get_medications()
get_lab_trends()
get_encounters()
```

The model can select and parameterize these tools, but it cannot arbitrarily execute SQL.

This creates a **tool-level security boundary** that is independent of the model's ability to follow system prompts.

---

## Security Model

The project treats the LLM as an **untrusted component**.

Security controls are therefore implemented outside the model wherever possible.

### Implemented controls

* **Tool-scoped database access**
* **Input validation**
* **Read-only operations**
* **Result-size limits**
* **Exclusion of identifying fields**
* **Audit logging**
* **Protocol-level testing**
* **Security test suite**

The design goal is:

> **Do not rely on the LLM to enforce security policies that can be enforced by the application layer.**

This is particularly important when language models are connected to sensitive data sources.

---

## MCP Tools

The MCP server currently exposes five controlled tools.

| Tool                  | Purpose                               |
| --------------------- | ------------------------------------- |
| `get_patient_summary` | Retrieve a restricted patient summary |
| `get_conditions`      | Retrieve patient conditions           |
| `get_medications`     | Retrieve medications                  |
| `get_lab_trends`      | Analyze laboratory observations       |
| `get_encounters`      | Retrieve encounter information        |

Each tool validates its arguments before accessing the database.

---

## LLM Integration

The project is designed to work with **local open-weight language models through Ollama**.

Example:

```text
User:
Find a female patient and summarize her conditions.

        ↓

Open-weight LLM

        ↓

MCP tool selection

        ↓

get_patient_summary(...)

        ↓

Validated database query

        ↓

Structured result

        ↓

LLM-generated response
```

The architecture keeps the clinical database local rather than requiring patient data to be sent to an external LLM API.

---

## Evaluation

The project includes an evaluation framework covering both **security and usability**.

### Security evaluation

The security test suite checks whether the MCP layer correctly prevents unauthorized or unsafe operations.

Examples include:

* Invalid tool parameters
* Excessive result requests
* Unauthorized field access
* Unsafe database operations
* Tool-boundary violations
* Audit logging behavior

Current security test status:

```text
7 / 7 security tests passing
```

### Usability evaluation

The evaluation framework is designed to measure:

* Question-answer accuracy
* Tool-selection accuracy
* Response latency
* Successful completion of clinical queries
* Failure cases

The goal is to evaluate not only whether the system works, but **how reliably an LLM can interact with structured clinical data through constrained tools**.

---

## Project Structure

```text
secure-clinical-llm-mcp/
│
├── agent/
│   ├── __init__.py
│   └── clinical_agent.py
│
├── eval/
│   ├── security_tests.py
│   ├── usability_benchmark.py
│   └── EVALUATION_REPORT.md
│
├── mcp_server/
│   └── server.py
│
├── pipeline/
│   └── ingest.py
│
├── synthea/
│   └── synthea.properties
│
├── requirements.txt
├── README.md
└── .gitignore
```

---

## Installation

### 1. Clone the repository

```bash
git clone https://github.com/eyasu11321238a/secure-clinical-llm-mcp.git
cd secure-clinical-llm-mcp
```

### 2. Install Python dependencies

```bash
pip install -r requirements.txt
```

### 3. Download Synthea

Synthea is not included in the repository.

Download the latest release:

```bash
cd synthea

curl -L -o synthea-with-dependencies.jar \
  https://github.com/synthetichealth/synthea/releases/download/master-branch-latest/synthea-with-dependencies.jar
```

Java 17 or newer is required.

### 4. Generate synthetic clinical data

```bash
java -jar synthea-with-dependencies.jar \
  -p 25 \
  -c synthea.properties \
  Massachusetts
```

### 5. Build the relational database

```bash
cd ../pipeline

python ingest.py \
  --csv-dir ../synthea/output/csv \
  --db ../data/clinical.db
```

### 6. Start the MCP server

```bash
cd ../mcp_server

python server.py
```

### 7. Run the security tests

```bash
cd ../eval

python security_tests.py
```

---

## Local LLM

The agent can be connected to an open-weight model running locally through Ollama.

For example:

```bash
ollama pull llama3.2:3b
```

Then run:

```bash
python agent/clinical_agent.py \
  "Find a female patient and summarize her conditions"
```

The agent communicates with the MCP server rather than accessing the database directly.

---

## Current Status

### Completed

* [x] Synthetic clinical data generation with Synthea
* [x] Heterogeneous clinical-data ingestion pipeline
* [x] Relational SQLite clinical database
* [x] MCP server
* [x] Five scoped clinical-data tools
* [x] Input validation
* [x] Result-size restrictions
* [x] Identifying-field exclusion
* [x] Audit logging
* [x] MCP protocol-level testing
* [x] Local Ollama agent integration
* [x] Security test suite
* [x] 7/7 security tests passing

### In Progress

* [ ] Schema-aware clinical query layer
* [ ] Expanded clinical question benchmark
* [ ] Tool-selection accuracy evaluation
* [ ] Latency evaluation
* [ ] Prompt-injection evaluation
* [ ] Expanded security analysis
* [ ] Research evaluation report

---

## Future Work

Several extensions are planned to move the prototype toward a more comprehensive research framework.

### 1. Schema-aware reasoning

Provide the LLM with structured descriptions of the relational schema and investigate how effectively it can select appropriate tools and parameters.

### 2. Expanded clinical benchmark

Develop a benchmark containing clinical questions ranging from simple lookups to multi-table analytical queries.

### 3. Security evaluation

Evaluate robustness against:

* Prompt injection
* Malicious clinical text
* Unauthorized data requests
* Tool manipulation
* Excessive data retrieval
* Attempts to bypass access controls

### 4. Heterogeneous medical data

Extend the pipeline to support additional healthcare data representations, including FHIR resources.

### 5. Model comparison

Compare multiple open-weight LLMs with respect to:

* Tool-selection accuracy
* Clinical query accuracy
* Latency
* Failure rate
* Security robustness

---

## Limitations

This project is a research prototype and **not a clinical decision-support system**.

The data is entirely synthetic and therefore does not represent the full complexity, noise, missingness, or distribution of real-world hospital data.

The system should not be used for medical diagnosis, treatment decisions, or real patient care.

Further validation would be required before applying the approach to real clinical environments.

---

## Research Relevance

This project sits at the intersection of:

```text
Large Language Models
        +
Structured Data
        +
Medical Informatics
        +
MCP / Tool Calling
        +
Data Engineering
        +
AI Security
        +
Evaluation
```

The central objective is to investigate how language models can provide a natural-language interface to structured clinical data **without giving the model unrestricted access to the underlying data source**.

---