Skip to main content
Glama
channico

MCP Knowledge Assistant

by channico
README.md
# MCP Knowledge Assistant

A read-only Model Context Protocol (MCP) server that lets an AI client search a
knowledge base and retrieve complete source documents. The project starts with a
small local prototype, then applies the same tool contract to semantic retrieval
from an OpenAI vector store.

## What this demonstrates

- MCP tool design with narrow `search` and `fetch` responsibilities
- Structured inputs and outputs using FastMCP and Pydantic
- Semantic retrieval over uploaded documents in an OpenAI vector store
- Document-level deduplication when vector search returns several matching chunks
- Streamable HTTP and `stdio` MCP transports
- End-to-end tool use through the OpenAI Responses API and a secure MCP tunnel
- Configuration through environment variables, with no credentials in source code
- Unit and protocol-level tests that do not make paid API calls

## Architecture

The following sequence shows the verified ChatGPT path. The tunnel client makes
an outbound connection, so the local MCP server does not need a publicly exposed
inbound port.

```mermaid
sequenceDiagram
    actor User
    participant ChatGPT
    participant Plugin as ChatGPT plugin
    participant Tunnel as OpenAI secure tunnel
    participant MCP as Local FastMCP server
    participant Store as OpenAI vector store

    User->>ChatGPT: Ask a question about the knowledge base
    ChatGPT->>Plugin: Call search(query)
    Plugin->>Tunnel: Send MCP tool request
    Tunnel->>MCP: Forward request over Streamable HTTP
    MCP->>Store: Run semantic search
    Store-->>MCP: Return matching chunks and file IDs
    MCP-->>Tunnel: Return deduplicated document results
    Tunnel-->>Plugin: Relay search results
    Plugin-->>ChatGPT: Provide document references

    ChatGPT->>Plugin: Call fetch(id)
    Plugin->>Tunnel: Send MCP tool request
    Tunnel->>MCP: Forward request over Streamable HTTP
    MCP->>Store: Retrieve the selected file content
    Store-->>MCP: Return parsed document content
    MCP-->>Tunnel: Return document and metadata
    Tunnel-->>Plugin: Relay fetched document
    Plugin-->>ChatGPT: Provide source content
    ChatGPT-->>User: Compose a grounded answer with a citation
```

`api_client.py` follows the same tunnel-to-server path, with the OpenAI
Responses API acting as the MCP client instead of the ChatGPT plugin.

The repository also includes a fully local learning path:

```text
Local demo client -> FastMCP server (stdio) -> data/documents.json
```

Both servers expose the same public tool contract:

| Tool | Input | Purpose |
| --- | --- | --- |
| `search` | `query: string` | Return compact, relevant document references. |
| `fetch` | `id: string` | Retrieve one complete document selected from search. |

Keeping discovery separate from retrieval avoids sending full documents before
they are needed and gives the model stable document IDs to use in later calls.

## Project structure

```text
.
├── data/documents.json                   # Sample local knowledge base
├── sample_data/cats.pdf                  # Public-domain vector-store sample
├── src/mcp_knowledge_assistant/
│   ├── knowledge_base.py                 # Local keyword retrieval
│   ├── models.py                         # Shared response schemas
│   ├── server.py                         # Local stdio MCP server
│   └── vector_store_server.py            # OpenAI vector-store MCP server
├── tests/                                # Offline unit and MCP tests
├── demo_client.py                        # Local stdio demonstration
├── vector_store_demo_client.py           # Direct HTTP MCP demonstration
└── api_client.py                         # Responses API + secure tunnel demonstration
```

## Requirements

- Python 3.11 or newer
- An OpenAI API project with billing enabled for the vector-store path
- The included sample PDF, or your own document, uploaded to an OpenAI vector store
- The OpenAI tunnel client only for the secure-tunnel demonstration

The local JSON server and the complete test suite require no API key.

### Sample document attribution

The vector-store demonstration uses
[*Cats: Their Points and Characteristics*](https://cdn.openai.com/API/docs/cats.pdf)
by W. Gordon Stables, Project Gutenberg eBook #43429. The sample PDF is hosted
by OpenAI and was produced from the Project Gutenberg edition. See the PDF for
the Project Gutenberg licence and applicable reuse terms.

## Setup

Clone the repository, create a virtual environment, and install the project:

```bash
python -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[dev]"
```

For OpenAI-backed examples, copy the environment template:

```bash
cp .env.example .env.local
```

Then add your own values to `.env.local`:

```dotenv
OPENAI_API_KEY=your_project_api_key
VECTOR_STORE_ID=vs_your_vector_store_id
```

`.env.local`, PyCharm settings, virtual environments, and local tunnel profiles
are excluded from Git.

## 1. Run the local prototype

The first server uses `stdio`, so the MCP client launches it as a child process
and communicates over standard input and output:

```bash
python demo_client.py
```

The demo discovers both tools, searches the sample JSON knowledge base, and
fetches the selected document.

You can also start the server through its installed command:

```bash
mcp-knowledge-assistant
```

A silent, waiting process is normal for a `stdio` server that has no connected
client.

## 2. Run the vector-store server

Upload `sample_data/cats.pdf` to an OpenAI vector store, then set
`OPENAI_API_KEY` and `VECTOR_STORE_ID` in `.env.local`. You may substitute your
own document and query if preferred. Start the Streamable HTTP server:

```bash
mcp-vector-store-assistant
```

By default, its MCP endpoint is:

```text
http://127.0.0.1:8000/mcp
```

In a second terminal, test the endpoint directly:

```bash
python vector_store_demo_client.py
```

Vector search operates on chunks, so a long document may produce several
matches with the same file ID. The MCP `search` tool deliberately collapses
those matches into one document result. The `fetch` tool then retrieves and
combines that document's parsed content for the model.

## 3. Call it through the Responses API

### Download and configure tunnel-client

1. Create a tunnel in [OpenAI Platform tunnel settings](https://platform.openai.com/settings/organization/tunnels)
   and copy its `tunnel_id`. Associate it with the Platform organization used
   by the API request and, for section 4, the intended ChatGPT workspace.
2. Download the archive for your operating system and CPU from the
   [latest official tunnel-client release](https://github.com/openai/tunnel-client/releases/latest).
   For example, an Apple Silicon Mac uses the `darwin-arm64` archive.
3. Extract the archive into the ignored `.tools/tunnel-client` directory and
   make the client executable:

```bash
mkdir -p .tools/tunnel-client
unzip /path/to/downloaded-tunnel-client.zip -d .tools/tunnel-client
chmod +x .tools/tunnel-client/tunnel-client
.tools/tunnel-client/tunnel-client --version
```

Create an ignored profile that connects the OpenAI-hosted tunnel to this
project's local Streamable HTTP endpoint:

```bash
mkdir -p .tools/tunnel-profiles
.tools/tunnel-client/tunnel-client init \
  --profile cats-vector-store \
  --profile-dir .tools/tunnel-profiles \
  --tunnel-id tunnel_your_tunnel_id \
  --mcp-server-url http://127.0.0.1:8000/mcp
```

This creates `.tools/tunnel-profiles/cats-vector-store.yaml`. Do not commit that
profile because it contains your tunnel ID. Add the same ID and a runtime API
key from [OpenAI Platform API keys](https://platform.openai.com/settings/organization/api-keys)
to `.env.local`:

```dotenv
MCP_TUNNEL_ID=tunnel_your_tunnel_id
CONTROL_PLANE_API_KEY=your_runtime_api_key
```

The tunnel ID identifies the tunnel. `CONTROL_PLANE_API_KEY` authenticates the
locally running tunnel client to OpenAI. Validate the configuration after the
MCP server is running:

```bash
.tools/tunnel-client/tunnel-client doctor \
  --profile-file .tools/tunnel-profiles/cats-vector-store.yaml
```

Run the demonstration from three terminals opened at the project root. Keep the
first two processes running while you start the third.

### Terminal 1: start the MCP server

```bash
mcp-vector-store-assistant
```

This starts the local Streamable HTTP endpoint at
`http://127.0.0.1:8000/mcp`.

### Terminal 2: start the secure tunnel

The Python programs load `.env.local` automatically, but the tunnel-client
binary does not. Export the variables from that file into this terminal first:

```bash
set -a
source .env.local
set +a
```

Then start the tunnel client using your local binary and profile. For example:

```bash
.tools/tunnel-client/tunnel-client run \
  --profile-file .tools/tunnel-profiles/cats-vector-store.yaml
```

The profile tells `tunnel-client` which tunnel and local MCP endpoint to use.
`CONTROL_PLANE_API_KEY` authenticates the running tunnel client to OpenAI.

### Terminal 3: call the Responses API

```bash
python api_client.py
```

`api_client.py` loads `OPENAI_API_KEY` and `MCP_TUNNEL_ID` from `.env.local`.
Its Responses API request declares only the read-only `search` and `fetch` MCP
tools. The request travels through the running tunnel to the local server, which
searches the vector store and returns the selected source for the model's answer.

## 4. Test it from ChatGPT

With the vector-store server and tunnel client still running, create or refresh
a developer-mode ChatGPT plugin associated with the same tunnel and ChatGPT
workspace. Add the plugin to a new conversation and ask ChatGPT to search the
knowledge base, fetch the most relevant document, and cite it in the answer.

This path was verified on September 11, 2026 with FastMCP 3.4.7 and
`tunnel-client` v0.0.14 using the Streamable HTTP endpoint. ChatGPT invoked both
`search` and `fetch`, retrieved content from `cats.pdf`, and produced a sourced
answer. An earlier v0.0.11 test entered a reconnect loop; the successful v0.0.14
retest is documented in
[openai/tunnel-client#41](https://github.com/openai/tunnel-client/issues/41).

## Tests

Run all tests with:

```bash
pytest
```

The tests cover local ranking and fetching, MCP tool discovery, vector-result
deduplication, content assembly, and input validation. OpenAI calls are mocked,
so the suite is repeatable and does not consume API credits.

## Design decisions and scope

- **Read-only first:** neither MCP tool changes files or external state.
- **Stable compatibility contract:** `search(query)` returns document references;
  `fetch(id)` returns complete content and metadata.
- **Document results, not chunk results:** chunks are retrieval evidence inside the
  vector store, while the MCP client receives stable file IDs.
- **MCP is an abstraction layer:** for a single OpenAI-hosted vector store, the
  Responses API's built-in File Search tool is simpler. MCP becomes useful when
  the same retrieval interface must serve multiple clients, hide backend details,
  or later add authorization and domain logic.
- **Verified integration boundary:** the local servers, direct MCP clients,
  Responses API path, and a developer-mode ChatGPT plugin were exercised through
  the secure tunnel. This repository does not claim a deployed public server or
  a published ChatGPT app.

## Security notes

- Never commit `.env.local`, API keys, tunnel runtime keys, or organization IDs.
- Use project-scoped credentials and grant only the permissions required.
- Keep the local MCP server bound behind the secure outbound tunnel rather than
  opening an inbound firewall port.
- Review tool permissions before adding any write or consequential action.

## References

- [OpenAI MCP guide](https://developers.openai.com/api/docs/mcp)
- [OpenAI File Search guide](https://developers.openai.com/api/docs/guides/tools-file-search)
- [OpenAI secure MCP tunnels guide](https://developers.openai.com/api/docs/guides/secure-mcp-tunnels)
- [FastMCP documentation](https://gofastmcp.com/)

TDQS

A3.6/5.0

Scored across 2 tools

Disambiguation5/5

search and fetch have clearly distinct purposes: one finds documents by query, the other retrieves a specific document by ID. There is no overlap or ambiguity.

Naming Consistency5/5

Both tools use simple, consistent single-verb names ('search' and 'fetch') that clearly indicate their actions. The pattern is uniform.

Tool Count3/5

With only 2 tools, the server is minimal but functionally complete for a simple search-and-retrieve pattern. It feels thin but is reasonable for a narrow purpose.

Completeness4/5

The two tools cover the core workflow of searching and retrieving documents. Minor gaps could include listing all documents or getting metadata, but the essential lifecycle is covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues