MCP Knowledge Assistant
This server is a read-only MCP knowledge assistant that lets you search a knowledge base and retrieve full source documents.
Search documents: Use the
searchtool with a natural-language query to get compact document references (id, title, url).Fetch complete documents: Use the
fetchtool with an ID returned bysearchto retrieve the full text, title, url, and optional metadata.Query local or vector-store backends: It supports a local JSON knowledge base over stdio and semantic retrieval from an OpenAI vector store over Streamable HTTP.
Get deduplicated results: Vector-store search returns one document-level result per file, not multiple chunk matches.
Integrate with MCP clients: Callable from ChatGPT, the OpenAI Responses API through a secure tunnel, or direct MCP demo clients.
Operate safely and offline-testable: All operations are read-only, and the test suite runs without paid API calls.
Provides semantic search and complete document retrieval against an OpenAI vector store, with tools to find relevant documents and fetch their full source content.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Knowledge Assistantsearch for content on MCP tool design"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Knowledge Assistant
A read-only Model Context Protocol (MCP) server that lets an AI client search a knowledge base and retrieve complete source documents. The project starts with a small local prototype, then applies the same tool contract to semantic retrieval from an OpenAI vector store.
What this demonstrates
MCP tool design with narrow
searchandfetchresponsibilitiesStructured inputs and outputs using FastMCP and Pydantic
Semantic retrieval over uploaded documents in an OpenAI vector store
Document-level deduplication when vector search returns several matching chunks
Streamable HTTP and
stdioMCP transportsEnd-to-end tool use through the OpenAI Responses API and a secure MCP tunnel
Configuration through environment variables, with no credentials in source code
Unit and protocol-level tests that do not make paid API calls
Related MCP server: File AI
Architecture
The following sequence shows the verified ChatGPT path. The tunnel client makes an outbound connection, so the local MCP server does not need a publicly exposed inbound port.
sequenceDiagram
actor User
participant ChatGPT
participant Plugin as ChatGPT plugin
participant Tunnel as OpenAI secure tunnel
participant MCP as Local FastMCP server
participant Store as OpenAI vector store
User->>ChatGPT: Ask a question about the knowledge base
ChatGPT->>Plugin: Call search(query)
Plugin->>Tunnel: Send MCP tool request
Tunnel->>MCP: Forward request over Streamable HTTP
MCP->>Store: Run semantic search
Store-->>MCP: Return matching chunks and file IDs
MCP-->>Tunnel: Return deduplicated document results
Tunnel-->>Plugin: Relay search results
Plugin-->>ChatGPT: Provide document references
ChatGPT->>Plugin: Call fetch(id)
Plugin->>Tunnel: Send MCP tool request
Tunnel->>MCP: Forward request over Streamable HTTP
MCP->>Store: Retrieve the selected file content
Store-->>MCP: Return parsed document content
MCP-->>Tunnel: Return document and metadata
Tunnel-->>Plugin: Relay fetched document
Plugin-->>ChatGPT: Provide source content
ChatGPT-->>User: Compose a grounded answer with a citationapi_client.py follows the same tunnel-to-server path, with the OpenAI
Responses API acting as the MCP client instead of the ChatGPT plugin.
The repository also includes a fully local learning path:
Local demo client -> FastMCP server (stdio) -> data/documents.jsonBoth servers expose the same public tool contract:
Tool | Input | Purpose |
|
| Return compact, relevant document references. |
|
| Retrieve one complete document selected from search. |
Keeping discovery separate from retrieval avoids sending full documents before they are needed and gives the model stable document IDs to use in later calls.
Project structure
.
├── data/documents.json # Sample local knowledge base
├── sample_data/cats.pdf # Public-domain vector-store sample
├── src/mcp_knowledge_assistant/
│ ├── knowledge_base.py # Local keyword retrieval
│ ├── models.py # Shared response schemas
│ ├── server.py # Local stdio MCP server
│ └── vector_store_server.py # OpenAI vector-store MCP server
├── tests/ # Offline unit and MCP tests
├── demo_client.py # Local stdio demonstration
├── vector_store_demo_client.py # Direct HTTP MCP demonstration
└── api_client.py # Responses API + secure tunnel demonstrationRequirements
Python 3.11 or newer
An OpenAI API project with billing enabled for the vector-store path
The included sample PDF, or your own document, uploaded to an OpenAI vector store
The OpenAI tunnel client only for the secure-tunnel demonstration
The local JSON server and the complete test suite require no API key.
Sample document attribution
The vector-store demonstration uses Cats: Their Points and Characteristics by W. Gordon Stables, Project Gutenberg eBook #43429. The sample PDF is hosted by OpenAI and was produced from the Project Gutenberg edition. See the PDF for the Project Gutenberg licence and applicable reuse terms.
Setup
Clone the repository, create a virtual environment, and install the project:
python -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[dev]"For OpenAI-backed examples, copy the environment template:
cp .env.example .env.localThen add your own values to .env.local:
OPENAI_API_KEY=your_project_api_key
VECTOR_STORE_ID=vs_your_vector_store_id.env.local, PyCharm settings, virtual environments, and local tunnel profiles
are excluded from Git.
1. Run the local prototype
The first server uses stdio, so the MCP client launches it as a child process
and communicates over standard input and output:
python demo_client.pyThe demo discovers both tools, searches the sample JSON knowledge base, and fetches the selected document.
You can also start the server through its installed command:
mcp-knowledge-assistantA silent, waiting process is normal for a stdio server that has no connected
client.
2. Run the vector-store server
Upload sample_data/cats.pdf to an OpenAI vector store, then set
OPENAI_API_KEY and VECTOR_STORE_ID in .env.local. You may substitute your
own document and query if preferred. Start the Streamable HTTP server:
mcp-vector-store-assistantBy default, its MCP endpoint is:
http://127.0.0.1:8000/mcpIn a second terminal, test the endpoint directly:
python vector_store_demo_client.pyVector search operates on chunks, so a long document may produce several
matches with the same file ID. The MCP search tool deliberately collapses
those matches into one document result. The fetch tool then retrieves and
combines that document's parsed content for the model.
3. Call it through the Responses API
Download and configure tunnel-client
Create a tunnel in OpenAI Platform tunnel settings and copy its
tunnel_id. Associate it with the Platform organization used by the API request and, for section 4, the intended ChatGPT workspace.Download the archive for your operating system and CPU from the latest official tunnel-client release. For example, an Apple Silicon Mac uses the
darwin-arm64archive.Extract the archive into the ignored
.tools/tunnel-clientdirectory and make the client executable:
mkdir -p .tools/tunnel-client
unzip /path/to/downloaded-tunnel-client.zip -d .tools/tunnel-client
chmod +x .tools/tunnel-client/tunnel-client
.tools/tunnel-client/tunnel-client --versionCreate an ignored profile that connects the OpenAI-hosted tunnel to this project's local Streamable HTTP endpoint:
mkdir -p .tools/tunnel-profiles
.tools/tunnel-client/tunnel-client init \
--profile cats-vector-store \
--profile-dir .tools/tunnel-profiles \
--tunnel-id tunnel_your_tunnel_id \
--mcp-server-url http://127.0.0.1:8000/mcpThis creates .tools/tunnel-profiles/cats-vector-store.yaml. Do not commit that
profile because it contains your tunnel ID. Add the same ID and a runtime API
key from OpenAI Platform API keys
to .env.local:
MCP_TUNNEL_ID=tunnel_your_tunnel_id
CONTROL_PLANE_API_KEY=your_runtime_api_keyThe tunnel ID identifies the tunnel. CONTROL_PLANE_API_KEY authenticates the
locally running tunnel client to OpenAI. Validate the configuration after the
MCP server is running:
.tools/tunnel-client/tunnel-client doctor \
--profile-file .tools/tunnel-profiles/cats-vector-store.yamlRun the demonstration from three terminals opened at the project root. Keep the first two processes running while you start the third.
Terminal 1: start the MCP server
mcp-vector-store-assistantThis starts the local Streamable HTTP endpoint at
http://127.0.0.1:8000/mcp.
Terminal 2: start the secure tunnel
The Python programs load .env.local automatically, but the tunnel-client
binary does not. Export the variables from that file into this terminal first:
set -a
source .env.local
set +aThen start the tunnel client using your local binary and profile. For example:
.tools/tunnel-client/tunnel-client run \
--profile-file .tools/tunnel-profiles/cats-vector-store.yamlThe profile tells tunnel-client which tunnel and local MCP endpoint to use.
CONTROL_PLANE_API_KEY authenticates the running tunnel client to OpenAI.
Terminal 3: call the Responses API
python api_client.pyapi_client.py loads OPENAI_API_KEY and MCP_TUNNEL_ID from .env.local.
Its Responses API request declares only the read-only search and fetch MCP
tools. The request travels through the running tunnel to the local server, which
searches the vector store and returns the selected source for the model's answer.
4. Test it from ChatGPT
With the vector-store server and tunnel client still running, create or refresh a developer-mode ChatGPT plugin associated with the same tunnel and ChatGPT workspace. Add the plugin to a new conversation and ask ChatGPT to search the knowledge base, fetch the most relevant document, and cite it in the answer.
This path was verified on September 11, 2026 with FastMCP 3.4.7 and
tunnel-client v0.0.14 using the Streamable HTTP endpoint. ChatGPT invoked both
search and fetch, retrieved content from cats.pdf, and produced a sourced
answer. An earlier v0.0.11 test entered a reconnect loop; the successful v0.0.14
retest is documented in
openai/tunnel-client#41.
Tests
Run all tests with:
pytestThe tests cover local ranking and fetching, MCP tool discovery, vector-result deduplication, content assembly, and input validation. OpenAI calls are mocked, so the suite is repeatable and does not consume API credits.
Design decisions and scope
Read-only first: neither MCP tool changes files or external state.
Stable compatibility contract:
search(query)returns document references;fetch(id)returns complete content and metadata.Document results, not chunk results: chunks are retrieval evidence inside the vector store, while the MCP client receives stable file IDs.
MCP is an abstraction layer: for a single OpenAI-hosted vector store, the Responses API's built-in File Search tool is simpler. MCP becomes useful when the same retrieval interface must serve multiple clients, hide backend details, or later add authorization and domain logic.
Verified integration boundary: the local servers, direct MCP clients, Responses API path, and a developer-mode ChatGPT plugin were exercised through the secure tunnel. This repository does not claim a deployed public server or a published ChatGPT app.
Security notes
Never commit
.env.local, API keys, tunnel runtime keys, or organization IDs.Use project-scoped credentials and grant only the permissions required.
Keep the local MCP server bound behind the secure outbound tunnel rather than opening an inbound firewall port.
Review tool permissions before adding any write or consequential action.
References
Available Tools
2 toolsfetchA
Retrieve one complete document using an ID returned by search.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| url | Yes | |
| text | Yes | |
| title | Yes | |
| metadata | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It reveals that this returns a complete document (not a summary) and that it depends on search-generated IDs, but doesn't mention error handling, permissions, or potential side effects. For a simple read operation this is adequate but has gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, no redundant details, and the key information (retrieve by ID) is front-loaded. Perfect conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema is present and there is only one parameter, the description sufficiently covers the essential context and its relationship to the sibling search tool. It could be improved by mentioning error cases or what happens if the ID is invalid, but for this simplicity it is complete enough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'id' has no schema description, so the description's phrase 'returned by search' provides crucial semantic context, explaining where the ID comes from and how it should be used to retrieve the document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Retrieve' with the resource 'one complete document' and the qualifier 'using an ID returned by search', which clearly defines the tool's scope and distinguishes it from the sibling tool 'search'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states that the ID must come from search, establishing a clear workflow (search first, then fetch). It doesn't explicitly mention alternatives, but the context is clear enough for an agent to know when to use this tool versus search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchC
Find documents relevant to a natural-language query.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| results | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the tool does semantic matching over a natural-language query, but says nothing about how results are ranked, paginated, limited, scoped, or whether interactions are read-only or have side effects. It is not misleading, but it adds very little beyond what the name and schema already imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with zero wasted words. It is appropriately brief for a minimal tool definition, though one could argue it is slightly under-specified—but the economy of expression is commendable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 1 parameter, no annotations, 0% schema coverage, and an unaddressed sibling 'fetch', one sentence is insufficient context. The description would be more complete by mentioning result limits, sort order, or any behavioral differences from fetch. The presence of an output schema helps explain return values, but the in/out behavior is still too thin for an agent to fully understand the tool's scope.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does add value by clarifying that 'query' accepts natural-language text rather than structured filters or keywords. However, it provides no examples, formatting hints, length limits, or syntax conventions—the description does the bare minimum to make the parameter actionable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
'Find documents relevant to a natural-language query' uses a clear verb ('find'), a resource ('documents'), and a scope qualifier ('natural-language query'). It is clear and specific, but it does not explicitly distinguish itself from its sibling 'fetch', so it earns a 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use search versus its sibling 'fetch', no exclusions, and no when-to-use context. The description implies usage through the term 'search', but a sibling tool with zero differentiator guidance means the agent is left without decision support. This is 'no guidance' and scores a 2.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
fetch - First observed
search
TDQS
Scored across 2 tools
search and fetch have clearly distinct purposes: one finds documents by query, the other retrieves a specific document by ID. There is no overlap or ambiguity.
Both tools use simple, consistent single-verb names ('search' and 'fetch') that clearly indicate their actions. The pattern is uniform.
With only 2 tools, the server is minimal but functionally complete for a simple search-and-retrieve pattern. It feels thin but is reasonable for a narrow purpose.
The two tools cover the core workflow of searching and retrieving documents. Minor gaps could include listing all documents or getting metadata, but the essential lifecycle is covered.
Maintenance
Related MCP Connectors
Read-only MCP server exposing a user ORANO library to their own AI agent.
Knowledge base MCP for AI agents on iknow.dev. Search, read, and maintain via OAuth.
Knowledge base for AI agents: search, read, tasks; write, correct, report outcomes via /mcp/write.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceAn MCP server that provides tools for retrieving and processing documentation through vector search, enabling AI assistants to augment their responses with relevant documentation context.32 npmMIT
- AlicenseBqualityDmaintenanceA read-only MCP server that provides document awareness for agents by parsing local files into structured profiles, blocks, chunks, and search results, enabling agents to understand and cite document content without dealing with raw file formats.538 npm3Apache 2.0
- FlicenseNot gradedqualityBmaintenanceAn MCP server that connects the Casio Plus knowledge base (playbooks, architecture, learning resources) to AI clients, offering read-only search and validation tools along with controlled feedback intake and review workflows.-
- AlicenseNot gradedqualityAmaintenanceMCP server that enables AI agents to search, fetch, and analyze a self-maintaining markdown knowledge base with provenance, drift detection, and canonical definitions.1MIT