Skip to main content
Glama

MCP Chat

MCP Chat is a command-line interface application that enables interactive chat capabilities with AI models. It can run against either the hosted Anthropic API or a fully local model (Ollama via a LiteLLM proxy). The application supports document retrieval, command-based prompts, and extensible tool integrations via the MCP (Model Control Protocol) architecture.

Prerequisites

  • Python 3.9+

  • Either an Anthropic API key (hosted) or Ollama installed (local, no key required)

Related MCP server: MCP Chat

Running with a local model (no API key)

The app talks to the Anthropic SDK; a small LiteLLM proxy translates that to a local model, so no code changes are needed to switch backends — only .env. Two local backends are supported (defined in litellm_config.yaml): llama.cpp (default) and Ollama. Both serve the same qwen2.5:14b weights.

Common to both:

  • .env is already set for local use:

    ANTHROPIC_BASE_URL="http://localhost:4000"
    ANTHROPIC_API_KEY="not-needed"
  • Start the LiteLLM proxy in its own terminal and leave it running:

    uv run litellm --config litellm_config.yaml
  • Run the app in another terminal:

    uv run main.py

Memory note: a 14B model at Q4 (~9 GB) runs comfortably on 24 GB+ of RAM. On 16–18 GB it works but is memory-tight — only one 14B can be GPU-resident at a time, so don't run llama.cpp and Ollama simultaneously. For lighter machines use a 7B variant.

Backend A — llama.cpp (default: CLAUDE_MODEL="qwen2.5-llamacpp")

  1. Install: brew install llama.cpp

  2. Get the weights. If you've already pulled the model with Ollama (below), llama.cpp can load that exact GGUF blob — no second download. Otherwise download a qwen2.5-14b-instruct GGUF from Hugging Face.

  3. Start llama-server (leave running). Flash-attention + quantized KV cache are required to fit a 14B on ~18 GB:

    llama-server \
      -m ~/.ollama/models/blobs/sha256-2049f5674b1e92b4464e5729975c9689fcfbf0b0e4443ccf10b5339f370f9a54 \
      --jinja -c 4096 -np 1 -fa on \
      --cache-type-k q8_0 --cache-type-v q8_0 -b 1024 -ub 512 -ngl 999 \
      --host 127.0.0.1 --port 8080 --alias qwen2.5-14b

    --jinja enables tool calling (the app relies on it). The blob path is the GGUF Ollama stored; check yours with ls -lhS ~/.ollama/models/blobs/.

Backend B — Ollama (set CLAUDE_MODEL="qwen2.5-local")

  1. Install and start: brew install ollama && brew services start ollama

  2. Download the model: ollama pull qwen2.5:14b

Ollama runs as a background service (no separate terminal needed) and handles flash-attention/KV settings itself.

Running with the hosted Anthropic API

In .env, comment out ANTHROPIC_BASE_URL, set a real key, and use a real model id:

CLAUDE_MODEL="claude-sonnet-4-5"
ANTHROPIC_API_KEY="sk-ant-..."

Then uv run main.py (no proxy needed).

Setup

Step 1: Configure the environment variables

  1. Create or edit the .env file in the project root and verify that the following variables are set correctly:

ANTHROPIC_API_KEY=""  # Enter your Anthropic API secret key

Step 2: Install dependencies

uv is a fast Python package installer and resolver.

  1. Install uv, if not already installed:

pip install uv
  1. Create and activate a virtual environment:

uv venv
source .venv/bin/activate  # On Windows: .venv\Scripts\activate
  1. Install dependencies:

uv pip install -e .
  1. Run the project

uv run main.py

Option 2: Setup without uv

  1. Create and activate a virtual environment:

python -m venv .venv
source .venv/bin/activate  # On Windows: .venv\Scripts\activate
  1. Install dependencies:

pip install anthropic python-dotenv prompt-toolkit "mcp[cli]==1.8.0"
  1. Run the project

python main.py

Usage

Basic Interaction

Simply type your message and press Enter to chat with the model.

Document Retrieval

Use the @ symbol followed by a document ID to include document content in your query:

> Tell me about @deposition.md

Commands

Use the / prefix to execute commands defined in the MCP server:

> /summarize deposition.md

Commands will auto-complete when you press Tab.

Development

Adding New Documents

Edit the mcp_server.py file to add new documents to the docs dictionary.

Implementing MCP Features

To fully implement the MCP features:

  1. Complete the TODOs in mcp_server.py

  2. Implement the missing functionality in mcp_client.py

Linting and Typing Check

There are no lint or type checks implemented.

Available Tools

2 tools
edit_documentA

Edit a document by replacing a string in the documents content with a new string

ParametersJSON Schema
NameRequiredDescriptionDefault
doc_idYesId of the document that will be edited
new_strYesThe new text to insert in place of the old text
old_strYesThe text to replace. Must match exactly, including whitespace

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavior. It states that a string is replaced, but does not clarify whether all occurrences are replaced or only the first, nor does it mention that the edit is permanent/mutating, any side effects, or error conditions. This ambiguity is a significant gap for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence with no filler or redundant restatement of the tool name. It is front-loaded with the action and resource, and every word contributes meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 3-parameter tool, the description plus full schema coverage is mostly sufficient for an agent to make a call. However, the missing occurrence behavior (all vs. first match) and the absence of an output schema leave some uncertainty about the tool's actual effect and return value, making it minimally complete rather than fully contextualized.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers 100% of the parameters with clear descriptions. The tool description essentially paraphrases the schema (old_str replaced by new_str) without adding new semantic meaning, such as constraints or examples. Baseline 3 is appropriate because the schema carries the burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Edit'), names the exact resource ('a document'), and specifies the mechanism ('replacing a string ... with a new string'). This clearly distinguishes it from the sibling read_doc_contents, which is a read operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: when a document's content must be modified via string replacement. However, it provides no explicit guidance about when NOT to use it or when to prefer the sibling read_doc_contents instead. The alternative is only inferable from the tool name, not the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_doc_contentsC

Read the contents of a doc and return it as a string

ParametersJSON Schema
NameRequiredDescriptionDefault
doc_idYesId of the document to read

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are absent, so the description carries full behavioral burden. It does disclose the return format ('as a string'), which is genuinely useful, but says nothing about permissions/access scoping, error behavior for invalid or inaccessible docs, or size limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler, and the return type is included economically. It is appropriately sized, though it is perhaps too terse to earn a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read tool with no output schema, stating the return type partially compensates. However, with no annotations to cover the safety/behavior profile, the lack of any access or failure context leaves a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter and schema coverage is 100% with a clear 'Id of the document to read' description, so the schema already does the work. The description adds no format or sourcing detail beyond the schema, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read) and resource (doc contents) and clarifies the return form ('as a string'). The read vs. edit contrast with the sibling edit_document is implied by the verb but never named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no mention of prerequisites, and no explicit routing to the sibling edit_document as the alternative for modifications. The agent must infer that this is the read-only counterpart from the verb alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.0
    • First observededit_document
    • First observedread_doc_contents

TDQS

B3.3/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: reading content versus modifying content. An agent can trivially tell them apart with no overlap in intent.

Naming Consistency4/5

Both names follow a verb_noun pattern (read_doc_contents, edit_document), which is largely consistent. Minor deviation: one uses 'doc' and the other 'document', an inconsistency in noun form.

Tool Count3/5

Two tools is thin for a document-oriented server; it covers only reading and in-place editing. It is borderline minimal but defensible for a tightly scoped use case.

Completeness3/5

Core read and edit operations exist, but the surface lacks create, delete, append, or list operations, which are common document lifecycle needs. The string-replacement edit also limits editing to exact-match substitutions.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers