Skip to main content
Glama

ask-llm-mcp

Give your coding agent a cheap LLM to delegate to.

ask-llm-mcp is a minimal MCP server that exposes raw text-in / text-out LLM calls. No agent runtime, no tools, no context injection — you send a prompt, you get text back.

The point: your main agent (Claude / Codex / CodeBuddy / …) is running on an expensive model. Summarising a file, classifying an issue, drafting a commit message, extracting a JSON blob — none of that needs the expensive model. Hand those subtasks to ask_llm and keep the good model for real reasoning.

WARNING

Unofficial project. This speaks an undocumented, reverse-engineered API used by the Devin / Windsurf client. It is not affiliated with, endorsed by, or supported by Cognition, Windsurf or Codeium. It requires your own Devin subscription and uses your own credentials. The protocol can change or stop working at any time. Use at your own risk and make sure your usage complies with the terms of service of the service you subscribe to.


Contents

English | 中文文档 | Install guide


Related MCP server: geminicli-mcp

Why

Running devin -p "<prompt>" spawns a full agent runner: it loads a system prompt, registers tools, and can burn minutes deciding it doesn't want to answer a one-line question. That's the wrong shape for "classify this log line".

ask-llm-mcp calls the backend GetChatMessage endpoint directly over Connect-RPC. One HTTP request, one text response. That's it.

The second half of the design is keeping the answer out of your context window. ask_llm does not return the LLM's text inline — it returns a path. Your agent reads the file only if it actually needs the content, and can ignore it, grep it, or pass it along without paying for the tokens twice.

Features

  • Three tools: ask_llm, list_models, list_sessions.

  • Model choice with pricing — pick a cheap model per call, or discover options at runtime.

  • Multi-turn conversations via session_id.

  • File-based results — result_file (answer) + log_file (timing, errors, prompt preview) under a configurable data directory.

  • No MCP SDK dependency — hand-rolled JSON-RPC over stdio; the only runtime dependency is requests.

  • Two stdio transports — newline-delimited JSON and Content-Length framing, auto-detected per message.

  • Offline test suite — protobuf/framing round-trip tests, no credentials needed.

Requirements

Requirement

Notes

Python

3.11 or newer (uses tomllib)

requests

Installed automatically by pip / the install script

devin CLI

Must be installed and logged in — the server reads its credentials

A Devin subscription

The API is called with your account

The server reads your API key from ~/.local/share/devin/credentials.toml (written by devin auth login / devin login). No key is ever passed through MCP arguments or env vars.

list_models and list_sessions additionally shell out to the devin CLI (devin models list --format json, devin list --format json). ask_llm does not.

Quick start

# 1. get the code
git clone https://github.com/agentming/ask-llm-mcp.git
cd ask-llm-mcp

# 2. install (creates an isolated venv + console script)
./install.sh

# 3. sanity check — should print the three tool definitions
echo '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | ask-llm-mcp

# 4. sanity check with real credentials
python3 devin_api.py "Say hello in one sentence."

./install.sh --help shows the options (--dev, --uninstall, --prefix DIR).

Alternative installs

# pipx — no venv management for you
pipx install git+https://github.com/agentming/ask-llm-mcp.git

# plain pip
pip install git+https://github.com/agentming/ask-llm-mcp.git

# run straight from the checkout, no install at all
python3 /path/to/ask-llm-mcp/server.py

Then register it with your MCP client — see below.

Client configuration

Claude Code / Codex CLI

claude mcp add ask-llm -- /absolute/path/to/ask-llm-mcp
# or, without installing:
claude mcp add ask-llm -- python3 /absolute/path/to/ask-llm-mcp/server.py

Claude Desktop

~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):

{
  "mcpServers": {
    "ask-llm": {
      "command": "/absolute/path/to/ask-llm-mcp",
      "env": {
        "DEVIN_MODEL": "swe-1-7"
      }
    }
  }
}

CodeBuddy Code / other JSON-config clients

.codebuddy/mcp.json or ~/.codebuddy/mcp.json:

{
  "mcpServers": {
    "ask-llm": {
      "command": "python3",
      "args": ["/absolute/path/to/ask-llm-mcp/server.py"]
    }
  }
}

More copy-pasteable variants live in examples/.

Use absolute paths. MCP clients launch the server with an arbitrary working directory.

Tools

ask_llm

Raw text-in / text-out LLM call.

Argument

Type

Default

Description

prompt

string

required

The prompt text to send.

model

string

swe-1-7

Model uid or alias. Ignored when resuming a session. Use list_models to discover values.

session_id

string

—

Resume a multi-turn conversation. Omit to start a new one.

system_prompt

string

You are a helpful assistant.

Optional override.

Returns structured data, not the answer text:

{
  "status": "ok",
  "session_id": "0f9c...",
  "result_file": "/home/you/.local/share/ask-llm/results/20260908_113000_ab12cd34.txt",
  "log_file": "/home/you/.local/share/ask-llm/logs/20260908_113000_ab12cd34.json",
  "error": null,
  "call_id": "20260908_113000_ab12cd34"
}

Read result_file for the answer. If the model emitted reasoning, it is prepended and separated by a --- divider. On failure status is "error", error carries the message, and result_file contains [ERROR] ….

Multi-turn example — just feed the returned session_id back in:

turn 1: ask_llm(prompt="My function returns None. Here it is: ...")
        -> session_id = "0f9c..."
turn 2: ask_llm(prompt="Now show the fix.", session_id="0f9c...")

list_models

Lists model families with per-1M-token pricing, context window and cost tier.

Argument

Type

Description

query

string

Case-insensitive filter against family name, uid, label or cost tier (e.g. swe, cheap, opus).

## SWE (aliases: swe)
  swe-1-7  —  SWE-1.7  [200K ctx, cheap, $0.25/$0.03/$1.00]

list_sessions

Lists recent Devin sessions so you can recover a session_id.

Argument

Type

Default

Description

limit

integer

20

Max sessions to return.

workdir

string

all

Filter by working directory; all or "" disables filtering.

Configuration

All configuration is via environment variables (set them in your MCP client's env block).

Variable

Default

Description

DEVIN_MODEL

swe-1-7

Default model when model is omitted.

DEVIN_TIMEOUT

300

Seconds to wait for the API response.

ASK_LLM_DATA_DIR

~/.local/share/ask-llm

Where results/ and logs/ are written.

Result files are never deleted automatically — clear $ASK_LLM_DATA_DIR/results yourself if it grows.

How it works

MCP client ──stdin (JSON-RPC)──► server.py ──► devin_api.py ──HTTPS──► server.codeium.com
                                     │                                  (Connect-RPC)
                                     └──► $ASK_LLM_DATA_DIR/{results,logs}/
  • devin_api.py hand-rolls the protobuf request (GetChatMessageRequest) and wraps it in a single uncompressed Connect-RPC frame. Responses are streamed back as gzipped frames whose delta_text / delta_thinking fields are concatenated into the final answer.

  • server.py implements just enough of the MCP protocol to advertise and serve the three tools: initialize, tools/list, tools/call. It accepts both newline-delimited JSON and Content-Length framing on stdin.

Because the request field numbers were calibrated against live traffic rather than a published schema, they are the most likely thing to break upstream. If calls start failing with an HTTP error, look there first.

Troubleshooting

Symptom

Cause / fix

Credentials file not found: ~/.local/share/devin/credentials.toml

devin CLI is not installed or you never logged in. Run devin login.

No windsurf_api_key found in credentials.toml

Same file, missing key — re-login.

HTTP 401 / HTTP 403

Expired or revoked key. Re-login with the devin CLI.

HTTP 500: an internal error occurred

The request payload was rejected — most often an unsupported model uid. Run list_models and use an exact uid.

Empty result file

The model returned no text. Check log_file for elapsed_seconds and the prompt preview.

devin models list failed

devin CLI not on PATH inside the MCP client's environment — use an absolute path or extend PATH in the client config.

Tools never appear in the client

Verify the absolute path to server.py, then run the echo … | python3 server.py check above.

Server hangs

Raise DEVIN_TIMEOUT, or the model is simply slow; check the log file for elapsed time.

Logs for every call live in $ASK_LLM_DATA_DIR/logs/*.json and include the timestamp, model, elapsed seconds, prompt length/preview and the error.

Project layout

server.py       MCP JSON-RPC server + the three tool implementations
devin_api.py    Connect-RPC / protobuf client for the backend (no MCP knowledge)
tests/          Offline tests for the wire helpers
examples/       MCP client config snippets
install.sh      Isolated-venv installer

Contributing

See CONTRIBUTING.md. The short version: keep changes small, never add tests that hit the real API, never commit credentials.

License

MIT © agentming

Available Tools

3 tools
ask_llmA

Pure text-in/text-out LLM call. No system prompt, no tools, no context. Use this for cheap LLM calls instead of expensive models. Pass model= to select a specific model (default: swe-1-7). Pass session_id= to resume a multi-turn conversation. Returns structured JSON: session_id, result_file (LLM output), log_file, status. Read the result_file to get the LLM's full text answer.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel uid or alias (default: swe-1-7). Ignored when resuming a session. Use list_models to discover available models.swe-1-7
promptYesThe raw prompt text to send to the LLM.
session_idNoSession ID to resume a multi-turn conversation. Omit to start a new conversation. The returned session_id can be passed back for subsequent turns.
system_promptNoOptional system prompt. Defaults to a minimal helper prompt.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does well: it discloses that there is no system prompt, no tools, no context, that session resumption is possible, and that the response is a structured JSON with session_id, result_file, log_file, and status. The only slight issue is that 'No system prompt' sits uneasily with the optional system_prompt parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core behavior, then moves to usage, parameter hints, and return format. Every sentence contributes useful information, though 'No system prompt' slightly conflicts with the schema and the text-in/text-out idea is repeated in different forms.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with no annotations and no output schema, the description covers the key operational details: model selection, session resumption, return envelope, and how to read the full answer. It does not explain status values or log_file usage, but those are minor for typical calls.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides 100% parameter coverage, so the description does not need to explain parameters in depth. It does restate model as passable with a default and session_id for resuming conversations, matching the schema rather than adding new semantic value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool performs a pure text-in/text-out LLM call and highlights its simplicity and low cost. It distinguishes this from expensive models and from the sibling list_models/list_sessions tools, leaving no ambiguity about the tool's purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says to use this for cheap LLM calls instead of expensive models, giving clear when-to-use guidance. It does not explicitly name alternative tools or exclusion criteria, but the context and sibling names make the intended usage clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_modelsA

List available Devin models with pricing (input/cached/output per 1M tokens), context window, and cost tier. Optionally filter with query=. Use this to discover cheap models before calling ask_llm.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNoOptional case-insensitive filter — matches against family name, model uid, label, or cost tier (e.g. 'swe', 'cheap', 'opus').

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden; it meets it by stating a read-only listing behavior, optional query filtering, and the exact data returned. There are no side-effect expectations, and 'List' clearly signals a safe read operation, though it omits any authorization/availability caveats.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler, front-loading the core purpose before the optional filter and a clear use directive. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-optional-parameter listing tool, the description covers what is returned, the optional filter, and the recommended call sequence with ask_llm. No output schema exists, but the description supplies the relevant return fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully documents 'query' with a thorough description of matching behavior. The tool description adds workflow context but no new parameter semantics, so the baseline of 3 for 100% schema coverage applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Starts with 'List available Devin models', giving a specific verb and resource, and enumerates exact output attributes (pricing input/cached/output, context window, cost tier). The closing guidance 'before calling ask_llm' differentiates it from the sibling inference tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit use context: 'Use this to discover cheap models before calling ask_llm.' It tells an agent when to invoke it, though it does not enumerate cases when it should not be used or mention list_sessions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_sessionsA

List recent Devin sessions (id, title, last activity, working directory). Use this to find session_id values for resuming multi-turn conversations.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax number of sessions to return (default: 20).
workdirNoFilter sessions by working directory. Defaults to all directories. Pass empty string or "all" to list all.all

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It clarifies that the tool is a listing operation and mentions 'recent' sessions plus the returned fields, but it does not describe ordering, default limits beyond the schema, pagination, or empty-result behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences deliver the resource, output fields, and primary use case without wasted words. The key purpose is front-loaded before the usage hint.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete enough for a low-complexity listing tool: it names the output fields and gives a concrete purpose. Slight vagueness around what 'recent' means in terms of ordering or time window prevents a 5, but the schema covers the parameters and defaults.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters, so the description adds no essential parameter semantics beyond what the schema already provides. The baseline 3 is appropriate because the description does not need to repeat schema content.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and a clear resource ('recent Devin sessions'), and enumerates the return fields. It is self-explanatory and easily distinguished from the sibling tools list_models and ask_llm, which operate on different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states the intended use: finding session_id values for resuming multi-turn conversations. It does not discuss exclusions or alternatives, but no alternative listing tool is present, so the context is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.5.0
    • First observedask_llm
    • First observedlist_models
    • First observedlist_sessions

TDQS

A4.3/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: ask_llm executes LLM calls, list_models provides model discovery, and list_sessions handles session lookup. There is no overlap between the action and the two list operations.

Naming Consistency5/5

All tool names follow a consistent verb_first snake_case pattern: ask_llm, list_models, list_sessions. The naming style is uniform and predictable.

Tool Count5/5

Three tools is well-scoped for this server's purpose. Each tool covers a necessary part of the LLM request workflow: choosing a model, making a call, and resuming sessions.

Completeness4/5

The core workflow is covered: discover models, call the LLM, and list/resume sessions. Minor gaps exist such as no dedicated session detail or deletion tool, but the main lifecycle is functional.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers