ask-llm-mcp
Provides tools to make raw text-in/text-out LLM calls using the Codeium (Windsurf) API, with model selection and multi-turn session support.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ask-llm-mcpSummarize the contents of README.md in three bullet points."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ask-llm-mcp
Give your coding agent a cheap LLM to delegate to.
ask-llm-mcp is a minimal MCP server that
exposes raw text-in / text-out LLM calls. No agent runtime, no tools, no
context injection — you send a prompt, you get text back.
The point: your main agent (Claude / Codex / CodeBuddy / …) is running on an
expensive model. Summarising a file, classifying an issue, drafting a commit
message, extracting a JSON blob — none of that needs the expensive model.
Hand those subtasks to ask_llm and keep the good model for real reasoning.
Unofficial project. This speaks an undocumented, reverse-engineered API used by the Devin / Windsurf client. It is not affiliated with, endorsed by, or supported by Cognition, Windsurf or Codeium. It requires your own Devin subscription and uses your own credentials. The protocol can change or stop working at any time. Use at your own risk and make sure your usage complies with the terms of service of the service you subscribe to.
Contents
English | 中文文档 | Install guide
Related MCP server: geminicli-mcp
Why
Running devin -p "<prompt>" spawns a full agent runner: it loads a system
prompt, registers tools, and can burn minutes deciding it doesn't want to
answer a one-line question. That's the wrong shape for "classify this log line".
ask-llm-mcp calls the backend GetChatMessage endpoint directly over
Connect-RPC. One HTTP request, one text response. That's it.
The second half of the design is keeping the answer out of your context
window. ask_llm does not return the LLM's text inline — it returns a path.
Your agent reads the file only if it actually needs the content, and can ignore
it, grep it, or pass it along without paying for the tokens twice.
Features
Three tools:
ask_llm,list_models,list_sessions.Model choice with pricing — pick a cheap model per call, or discover options at runtime.
Multi-turn conversations via
session_id.File-based results —
result_file(answer) +log_file(timing, errors, prompt preview) under a configurable data directory.No MCP SDK dependency — hand-rolled JSON-RPC over stdio; the only runtime dependency is
requests.Two stdio transports — newline-delimited JSON and
Content-Lengthframing, auto-detected per message.Offline test suite — protobuf/framing round-trip tests, no credentials needed.
Requirements
Requirement | Notes |
Python | 3.11 or newer (uses |
| Installed automatically by pip / the install script |
| Must be installed and logged in — the server reads its credentials |
A Devin subscription | The API is called with your account |
The server reads your API key from
~/.local/share/devin/credentials.toml (written by devin auth login /
devin login). No key is ever passed through MCP arguments or env vars.
list_models and list_sessions additionally shell out to the devin CLI
(devin models list --format json, devin list --format json). ask_llm
does not.
Quick start
# 1. get the code
git clone https://github.com/agentming/ask-llm-mcp.git
cd ask-llm-mcp
# 2. install (creates an isolated venv + console script)
./install.sh
# 3. sanity check — should print the three tool definitions
echo '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | ask-llm-mcp
# 4. sanity check with real credentials
python3 devin_api.py "Say hello in one sentence."./install.sh --help shows the options (--dev, --uninstall,
--prefix DIR).
Alternative installs
# pipx — no venv management for you
pipx install git+https://github.com/agentming/ask-llm-mcp.git
# plain pip
pip install git+https://github.com/agentming/ask-llm-mcp.git
# run straight from the checkout, no install at all
python3 /path/to/ask-llm-mcp/server.pyThen register it with your MCP client — see below.
Client configuration
Claude Code / Codex CLI
claude mcp add ask-llm -- /absolute/path/to/ask-llm-mcp
# or, without installing:
claude mcp add ask-llm -- python3 /absolute/path/to/ask-llm-mcp/server.pyClaude Desktop
~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or
%APPDATA%\Claude\claude_desktop_config.json (Windows):
{
"mcpServers": {
"ask-llm": {
"command": "/absolute/path/to/ask-llm-mcp",
"env": {
"DEVIN_MODEL": "swe-1-7"
}
}
}
}CodeBuddy Code / other JSON-config clients
.codebuddy/mcp.json or ~/.codebuddy/mcp.json:
{
"mcpServers": {
"ask-llm": {
"command": "python3",
"args": ["/absolute/path/to/ask-llm-mcp/server.py"]
}
}
}More copy-pasteable variants live in examples/.
Use absolute paths. MCP clients launch the server with an arbitrary working directory.
Tools
ask_llm
Raw text-in / text-out LLM call.
Argument | Type | Default | Description |
| string | required | The prompt text to send. |
| string |
| Model uid or alias. Ignored when resuming a session. Use |
| string | — | Resume a multi-turn conversation. Omit to start a new one. |
| string |
| Optional override. |
Returns structured data, not the answer text:
{
"status": "ok",
"session_id": "0f9c...",
"result_file": "/home/you/.local/share/ask-llm/results/20260908_113000_ab12cd34.txt",
"log_file": "/home/you/.local/share/ask-llm/logs/20260908_113000_ab12cd34.json",
"error": null,
"call_id": "20260908_113000_ab12cd34"
}Read result_file for the answer. If the model emitted reasoning, it is
prepended and separated by a --- divider. On failure status is "error",
error carries the message, and result_file contains [ERROR] ….
Multi-turn example — just feed the returned session_id back in:
turn 1: ask_llm(prompt="My function returns None. Here it is: ...")
-> session_id = "0f9c..."
turn 2: ask_llm(prompt="Now show the fix.", session_id="0f9c...")list_models
Lists model families with per-1M-token pricing, context window and cost tier.
Argument | Type | Description |
| string | Case-insensitive filter against family name, uid, label or cost tier (e.g. |
## SWE (aliases: swe)
swe-1-7 — SWE-1.7 [200K ctx, cheap, $0.25/$0.03/$1.00]list_sessions
Lists recent Devin sessions so you can recover a session_id.
Argument | Type | Default | Description |
| integer |
| Max sessions to return. |
| string |
| Filter by working directory; |
Configuration
All configuration is via environment variables (set them in your MCP client's
env block).
Variable | Default | Description |
|
| Default model when |
|
| Seconds to wait for the API response. |
|
| Where |
Result files are never deleted automatically — clear
$ASK_LLM_DATA_DIR/results yourself if it grows.
How it works
MCP client ──stdin (JSON-RPC)──► server.py ──► devin_api.py ──HTTPS──► server.codeium.com
│ (Connect-RPC)
└──► $ASK_LLM_DATA_DIR/{results,logs}/devin_api.pyhand-rolls the protobuf request (GetChatMessageRequest) and wraps it in a single uncompressed Connect-RPC frame. Responses are streamed back as gzipped frames whosedelta_text/delta_thinkingfields are concatenated into the final answer.server.pyimplements just enough of the MCP protocol to advertise and serve the three tools:initialize,tools/list,tools/call. It accepts both newline-delimited JSON andContent-Lengthframing on stdin.
Because the request field numbers were calibrated against live traffic rather than a published schema, they are the most likely thing to break upstream. If calls start failing with an HTTP error, look there first.
Troubleshooting
Symptom | Cause / fix |
|
|
| Same file, missing key — re-login. |
| Expired or revoked key. Re-login with the |
| The request payload was rejected — most often an unsupported |
Empty result file | The model returned no text. Check |
|
|
Tools never appear in the client | Verify the absolute path to |
Server hangs | Raise |
Logs for every call live in $ASK_LLM_DATA_DIR/logs/*.json and include the
timestamp, model, elapsed seconds, prompt length/preview and the error.
Project layout
server.py MCP JSON-RPC server + the three tool implementations
devin_api.py Connect-RPC / protobuf client for the backend (no MCP knowledge)
tests/ Offline tests for the wire helpers
examples/ MCP client config snippets
install.sh Isolated-venv installerContributing
See CONTRIBUTING.md. The short version: keep changes small, never add tests that hit the real API, never commit credentials.
License
MIT © agentming
Available Tools
3 toolsask_llmA
Pure text-in/text-out LLM call. No system prompt, no tools, no context. Use this for cheap LLM calls instead of expensive models. Pass model= to select a specific model (default: swe-1-7). Pass session_id= to resume a multi-turn conversation. Returns structured JSON: session_id, result_file (LLM output), log_file, status. Read the result_file to get the LLM's full text answer.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model uid or alias (default: swe-1-7). Ignored when resuming a session. Use list_models to discover available models. | swe-1-7 |
| prompt | Yes | The raw prompt text to send to the LLM. | |
| session_id | No | Session ID to resume a multi-turn conversation. Omit to start a new conversation. The returned session_id can be passed back for subsequent turns. | |
| system_prompt | No | Optional system prompt. Defaults to a minimal helper prompt. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does well: it discloses that there is no system prompt, no tools, no context, that session resumption is possible, and that the response is a structured JSON with session_id, result_file, log_file, and status. The only slight issue is that 'No system prompt' sits uneasily with the optional system_prompt parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core behavior, then moves to usage, parameter hints, and return format. Every sentence contributes useful information, though 'No system prompt' slightly conflicts with the schema and the text-in/text-out idea is repeated in different forms.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with no annotations and no output schema, the description covers the key operational details: model selection, session resumption, return envelope, and how to read the full answer. It does not explain status values or log_file usage, but those are minor for typical calls.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides 100% parameter coverage, so the description does not need to explain parameters in depth. It does restate model as passable with a default and session_id for resuming conversations, matching the schema rather than adding new semantic value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs a pure text-in/text-out LLM call and highlights its simplicity and low cost. It distinguishes this from expensive models and from the sibling list_models/list_sessions tools, leaving no ambiguity about the tool's purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to use this for cheap LLM calls instead of expensive models, giving clear when-to-use guidance. It does not explicitly name alternative tools or exclusion criteria, but the context and sibling names make the intended usage clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsA
List available Devin models with pricing (input/cached/output per 1M tokens), context window, and cost tier. Optionally filter with query=. Use this to discover cheap models before calling ask_llm.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Optional case-insensitive filter — matches against family name, model uid, label, or cost tier (e.g. 'swe', 'cheap', 'opus'). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden; it meets it by stating a read-only listing behavior, optional query filtering, and the exact data returned. There are no side-effect expectations, and 'List' clearly signals a safe read operation, though it omits any authorization/availability caveats.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler, front-loading the core purpose before the optional filter and a clear use directive. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-optional-parameter listing tool, the description covers what is returned, the optional filter, and the recommended call sequence with ask_llm. No output schema exists, but the description supplies the relevant return fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already fully documents 'query' with a thorough description of matching behavior. The tool description adds workflow context but no new parameter semantics, so the baseline of 3 for 100% schema coverage applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Starts with 'List available Devin models', giving a specific verb and resource, and enumerates exact output attributes (pricing input/cached/output, context window, cost tier). The closing guidance 'before calling ask_llm' differentiates it from the sibling inference tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use context: 'Use this to discover cheap models before calling ask_llm.' It tells an agent when to invoke it, though it does not enumerate cases when it should not be used or mention list_sessions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_sessionsA
List recent Devin sessions (id, title, last activity, working directory). Use this to find session_id values for resuming multi-turn conversations.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max number of sessions to return (default: 20). | |
| workdir | No | Filter sessions by working directory. Defaults to all directories. Pass empty string or "all" to list all. | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It clarifies that the tool is a listing operation and mentions 'recent' sessions plus the returned fields, but it does not describe ordering, default limits beyond the schema, pagination, or empty-result behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences deliver the resource, output fields, and primary use case without wasted words. The key purpose is front-loaded before the usage hint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is complete enough for a low-complexity listing tool: it names the output fields and gives a concrete purpose. Slight vagueness around what 'recent' means in terms of ordering or time window prevents a 5, but the schema covers the parameters and defaults.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters, so the description adds no essential parameter semantics beyond what the schema already provides. The baseline 3 is appropriate because the description does not need to repeat schema content.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and a clear resource ('recent Devin sessions'), and enumerates the return fields. It is self-explanatory and easily distinguished from the sibling tools list_models and ask_llm, which operate on different resources.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states the intended use: finding session_id values for resuming multi-turn conversations. It does not discuss exclusions or alternatives, but no alternative listing tool is present, so the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.5.0- First observed
ask_llm - First observed
list_models - First observed
list_sessions
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: ask_llm executes LLM calls, list_models provides model discovery, and list_sessions handles session lookup. There is no overlap between the action and the two list operations.
All tool names follow a consistent verb_first snake_case pattern: ask_llm, list_models, list_sessions. The naming style is uniform and predictable.
Three tools is well-scoped for this server's purpose. Each tool covers a necessary part of the LLM request workflow: choosing a model, making a call, and resuming sessions.
The core workflow is covered: discover models, call the LLM, and list/resume sessions. Minor gaps exist such as no dedicated session detail or deletion tool, but the main lifecycle is functional.
Maintenance
Related MCP Connectors
MCP server for generating rough-draft project plans from natural-language prompts.
MCP server for AI dialogue using various LLM models via AceDataCloud
MCP-Native LLM Orchestration Agent
Generate contextual prompts and reusable agent skills, evaluate prompts with the 16-dimension Prompt Score, and manage saved work in PromptDrive. Twelve MCP tools also provide authorized access to private Memory for source-grounded answers. Connect over Streamable HTTP using OAuth 2.1 and PKCE. Generation consumes account quota and automatically saves successful results; Memory access follows account permissions and plan limits.
Related MCP Servers
- AlicenseBqualityBmaintenanceMCP server that exposes any LynxPrompt instance to LLMs, enabling browsing, searching, and managing AI configuration blueprints and prompt hierarchies.628 npm2GPL 3.0
- AlicenseAqualityBmaintenanceA stateless MCP server that wraps the headless Gemini CLI, providing tools to send prompts to Google's Gemini models and receive text responses. It supports both simple prompts and prompts with contextual information.2MIT
- FlicenseCqualityDmaintenanceEnables MCP clients to interact with local LLMs via LM Studio, supporting dynamic chat, vision, RAG, file interaction, and model orchestration.28-
- FlicenseNot gradedqualityCmaintenanceEnables developers to explain error messages, validate and format JSON, generate regex patterns from descriptions, and summarize text via an LLM, all through MCP-connected clients.-