mcp-mistral-queue
This server provides a rate-limited, queue-based interface to the Mistral AI API, enabling safe multi-process and multi-client access to the Mistral free tier. It uses a shared SQLite database (WAL mode, stored in a secure per-user directory, path overridable) for coordination.
Key capabilities:
Single-shot prompts: Submit a plain text prompt for a text response.
Multi-turn conversations: Provide a full
messagesarray for context-aware interactions.Flexible model selection: Choose any Mistral chat model (default:
mistral-small-latest).Custom system prompt: Set a system prompt when using the single-prompt mode.
Priority queueing: Assign priority (1=high, 2=normal, 3=low) to influence processing order.
Automatic rate-limit handling: Enforces a ~31‑second start interval between requests and backs off on 429 errors, coordinating across processes.
Internal streaming & cancellation: Streams responses internally and handles client cancellation gracefully, updating the database on cancel.
MCP integration: Exposes an
ask_mistraltool for MCP-compatible clients (e.g., Claude Desktop, Vibe, Goose, OpenCode) and includes diagnostic commands (get_queue_status,clear_queue).CLI usage: Offers a direct command-line interface for ad‑hoc requests.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-mistral-queueExplain Python list comprehensions briefly"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-mistral-queue
An MCP (Model Context Protocol) server and CLI tool that coordinates local and multi-process / multi-client calls to the Mistral free tier (~1 request / 30 seconds) via a shared SQLite queue. It uses SQLite (WAL mode) and async queueing with a single in-flight task to space request starts. This is best-effort traffic control, not an official SLA.
Package: mcp-mistral-queue on PyPI · console script: mmq (not the package name) · current release: 0.1.2
Features
Automatic rate-limit coordination: Shared ~31s start interval; on 429, shared backoff then re-enter the gate. Resets to the base interval on success.
Multi-process & priority control: Multiple processes/tasks can enqueue work. Priority (1–3) plus single in-flight processing order the queue.
Flexible model & message options: Any Mistral chat model name (defaults to
mistral-small-latest; e.g.mistral-large-latest,codestral-latest), plus full conversation history via amessagesarray.Streaming & cancel handling: Streams the Mistral API response internally (tool returns the full text); on client cancel (
CancelledError) updates task status in the DB.Local control DB: Temp DB under a per-user directory with mode
0700(path overridable viaMMQ_TEMP_DB_PATH).PyPI / uvx: Install once or run ephemerally; entry point is
mmq.Mistral Vibe / Grok / Claude Desktop: Register as an MCP server (
mmq --mcp). Do not usevibe mmq.py "..."— that runs Vibe’s agent CLI, not this tool.Good free-tier fit: Occasional jobs (e.g. translating docs) that can wait ~31s between calls without burning a dedicated rate-limit stack.
AI-friendly CLI: Built for coding agents (Vibe, Claude Code, etc.) with
docs list/docs showsubcommands, agent guidance in help text, and JSON outputs for easy parsing.Stdin pipe support: Pipe
git diff --stagedoutput directly intommqto generate commit messages.
Related MCP server: claude-sync
Prerequisites
Python 3.10+
uv recommended (
uvx/uv run);pipalso worksA Mistral API key (
MISTRAL_API_KEY)
export MISTRAL_API_KEY="your-mistral-api-key"Install (PyPI)
Published and verified on PyPI.
# One-shot (no permanent install) — recommended for MCP hosts
uvx --from mcp-mistral-queue mmq --help
# Or install into an environment
uv pip install mcp-mistral-queue
# pip install mcp-mistral-queue
mmq --helpQuick smoke (needs MISTRAL_API_KEY; counts against free-tier quota):
uvx --from mcp-mistral-queue mmq "Reply with pong only."Notes:
Console script name is
mmq. Wrong:uvx mcp-mistral-queue --mcp. Right:uvx --from mcp-mistral-queue mmq --mcp.Dependencies:
mcp[cli]>=1.0.0,<2,mistralai>=1.0.0,<2(pulled in by the package).
Usage
1. CLI mode
After PyPI install / via uvx, invoke mmq.
From a git checkout you can still use uv run mmq.py ... (PEP 723).
# Basic run (default model: mistral-small-latest)
uvx --from mcp-mistral-queue mmq "Explain Python list comprehensions briefly"
# or: mmq "Explain Python list comprehensions briefly"
# Choose a model (e.g. mistral-large-latest, codestral-latest)
mmq -m mistral-large-latest "Explain a complex algorithm"
# Custom system prompt
mmq -s "You are an AI that speaks casually." "How is the weather today?"
# Priority (1: high, 2: normal, 3: low)
mmq --priority 1 "Urgent question"
# Full conversation context as a messages JSON array
# (specify either prompt or --messages, not both)
mmq --messages '[{"role":"system","content":"Strict programmer"},{"role":"user","content":"What is ownership in Rust?"}]'
# Emergency brake: cancel queued / stuck work (no API call)
mmq --purge # cancel all pending
mmq --purge-all # cancel pending + processing
mmq --purge-id 42 # cancel one task by ID
# New structured purge subcommand (recommended for scripts/AI)
mmq purge --pending # cancel all pending tasks
mmq purge --all # cancel all pending + processing tasks
mmq purge --id 42 # cancel specific task by ID
# Pipe stdin to generate a commit message from staged changes
git diff --staged | mmq
git diff --staged | mmq -
git diff --staged | mmq --stdin
git diff --staged | mmq -s "Generate a concise commit message"AI-Friendly Documentation Commands
For coding agents (Vibe, Claude Code, etc.):
# List all available documentation
mmq docs list
# Show specific documentation (returns markdown content)
mmq docs show usage
mmq docs show install
mmq docs show mcp
mmq docs show rate-limit
mmq docs show troubleshooting
mmq docs show examplesThe docs list command outputs JSON with descriptions for easy parsing:
{
"results": [
{"name": "usage", "description": "Usage guide and examples for mcp-mistral-queue CLI"},
{"name": "install", "description": "Installation instructions for mcp-mistral-queue"}
],
"help": "If you are a coding agent, run `mmq docs show {name}` to see details."
}2. MCP server mode (Vibe / Grok / Claude Desktop / …)
Expose ask_mistral and get_queue_status to MCP hosts.
Separate path from CLI prompts.
PyPI / uvx (recommended)
{
"mcpServers": {
"mistral-queue": {
"command": "uvx",
"args": ["--from", "mcp-mistral-queue", "mmq", "--mcp"],
"env": {
"MISTRAL_API_KEY": "your-mistral-api-key"
}
}
}
}If mmq is already on PATH (venv / uv pip install):
{
"mcpServers": {
"mistral-queue": {
"command": "mmq",
"args": ["--mcp"],
"env": {
"MISTRAL_API_KEY": "your-mistral-api-key"
}
}
}
}Local checkout (development)
{
"mcpServers": {
"mistral-queue": {
"command": "uv",
"args": [
"run",
"--with", "mcp[cli]>=1.0.0,<2",
"--with", "mistralai>=1.0.0,<2",
"--no-project",
"/absolute/path/to/mmq.py",
"--mcp"
],
"env": {
"MISTRAL_API_KEY": "your-mistral-api-key"
}
}
}
}After changing config, restart the client. Manual Vibe checklist: docs/SMOKE_VIBE.md.
3. Environment variables (optional)
Variable | Default | Purpose |
| (required) | Mistral API key |
| per-user under tempdir | Shared queue DB file path |
|
| Seconds between starts (free-tier pacing) |
|
| Default model name |
| off | Offline / e2e: fake client ( |
Other knobs (MMQ_MAX_WAIT_TIME, MMQ_MAX_RETRIES, …) exist for tuning; see mmq.py.
MCP tools
When the server is running, clients can use the following tools:
ask_mistral
Argument | Type | Default | Description |
prompt | string | null | Single-shot user prompt text |
messages | array | null | Conversation history ( |
model | string |
| Mistral model name |
system_prompt | string | null | Custom system prompt (only when using |
priority | number | 2 | Task priority (1: high, 2: normal, 3: low) |
get_queue_status
Returns current shared queue / rate-limit status as JSON:
Field | Type | Description |
pending | number | Tasks waiting in the queue |
processing | number | Tasks currently claimed / running |
seconds_until_next_slot | number | Seconds until the shared API gate opens |
current_wait_interval | number | Active shared wait interval (seconds) |
in_flight | boolean | Whether any task is currently processing |
Control data location
The coordination temp DB is stored in a per-user directory created with mode 0700:
Default:
<tempdir>/mcp_mistral_queue_<USER>/mcp_mistral_flow_control.db
(tempfile.gettempdir(), often/tmpon Linux)Override: set
MMQ_TEMP_DB_PATHto a full file path (parent dir is created with0700)
Tests
# Unit + e2e (fake API; no network required)
uv run --with 'mcp[cli]>=1.0.0,<2' --with 'mistralai>=1.0.0,<2' \
--with pytest --with pytest-asyncio --no-project \
python -m pytest tests/ -v -m "not live"
# e2e only
uv run --with 'mcp[cli]>=1.0.0,<2' --with 'mistralai>=1.0.0,<2' \
--with pytest --with pytest-asyncio --no-project \
python -m pytest tests/e2e -v -m "not live"
# Live API (optional; consumes free-tier quota)
export MISTRAL_API_KEY=...
uv run --with 'mcp[cli]>=1.0.0,<2' --with 'mistralai>=1.0.0,<2' \
--with pytest --with pytest-asyncio --no-project \
python -m pytest tests/e2e/test_live_api.py -v -m livee2e uses MMQ_FAKE_API=1 and a short MMQ_BASE_WAIT_TIME to exercise process boundaries (CLI / MCP stdio).
For a manual Vibe UI check, see docs/SMOKE_VIBE.md.
Example: batch-style use of mmq (scripts/translate_readme.py)
Besides the CLI and MCP server, you can call the queue from Python. This repo ships a small sample:
scripts/translate_readme.py — regenerate locale READMEs from the English source via the same free-tier queue as mmq / ask_mistral.
Idea | Why it fits mmq |
Occasional job | Docs change far less often than chat traffic |
Can wait ~31s | ja then fr each take a gated slot |
Shared DB | Does not bypass other free-tier clients on the machine |
Programmatic API | Uses |
Locales workflow: edit README.md (English) only; do not hand-maintain README.ja.md / README.fr.md.
export MISTRAL_API_KEY=...
# optional: TRANSLATE_MODEL=mistral-small-latest
# From a git checkout (imports mmq.py on PYTHONPATH via the script)
python scripts/translate_readme.py # → README.ja.md + README.fr.md
python scripts/translate_readme.py --lang ja # one language
python scripts/translate_readme.py --dry-run # preview, no writeWhat the sample does:
Protects fenced code blocks (line FSM) and inline
codewith placeholdersEnqueues one translation job per language through
execute_mistral_queue_asyncRestores placeholders, fixes the language switcher, validates (e.g. balanced fences)
Writes outputs atomically
Use it as a template for other infrequent batch jobs (summaries, structured extraction) that should share the free-tier gate.
Acknowledgments
sioois for sharing information about the Mistral API free tier (link).
@fujibee for providing insights on using queues with SQLite WAL mode (#agmsg).
shunsuke_suzuki for the AI-friendly CLI development methodology (link).
Thank you all!
Further docs
docs/SMOKE_VIBE.md — Vibe / MCP manual smoke
docs/SEARCH_POSITIONING.md — where web search belongs (outside mmq base)
docs/tasks.md — backlog
docs/NOTES.md — design notes
License
MIT License
Copyright (c) 2026 utenadev
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Tools
Related MCP Servers
- AlicenseAqualityCmaintenanceMCP server for "taming the Claude" with structured task queues.144670MIT
- Alicense-qualityCmaintenanceLocal-first MCP server that enables multiple Claude agents to coordinate through a shared message bus with SQLite persistence and real-time clock anchoring.MIT
- Flicense-qualityCmaintenanceMCP server enabling natural-language querying of SQLite databases via schema discovery, GraphRAG retrieval, and safely guarded read-only SQL execution.
- Flicense-qualityCmaintenanceAn MCP server that provides an AI LLM orchestrator supporting multiple providers (LM Studio, Ollama, OpenAI, generic) plus SQLite-backed memory, kanban, and todo databases for persistent task and knowledge management.
Related MCP Connectors
Hosted MCP server for LLM cost estimation, model comparison, and budget-aware routing.
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/utenadev/mistral-managed-queue'
If you have feedback or need assistance with the MCP directory API, please join our Discord server