BlindWrite MCP
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| DB_PATH | No | Absolute path to the SQLite database file (e.g., /absolute/path/to/BlindWrite_MCP/data/blindwrite.sqlite). Defaults to a local data directory if not specified. | |
| LOG_LEVEL | No | Logging level for the server (e.g., 'info', 'debug', 'error'). Defaults to 'info'. | |
| OPENROUTER_API_KEY | Yes | Your OpenRouter API key (e.g., sk-or-v1-...). Required for all OpenRouter requests. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
| prompts | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| benchmark_create_taskA | Create a new blind AI writing benchmark task with a category, prompt, and optional evaluation criteria. |
| benchmark_list_modelsB | List all AI writing models available in the benchmark registry with pricing information. |
| benchmark_generate_outputsA | Generate writing outputs for a benchmark task across competing models via OpenRouter. Models remain strictly anonymous. |
| benchmark_start_duelB | Start a randomized, blind A/B battle between two outputs for evaluation. Model identities remain masked. |
| benchmark_submit_voteA | Submit a blind user preference vote ('A', 'B', or 'tie') with optional reasoning and dimension ratings. |
| benchmark_get_resultsB | Retrieve benchmark results for a task. When reveal is true, model identities are unmasked. |
| benchmark_get_leaderboardA | Get personal or global model rankings calculated with Bradley-Terry Maximum Likelihood Estimation and Elo. |
| benchmark_compare_modelsB | Compare two AI models head-to-head using accumulated pairwise benchmark battle history. |
| benchmark_analyze_preferencesB | Analyze empirical user preference patterns (conciseness, structure, tone) from battle voting history. |
| benchmark_get_model_statsA | Get comprehensive performance metrics, win rates, and ranking score for a specific model. |
| writer_generateA | PRIMARY WRITING TOOL. Use this tool whenever the user asks to write, draft, or compose content (emails, articles, proposals, essays, sales copy, social posts). Instead of generating long-form drafts with Claude output tokens, first outline the strategy and key arguments, then call this tool to delegate the draft generation to OpenRouter models (DeepSeek V3, Llama 3.3, GPT-4o, etc.). Automatically selects the user's #1 ranked model from their personal leaderboard or cost-effective DeepSeek V3. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
| writing-orchestrator | Orchestrate token-saving AI writing: Claude outlines the strategy, OpenRouter models draft the content via writer_generate. |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 11 tools
Most benchmark tools have clearly distinct resource-action boundaries, but benchmark_compare_models, benchmark_get_model_stats, and benchmark_get_leaderboard all describe performance summaries and could be confused by an agent. benchmark_generate_outputs and writer_generate are somewhat related, though writer_generate is explicitly marked as the primary writing tool.
Ten tools consistently follow a benchmark_<verb>_<noun> naming pattern, which is highly predictable. writer_generate breaks the pattern by using noun_verb form and dropping the benchmark_ prefix, making the set slightly inconsistent.
Eleven tools is well-scoped for a blind benchmark workflow plus an integrated writing generation tool. Each tool serves a recognizable step or query in the system, and the count is neither bloated nor thin.
The core benchmark lifecycle is covered: create task, generate outputs, start duel, submit vote, get results, and retrieve leaderboards. Obvious gaps include no list/update/delete for tasks and no way to enumerate active duels, but these are workable minor omissions.