Code Reasoning MCP Server
Servidor MCP de razonamiento de código
Un servidor de Protocolo de Contexto de Modelo (MCP) que mejora la capacidad de Claude para resolver tareas de programación complejas a través del pensamiento estructurado, paso a paso.
Instalación rápida
Configure Claude Desktop editando:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonVentanas:
%APPDATA%\Claude\claude_desktop_config.jsonLinux:
~/.config/Claude/claude_desktop_config.json
{ "mcpServers": { "code-reasoning": { "command": "npx", "args": ["-y", "@mettamatt/code-reasoning"] } } }Configurar VS Code:
{
"mcp": {
"servers": {
"code-reasoning": {
"command": "npx",
"args": ["-y", "@mettamatt/code-reasoning"]
}
}
}
}Related MCP server: Sequential Thinking MCP Server
Uso
Para activar este MCP, agregue lo siguiente a sus mensajes de chat:
Use sequential thinking to reason about this.Utilice indicaciones listas para usar que activen el razonamiento de código:

Haga clic en el ícono "+" en la ventana de chat de Claude Desktop o, en Claude Code, escriba
/helppara ver los comandos específicos.Seleccione "Agregar desde Razonamiento de código" de las herramientas disponibles
Elija una plantilla de solicitud y complete la información requerida
Envíe el formulario para agregar el mensaje a su mensaje de chat y presione Enter.
Consulte la Guía de indicaciones para obtener detalles sobre el uso de las plantillas de indicaciones.
Opciones de línea de comandos
--debug: Habilitar el registro detallado--helpo-h: Mostrar información de ayuda
Características principales
Enfoque de programación : Optimizado para tareas de codificación y resolución de problemas.
Pensamiento estructurado : Dividir problemas complejos en pasos manejables
Ramificación del pensamiento : explorar múltiples caminos de solución en paralelo
Revisión del pensamiento : refinar el razonamiento anterior a medida que mejora la comprensión
Límites de seguridad : se detiene automáticamente después de 20 pasos de reflexión para evitar bucles
Avisos listos para usar : plantillas predefinidas para tareas de desarrollo comunes
Documentación
Documentación detallada disponible en el directorio docs:
Ejemplos de uso : Ejemplos de pensamiento secuencial con el servidor MCP
Guía de configuración : Todas las opciones de configuración para el servidor MCP
Guía de indicaciones : uso y personalización de indicaciones con el servidor MCP
Marco de pruebas : información de pruebas
Estructura del proyecto
├── index.ts # Entry point
├── src/ # Implementation source files
└── test/ # Testing frameworkEvaluación rápida
El servidor MCP de razonamiento de código incluye un sistema de evaluación de indicaciones que evalúa la capacidad de Claude para seguir las indicaciones de razonamiento de código. Este sistema permite:
Probar diferentes variaciones de indicaciones frente a problemas de escenarios
Verificación de la adherencia al formato de parámetros
Puntuación de la calidad de la solución
Para utilizar el sistema de evaluación rápida, ejecute:
npm run evalComparación y desarrollo rápidos
Se dedicó un gran esfuerzo al desarrollo del indicador óptimo para el servidor de razonamiento de código. La implementación actual utiliza el indicador HYBRID_DESIGN, que resultó ganador en nuestro proceso de evaluación.
Comparamos cuatro diseños de indicaciones diferentes:
Diseño de avisos | Descripción |
SECUENCIAL | El diseño original de la propuesta de pensamiento secuencial |
POR DEFECTO | El mensaje de línea base utilizado anteriormente en el servidor |
CÓDIGO_RAZONAMIENTO_0_30 | Una variante experimental centrada en el razonamiento específico del código |
DISEÑO HÍBRIDO | Un diseño refinado que incorpora los mejores elementos de otros enfoques. |
Nuestra evaluación en siete escenarios de programación diferentes mostró que HYBRID_DESIGN superó a otras indicaciones:
Guión | DISEÑO HÍBRIDO | CÓDIGO_RAZONAMIENTO_0_30 | POR DEFECTO | SECUENCIAL |
Selección de algoritmos | 87% | 82% | 88% | 82% |
Identificación de errores | 87% | 91% | 88% | 92% |
Implementación en múltiples etapas | 83% | 67% | 79% | 82% |
Análisis del diseño del sistema | 82% | 87% | 78% | 82% |
Tarea de depuración de código | 92% | 87% | 92% | 92% |
Optimización del compilador | 83% | 78% | 67% | 73% |
Estrategia de caché | 86% | 88% | 82% | 87% |
Promedio | 86% | 83% | 82% | 84% |
El mensaje HYBRID_DESIGN demostró, por un margen, la calidad de solución promedio más alta (86%) y el rendimiento más consistente en todos los escenarios, sin puntuaciones inferiores al 80%. También generó la mayor cantidad de reflexiones. El archivo src/server.ts se ha actualizado para utilizar este diseño óptimo de mensaje.
Personalmente, creo que la mayor mejora fue agregar esto al final del mensaje: "✍️ Termine cada pensamiento preguntando: "¿Qué me estoy perdiendo o necesito reconsiderar?"
Consulte el marco de pruebas para obtener más detalles sobre el sistema de evaluación rápida.
Licencia
Este proyecto está licenciado bajo la Licencia MIT. Consulte el archivo de LICENCIA para más detalles.
Available Tools
1 toolcode-reasoningA
🧠 Code Reasoning Tool (using sequential thinking)
Purpose → break complex problems into self-auditing, exploratory thought steps that can branch, revise, or back-track until a single, well-supported answer emerges.
WHEN TO CALL
• Multi-step planning, design, debugging, or open-ended analysis
• Whenever further private reasoning or hypothesis testing is required before replying to the user
ENCOURAGED PRACTICES
🔍 Question aggressively – ask "What am I missing?" after each step
🔄 Revise freely – mark is_revision=true even late in the chain
🌿 Branch often – explore plausible alternatives in parallel; you can merge or discard branches later
↩️ Back-track – if a path looks wrong, start a new branch from an earlier thought
❓ Admit uncertainty – explicitly note unknowns and schedule extra thoughts to resolve them
MUST DO
✅ Put every private reasoning step in thought
✅ Keep thought_number correct; update total_thoughts when scope changes
✅ Use is_revision & branch_from_thought/branch_id precisely
✅ Set next_thought_needed=false only when all open questions are resolved
✅ Abort and summarise if thought_number > 20
DO NOT
⛔️ Reveal the content of thought to the end-user
⛔️ Continue thinking once next_thought_needed=false
⛔️ Assume thoughts must proceed strictly linearly – branching is first-class
PARAMETER CHEAT-SHEET
• thought (string) – current reasoning step
• next_thought_needed (boolean) – request further thinking?
• thought_number (int ≥ 1) – 1-based counter
• total_thoughts (int ≥ 1) – mutable estimate
• is_revision, revises_thought (int) – mark corrections
• branch_from_thought, branch_id – manage alternative paths
• needs_more_thoughts (boolean) – optional hint that more thoughts may follow
All JSON keys must use lower_snake_case.
EXAMPLE ✔️
{
"thought": "List solution candidates and pick the most promising",
"thought_number": 1,
"total_thoughts": 4,
"next_thought_needed": true
}EXAMPLE ✔️ (branching late)
{
"thought": "Alternative approach: treat it as a graph-search problem",
"thought_number": 6,
"total_thoughts": 8,
"branch_from_thought": 3,
"branch_id": "B1",
"next_thought_needed": true
}| Name | Required | Description | Default |
|---|---|---|---|
| branch_from_thought | No | Branching point thought number | |
| branch_id | No | Identifier for the current branch | |
| is_revision | No | Whether this is a revision of a previous thought | |
| needs_more_thoughts | No | Optional hint that more thoughts may follow | |
| next_thought_needed | Yes | Whether another thought step is needed | |
| revises_thought | No | Which thought is being revised | |
| thought | Yes | Your current reasoning step | |
| thought_number | Yes | Current thought number (1-based) | |
| total_thoughts | Yes | Estimated total thoughts needed (can be adjusted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure and excels at this. It provides extensive behavioral guidance including 'ENCOURAGED PRACTICES' (questioning, revising, branching, backtracking, admitting uncertainty), 'MUST DO' rules (put every step in thought, keep counters correct, use branching/revision flags precisely, set next_thought_needed=false only when resolved, abort after 20 thoughts), and 'DO NOT' prohibitions (don't reveal thoughts to user, don't continue after next_thought_needed=false, don't assume linear thinking). This comprehensively describes how the tool should be used.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (Purpose, WHEN TO CALL, ENCOURAGED PRACTICES, MUST DO, DO NOT, PARAMETER CHEAT-SHEET, EXAMPLES) that make it easy to navigate. While comprehensive, it maintains focus with each section serving a clear purpose. Some sections could be slightly more concise, but overall the structure enhances readability and information retrieval.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (9 parameters, no annotations, no output schema), the description provides exceptional contextual completeness. It covers purpose, usage guidelines, behavioral patterns, parameter semantics, and practical examples. The description fully compensates for the lack of annotations and output schema by providing comprehensive guidance on how to use this complex reasoning tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds significant value through the 'PARAMETER CHEAT-SHEET' section that provides practical guidance on parameter usage beyond the schema's basic descriptions. It explains the relationships between parameters (e.g., how is_revision and revises_thought work together, how branching parameters relate) and includes important implementation notes like 'All JSON keys must use lower_snake_case.' The examples further illustrate parameter usage in context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as 'break complex problems into self-auditing, exploratory thought steps that can branch, revise, or back-track until a single, well-supported answer emerges.' This is specific (verb+resource+methodology) and distinguishes it from any potential alternatives. The 'Purpose →' section provides a concise, accurate summary of what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'WHEN TO CALL' section explicitly lists scenarios for using this tool: 'Multi-step planning, design, debugging, or open-ended analysis' and 'Whenever further private reasoning or hypothesis testing is required before replying to the user.' It provides clear guidance on when this tool should be invoked versus when to respond directly to the user.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.0.0- First observed
code-reasoning
TDQS
Scored across 1 tool
With only one tool, there is no possibility of confusion or overlap between tools. The single tool has a clearly defined purpose for code reasoning and problem-solving, so agents cannot misselect between multiple options.
Since there is only one tool named 'code-reasoning', naming consistency is inherently perfect. There are no other tools to compare against, so no inconsistencies can exist in the tool set.
A single tool for a 'Code Reasoning MCP Server' feels too minimal for the apparent scope. While the tool is feature-rich internally, the server's purpose suggests it should offer multiple specialized reasoning tools (e.g., for debugging, design, analysis) rather than one monolithic tool, making the count inappropriate.
The server claims to handle 'code reasoning' but provides only one general-purpose tool. This creates significant gaps: there are no specialized tools for different reasoning tasks (e.g., debugging vs. design), no tools for input/output handling, and no way to manage reasoning sessions independently, leading to potential agent failures in complex workflows.
Maintenance
Related MCP Connectors
Adaptive plan/build/review cycles for AI coding assistants, persisted across sessions.
GodPrompt MCP + Agent Skill for coding agents with TDD, debugging, verification, and task routing.
Source-checked CLI guides and model-aware planning for Claude Code, Codex, and Grok Build.
Deterministic AI code review, with an audit record. Governance inside the agent loop.
Related MCP Servers
- AlicenseBqualityNot gradedmaintenanceProvides structured sequential thinking capabilities for AI assistants to break down complex problems into manageable steps, revise thoughts, and explore alternative reasoning paths.29-
- FlicenseAqualityDmaintenanceEnables structured, step-by-step problem-solving with dynamic revision and branching capabilities. Supports breaking down complex problems into manageable steps while allowing course corrections and alternative reasoning paths.1136,644 npm1-
- AlicenseAqualityBmaintenanceEnables structured step-by-step reasoning with branching, revisions, and self-critique to help break down complex problems into manageable steps with confidence tracking and thought history search.719 npm7MIT
- FlicenseAqualityDmaintenanceEnables structured, step-by-step problem-solving through dynamic thinking processes that can be revised, branched, and adjusted as understanding deepens. Supports breaking down complex problems into manageable steps with the ability to revise previous thoughts and explore alternative reasoning paths.1136,644 npm-
Appeared in Searches
- A server for learning and finding resources about SAS programming
- Tools and frameworks for thinking about software development
- A server for finding information about sequential thinking
- A tool for critical thinking and devil's advocate analysis of AI model plans
- Tools for slow thinking, step-back reasoning, and contextual memory capabilities