Chain of Draft (CoD) MCP Server
초안 체인(CoD) MCP 서버
개요
이 MCP 서버는 연구 논문 "Chain of Draft: Thinking Faster by Writing Less"에 설명된 초안 체인(CoD) 추론 방식을 구현합니다. CoD는 LLM이 과제를 해결하는 동안 간결하면서도 유익한 중간 추론 결과를 생성할 수 있도록 하는 새로운 패러다임으로, 토큰 사용량을 크게 줄이는 동시에 정확성을 유지합니다.
Related MCP server: Visum Thinker MCP Server
주요 이점
효율성 : 토큰 사용량이 크게 감소(표준 CoT의 7.6%에 불과)
속도 : 생성 시간이 짧아 응답 속도가 빠릅니다.
비용 절감 : LLM 호출에 대한 API 비용 절감
유지된 정확도 : CoT와 비교했을 때 유사하거나 더 향상된 정확도
유연성 : 다양한 추론 작업 및 도메인에 적용 가능
특징
초안 구현의 핵심 체인
간결한 추론 단계(일반적으로 5단어 이하)
형식 적용
답변 추출
성과 분석
토큰 사용 추적
솔루션 정확도 모니터링
실행 시간 측정
도메인별 성능 측정 항목
적응형 단어 제한
자동 복잡도 추정
단어 제한의 동적 조정
도메인별 교정
포괄적인 예제 데이터베이스
CoT에서 CoD로의 변환
도메인별 예시(수학, 코드, 생물학, 물리학, 화학, 퍼즐)
문제 유사성에 기반한 검색 예시
형식 시행
단어 제한 준수를 보장하기 위한 사후 처리
계단 구조 보존
준수 분석
하이브리드 추론 접근 방식
CoD와 CoT 간 자동 선택
도메인별 최적화
과거 성과 기반 선택
OpenAI API 호환성
표준 OpenAI 클라이언트를 위한 드롭인 교체
완성 및 채팅 인터페이스 모두 지원
기존 워크플로에 쉽게 통합
설정 및 설치
필수 조건
Python 3.10+(Python 구현용)
Node.js 18+(JavaScript 구현용)
Anthropic API 키
파이썬 설치
저장소를 복제합니다
종속성 설치:
지엑스피1
.env파일에서 API 키를 구성합니다.ANTHROPIC_API_KEY=your_api_key_here서버를 실행합니다:
python server.py
자바스크립트 설치
저장소를 복제합니다
종속성 설치:
npm install.env파일에서 API 키를 구성합니다.ANTHROPIC_API_KEY=your_api_key_here서버를 실행합니다:
node index.js
Claude 데스크톱 통합
Claude Desktop과 통합하려면:
claude.ai/download 에서 Claude Desktop을 설치하세요
Claude Desktop 구성 파일을 만들거나 편집합니다.
~/Library/Application Support/Claude/claude_desktop_config.json서버 구성을 추가합니다(Python 버전):
{ "mcpServers": { "chain-of-draft": { "command": "python3", "args": ["/absolute/path/to/cod/server.py"], "env": { "ANTHROPIC_API_KEY": "your_api_key_here" } } } }또는 JavaScript 버전의 경우:
{ "mcpServers": { "chain-of-draft": { "command": "node", "args": ["/absolute/path/to/cod/index.js"], "env": { "ANTHROPIC_API_KEY": "your_api_key_here" } } } }Claude Desktop을 다시 시작하세요
Claude CLI를 사용하여 서버를 추가할 수도 있습니다.
# For Python implementation
claude mcp add chain-of-draft -e ANTHROPIC_API_KEY="your_api_key_here" "python3 /absolute/path/to/cod/server.py"
# For JavaScript implementation
claude mcp add chain-of-draft -e ANTHROPIC_API_KEY="your_api_key_here" "node /absolute/path/to/cod/index.js"사용 가능한 도구
Chain of Draft 서버는 다음과 같은 도구를 제공합니다.
도구 | 설명 |
| 초안 추론을 사용하여 문제 해결 |
| CoD로 수학 문제를 풀어보세요 |
| CoD를 사용하여 코딩 문제 해결 |
| CoD를 사용하여 논리 문제를 해결하세요 |
| CoD와 CoT의 성능 통계를 확인하세요 |
| 토큰 감소 통계 가져오기 |
| 문제 복잡성 분석 |
개발자 사용
파이썬 클라이언트
Python 코드에서 Chain of Draft 클라이언트를 직접 사용하려면 다음을 수행하세요.
from client import ChainOfDraftClient
# Create client
cod_client = ChainOfDraftClient()
# Use directly
result = await cod_client.solve_with_reasoning(
problem="Solve: 247 + 394 = ?",
domain="math"
)
print(f"Answer: {result['final_answer']}")
print(f"Reasoning: {result['reasoning_steps']}")
print(f"Tokens used: {result['token_count']}")JavaScript 클라이언트
JavaScript/Node.js 애플리케이션의 경우:
import { Anthropic } from "@anthropic-ai/sdk";
import dotenv from "dotenv";
// Load environment variables
dotenv.config();
// Create the Anthropic client
const anthropic = new Anthropic({
apiKey: process.env.ANTHROPIC_API_KEY,
});
// Import the Chain of Draft client
import chainOfDraftClient from './lib/chain-of-draft-client.js';
// Use the client
async function solveMathProblem() {
const result = await chainOfDraftClient.solveWithReasoning({
problem: "Solve: 247 + 394 = ?",
domain: "math",
max_words_per_step: 5
});
console.log(`Answer: ${result.final_answer}`);
console.log(`Reasoning: ${result.reasoning_steps}`);
console.log(`Tokens used: ${result.token_count}`);
}
solveMathProblem();구현 세부 사항
서버는 Python과 JavaScript 구현 모두에서 사용 가능하며, 두 구현 모두 여러 가지 통합 구성 요소로 구성됩니다.
파이썬 구현
AnalyticsService : 다양한 문제 도메인과 추론 접근 방식에 걸쳐 성과 측정 항목을 추적합니다.
ComplexityEstimator : 문제를 분석하여 적절한 단어 제한을 결정합니다.
ExampleDatabase : CoT 예제를 CoD 형식으로 변환하여 예제를 관리하고 검색합니다.
FormatEnforcer : 추론 단계가 단어 제한을 준수하도록 보장합니다.
ReasoningSelector : 문제 특성에 따라 CoD와 CoT 중에서 지능적으로 선택합니다.
JavaScript 구현
analyticsDb : 성능 지표 추적을 위한 메모리 내 데이터베이스
복잡성 추정기 : 문제를 분석하여 복잡성과 적절한 단어 제한을 결정합니다.
formatEnforcer : 추론 단계가 단어 제한을 준수하도록 보장합니다.
reasoningSelector : 문제 특성과 과거 성과를 기반으로 CoD와 CoT 중에서 자동으로 선택합니다.
두 구현 모두 동일한 핵심 원칙을 따르고 동일한 MCP 도구를 제공하므로 대부분의 사용 사례에서 상호 교환이 가능합니다.
특허
이 프로젝트는 오픈 소스이며 MIT 라이선스에 따라 제공됩니다.
Available Tools
7 toolsanalyze_problem_complexityC
Analyze the complexity of a problem
| Name | Required | Description | Default |
|---|---|---|---|
| problem | Yes | The problem to analyze | |
| domain | No | Problem domain |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It doesn't disclose behavioral traits such as whether this is a read-only analysis, if it requires specific inputs beyond the schema, what the output format might be, or any rate limits. The description is too minimal to offer meaningful context beyond the basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no wasted words, making it front-loaded and easy to parse. However, it's so brief that it under-specifies the tool's purpose, slightly reducing its effectiveness despite the conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete. It doesn't explain what 'analyze' entails, what results to expect, or how it fits with sibling tools. For a tool with 2 parameters and no structured behavioral hints, more context is needed to guide an agent effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with clear documentation for both parameters ('problem' and 'domain'). The description doesn't add any meaning beyond what the schema provides, such as explaining how 'domain' influences the analysis. Since schema coverage is high, the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Analyze the complexity of a problem' states a vague purpose with the verb 'analyze' and resource 'complexity of a problem', but it doesn't specify what complexity means (e.g., computational, conceptual, time) or how it differs from siblings like 'logic_solve' or 'math_solve'. It's not tautological but lacks specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The description doesn't mention context, prerequisites, or exclusions, and with siblings like 'logic_solve' or 'code_solve' that might handle related tasks, there's no differentiation to help an agent choose appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chain_of_draft_solveC
Solve a reasoning problem using Chain of Draft approach
| Name | Required | Description | Default |
|---|---|---|---|
| problem | Yes | The problem to solve | |
| domain | No | Domain for context (math, logic, code, common-sense, etc.) | |
| max_words_per_step | No | Maximum words per reasoning step | |
| approach | No | Force 'CoD' or 'CoT' approach | |
| enforce_format | No | Whether to enforce the word limit | |
| adaptive_word_limit | No | Adjust word limits based on complexity |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions 'Chain of Draft approach' but doesn't disclose behavioral traits such as how it handles reasoning steps, output format, error conditions, or computational requirements. For a tool with 6 parameters and no output schema, this lack of detail is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose. It's appropriately sized and front-loaded, with no wasted words. However, it could be more structured by briefly hinting at key parameters or outcomes to improve clarity without adding bulk.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 parameters, no annotations, no output schema), the description is incomplete. It doesn't explain what 'Chain of Draft' means, how the output is structured, or provide context for parameter usage. Without this, an AI agent might struggle to invoke the tool correctly or interpret results, especially compared to more specific siblings like 'math_solve'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema documents all parameters thoroughly. The description adds no meaning beyond the schema, as it doesn't explain parameter interactions (e.g., how 'max_words_per_step' relates to 'adaptive_word_limit') or provide examples. Baseline score of 3 is appropriate since the schema does the heavy lifting, but the description doesn't compensate with additional insights.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool 'Solve[s] a reasoning problem using Chain of Draft approach', which provides a verb ('solve') and resource ('reasoning problem'), but it's vague about what 'Chain of Draft' entails compared to alternatives like Chain of Thought (CoT). It doesn't distinguish from siblings like 'math_solve' or 'logic_solve', leaving ambiguity about when to use this over those specific tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives is provided. The description mentions 'Chain of Draft approach' but doesn't explain its advantages over CoT or when to choose it over sibling tools like 'code_solve' or 'analyze_problem_complexity'. Usage is implied by the tool name and approach, but no clear context or exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
code_solveC
Solve a coding problem using Chain of Draft reasoning
| Name | Required | Description | Default |
|---|---|---|---|
| problem | Yes | The coding problem to solve | |
| approach | No | Force 'CoD' or 'CoT' approach | |
| max_words_per_step | No | Maximum words per reasoning step |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'Chain of Draft reasoning' but does not explain what this entails, such as step-by-step reasoning, potential outputs, error handling, or computational limits. This leaves significant gaps in understanding how the tool behaves beyond its basic function.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose. It is appropriately sized and front-loaded, with no unnecessary words, though it could benefit from more detail to improve clarity and completeness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (solving coding problems with a specific reasoning approach), no annotations, and no output schema, the description is incomplete. It fails to explain what 'Chain of Draft reasoning' is, what the output looks like, or any behavioral traits, leaving significant gaps for effective tool use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all parameters ('problem', 'approach', 'max_words_per_step') with descriptions. The description does not add any additional meaning or context beyond what the schema provides, such as examples or usage tips for parameters, resulting in a baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool 'Solve[s] a coding problem using Chain of Draft reasoning', which provides a verb ('Solve') and resource ('coding problem') but is vague about what 'Chain of Draft reasoning' entails. It distinguishes from some siblings like 'logic_solve' or 'math_solve' by specifying 'coding problem', but the distinction from 'chain_of_draft_solve' is unclear, making the purpose somewhat ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is provided on when to use this tool versus alternatives. The description mentions 'Chain of Draft reasoning' but does not explain when this approach is preferred over other methods or tools like 'analyze_problem_complexity' or 'logic_solve'. This lack of context leaves usage unclear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_performance_statsC
Get performance statistics for CoD vs CoT approaches
| Name | Required | Description | Default |
|---|---|---|---|
| domain | No | Filter for specific domain |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool 'Get[s] performance statistics,' implying a read-only operation, but doesn't specify whether it requires authentication, has rate limits, returns real-time or historical data, or what format the statistics are in. For a tool with no annotations, this is a significant gap in behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence: 'Get performance statistics for CoD vs CoT approaches.' It's front-loaded with the core purpose, has zero wasted words, and is appropriately sized for the tool's apparent complexity. Every part of the sentence contributes to understanding the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (performance statistics comparison), lack of annotations, and no output schema, the description is incomplete. It doesn't explain what 'performance statistics' entail (e.g., metrics like accuracy, speed, cost), how CoD vs CoT are defined, or what the return values look like. For a tool that likely involves nuanced data analysis, this leaves too many gaps for effective agent use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with one parameter 'domain' documented as 'Filter for specific domain.' The description doesn't add any meaning beyond this, such as examples of domains or how filtering affects the results. Since the schema does the heavy lifting, the baseline score of 3 is appropriate, as the description neither compensates nor detracts from the schema's information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get performance statistics for CoD vs CoT approaches.' It specifies the verb ('Get') and resource ('performance statistics'), and distinguishes the scope (CoD vs CoT approaches). However, it doesn't explicitly differentiate from sibling tools like 'get_token_reduction' or 'analyze_problem_complexity', which might also relate to performance metrics, so it's not a perfect 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any prerequisites, context, or exclusions, and with sibling tools like 'get_token_reduction' that might overlap in performance analysis, there's no explicit comparison or usage rules. This leaves the agent without clear direction on tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_token_reductionB
Get token reduction statistics for CoD vs CoT
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It only states what the tool does (get statistics) without revealing any behavioral traits such as whether it's read-only, requires authentication, has rate limits, or what the output format might be. This is a significant gap for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It is appropriately sized and front-loaded, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete. It doesn't explain what the statistics include (e.g., metrics, timeframes, or data sources), how the results are structured, or any behavioral context. For a tool that likely returns data, this leaves significant gaps in understanding its full functionality.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, and the schema description coverage is 100%, so there are no parameters to document. The description doesn't need to add parameter semantics, and it appropriately avoids mentioning any. A baseline of 4 is applied since no parameters exist, and the description doesn't introduce unnecessary complexity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as retrieving 'token reduction statistics for CoD vs CoT' (Chain-of-Draft vs Chain-of-Thought), which is a specific verb+resource combination. However, it doesn't differentiate this tool from sibling tools like 'get_performance_stats' or 'analyze_problem_complexity', which might provide related metrics, so it falls short of a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any context, prerequisites, or exclusions, nor does it reference sibling tools that might offer overlapping or complementary functionality, leaving the agent with no usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
logic_solveC
Solve a logic problem using Chain of Draft reasoning
| Name | Required | Description | Default |
|---|---|---|---|
| problem | Yes | The logic problem to solve | |
| approach | No | Force 'CoD' or 'CoT' approach | |
| max_words_per_step | No | Maximum words per reasoning step |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'Chain of Draft reasoning' but doesn't explain what this entails, how it differs from other approaches, or any operational constraints like rate limits, error handling, or output format. The description is too vague to inform the agent about the tool's behavior beyond its basic purpose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It is appropriately sized and front-loaded, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a logic-solving tool with no annotations and no output schema, the description is insufficient. It doesn't explain the 'Chain of Draft reasoning' method, how results are returned, or any behavioral traits. The agent lacks critical context to use this tool effectively compared to its siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with clear documentation for all three parameters. The description adds no additional semantic information about the parameters beyond what's in the schema. According to the rules, with high schema coverage (>80%), the baseline score is 3 even without param info in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Solve a logic problem using Chain of Draft reasoning'. It specifies the verb ('solve'), resource ('logic problem'), and method ('Chain of Draft reasoning'). However, it doesn't explicitly differentiate from sibling tools like 'chain_of_draft_solve' or 'math_solve', which appear related.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. With siblings like 'chain_of_draft_solve', 'code_solve', and 'math_solve', there's no indication of when this specific 'logic_solve' tool is appropriate, nor any mention of prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
math_solveC
Solve a math problem using Chain of Draft reasoning
| Name | Required | Description | Default |
|---|---|---|---|
| problem | Yes | The math problem to solve | |
| approach | No | Force 'CoD' or 'CoT' approach | |
| max_words_per_step | No | Maximum words per reasoning step |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'Chain of Draft reasoning' but does not explain what this means in practice, such as how it processes the problem, what output to expect, or any limitations (e.g., accuracy, computational constraints). This lack of detail makes it inadequate for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It is front-loaded and appropriately sized, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a math-solving tool with no annotations and no output schema, the description is incomplete. It lacks details on how the tool behaves, what the output looks like, or any error conditions. This makes it insufficient for an agent to understand the tool's full context and usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters ('problem', 'approach', 'max_words_per_step') with descriptions. The description does not add any meaning beyond this, such as clarifying the 'approach' parameter's 'CoD' or 'CoT' options or providing examples. Baseline 3 is appropriate when the schema handles parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool 'Solve[s] a math problem using Chain of Draft reasoning', which provides a verb ('Solve') and resource ('math problem') but is vague about what 'Chain of Draft reasoning' entails. It distinguishes from some siblings like 'code_solve' or 'logic_solve' by specifying 'math problem', but the distinction from 'chain_of_draft_solve' is unclear, as both mention 'Chain of Draft'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is provided on when to use this tool versus alternatives. The description mentions 'Chain of Draft reasoning' but does not explain when this approach is preferred over other methods or tools like 'analyze_problem_complexity' or 'chain_of_draft_solve'. This leaves the agent without clear usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
- First observed
analyze_problem_complexity - First observed
chain_of_draft_solve - First observed
code_solve - First observed
get_performance_stats - First observed
get_token_reduction - First observed
logic_solve - First observed
math_solve
TDQS
Scored across 7 tools
The tools have clear distinctions in problem domains (math, logic, code, general reasoning) but there is significant overlap between 'chain_of_draft_solve' and the domain-specific solvers (math_solve, logic_solve, code_solve), which could cause confusion about when to use the general versus specific versions. The analysis and statistics tools are clearly distinct from the solving tools.
Most tools follow a consistent snake_case pattern with descriptive names (e.g., 'analyze_problem_complexity', 'get_performance_stats'), but there's a minor inconsistency with 'chain_of_draft_solve' using a longer prefix while others use simpler domain names. The verb usage is reasonably consistent with 'solve', 'analyze', and 'get' patterns.
With 7 tools, this is well-scoped for a server focused on Chain of Draft problem-solving and analysis. The count covers solving across multiple domains plus performance analytics, without being overwhelming or too sparse for the apparent purpose.
The toolset provides good coverage for solving problems (math, logic, code, general) and analyzing performance/token usage, which aligns well with the CoD domain. A minor gap exists in not having tools for configuring or customizing the CoD approach (e.g., setting parameters), but core workflows are well supported.
Maintenance
Related MCP Connectors
Reduces AI Agent token usage by 40% via three-stage SOP workflow.
Agent-to-agent reasoning-as-a-service: chain-of-thought, analysis, and decision support.
Memory that reasons: continual learning for stateful agents. Better context, fewer tokens.
Same functionality, consuming only 1/20 of the context window tokens.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnhances AI model capabilities with structured, retrieval-augmented thinking processes that enable dynamic thought chains, parallel exploration paths, and recursive refinement cycles for improved reasoning.124MIT
- AlicenseBqualityNot gradedmaintenanceProvides structured sequential thinking capabilities for AI assistants to break down complex problems into manageable steps, revise thoughts, and explore alternative reasoning paths.29-
- AlicenseNot gradedqualityCmaintenanceTransforms prompts into Chain of Draft (CoD) or Chain of Thought (CoT) format to enhance LLM reasoning quality while reducing token usage by up to 92.4%, supporting multiple LLM providers including Claude, GPT, Ollama, and local models.31 npm19MIT
- AlicenseBqualityNot gradedmaintenanceEnables AI assistants to perform structured, step-by-step reasoning by breaking down complex problems into numbered thoughts, with support for revising previous steps and exploring alternative reasoning paths.5-