Skip to main content
Glama
stat-guy

Chain of Draft (CoD) MCP Server

by stat-guy

초안 체인(CoD) MCP 서버

개요

이 MCP 서버는 연구 논문 "Chain of Draft: Thinking Faster by Writing Less"에 설명된 초안 체인(CoD) 추론 방식을 구현합니다. CoD는 LLM이 과제를 해결하는 동안 간결하면서도 유익한 중간 추론 결과를 생성할 수 있도록 하는 새로운 패러다임으로, 토큰 사용량을 크게 줄이는 동시에 정확성을 유지합니다.

Related MCP server: Visum Thinker MCP Server

주요 이점

  • 효율성 : 토큰 사용량이 크게 감소(표준 CoT의 7.6%에 불과)

  • 속도 : 생성 시간이 짧아 응답 속도가 빠릅니다.

  • 비용 절감 : LLM 호출에 대한 API 비용 절감

  • 유지된 정확도 : CoT와 비교했을 때 유사하거나 더 향상된 정확도

  • 유연성 : 다양한 추론 작업 및 도메인에 적용 가능

특징

  1. 초안 구현의 핵심 체인

    • 간결한 추론 단계(일반적으로 5단어 이하)

    • 형식 적용

    • 답변 추출

  2. 성과 분석

    • 토큰 사용 추적

    • 솔루션 정확도 모니터링

    • 실행 시간 측정

    • 도메인별 성능 측정 항목

  3. 적응형 단어 제한

    • 자동 복잡도 추정

    • 단어 제한의 동적 조정

    • 도메인별 교정

  4. 포괄적인 예제 데이터베이스

    • CoT에서 CoD로의 변환

    • 도메인별 예시(수학, 코드, 생물학, 물리학, 화학, 퍼즐)

    • 문제 유사성에 기반한 검색 예시

  5. 형식 시행

    • 단어 제한 준수를 보장하기 위한 사후 처리

    • 계단 구조 보존

    • 준수 분석

  6. 하이브리드 추론 접근 방식

    • CoD와 CoT 간 자동 선택

    • 도메인별 최적화

    • 과거 성과 기반 선택

  7. OpenAI API 호환성

    • 표준 OpenAI 클라이언트를 위한 드롭인 교체

    • 완성 및 채팅 인터페이스 모두 지원

    • 기존 워크플로에 쉽게 통합

설정 및 설치

필수 조건

  • Python 3.10+(Python 구현용)

  • Node.js 18+(JavaScript 구현용)

  • Anthropic API 키

파이썬 설치

  1. 저장소를 복제합니다

  2. 종속성 설치:

    지엑스피1

  3. .env 파일에서 API 키를 구성합니다.

    ANTHROPIC_API_KEY=your_api_key_here
  4. 서버를 실행합니다:

    python server.py

자바스크립트 설치

  1. 저장소를 복제합니다

  2. 종속성 설치:

    npm install
  3. .env 파일에서 API 키를 구성합니다.

    ANTHROPIC_API_KEY=your_api_key_here
  4. 서버를 실행합니다:

    node index.js

Claude 데스크톱 통합

Claude Desktop과 통합하려면:

  1. claude.ai/download 에서 Claude Desktop을 설치하세요

  2. Claude Desktop 구성 파일을 만들거나 편집합니다.

    ~/Library/Application Support/Claude/claude_desktop_config.json
  3. 서버 구성을 추가합니다(Python 버전):

    {
        "mcpServers": {
            "chain-of-draft": {
                "command": "python3",
                "args": ["/absolute/path/to/cod/server.py"],
                "env": {
                    "ANTHROPIC_API_KEY": "your_api_key_here"
                }
            }
        }
    }

    또는 JavaScript 버전의 경우:

    {
        "mcpServers": {
            "chain-of-draft": {
                "command": "node",
                "args": ["/absolute/path/to/cod/index.js"],
                "env": {
                    "ANTHROPIC_API_KEY": "your_api_key_here"
                }
            }
        }
    }
  4. Claude Desktop을 다시 시작하세요

Claude CLI를 사용하여 서버를 추가할 수도 있습니다.

# For Python implementation
claude mcp add chain-of-draft -e ANTHROPIC_API_KEY="your_api_key_here" "python3 /absolute/path/to/cod/server.py"

# For JavaScript implementation
claude mcp add chain-of-draft -e ANTHROPIC_API_KEY="your_api_key_here" "node /absolute/path/to/cod/index.js"

사용 가능한 도구

Chain of Draft 서버는 다음과 같은 도구를 제공합니다.

도구

설명

chain_of_draft_solve

초안 추론을 사용하여 문제 해결

math_solve

CoD로 수학 문제를 풀어보세요

code_solve

CoD를 사용하여 코딩 문제 해결

logic_solve

CoD를 사용하여 논리 문제를 해결하세요

get_performance_stats

CoD와 CoT의 성능 통계를 확인하세요

get_token_reduction

토큰 감소 통계 가져오기

analyze_problem_complexity

문제 복잡성 분석

개발자 사용

파이썬 클라이언트

Python 코드에서 Chain of Draft 클라이언트를 직접 사용하려면 다음을 수행하세요.

from client import ChainOfDraftClient

# Create client 
cod_client = ChainOfDraftClient()

# Use directly
result = await cod_client.solve_with_reasoning(
    problem="Solve: 247 + 394 = ?",
    domain="math"
)

print(f"Answer: {result['final_answer']}")
print(f"Reasoning: {result['reasoning_steps']}")
print(f"Tokens used: {result['token_count']}")

JavaScript 클라이언트

JavaScript/Node.js 애플리케이션의 경우:

import { Anthropic } from "@anthropic-ai/sdk";
import dotenv from "dotenv";

// Load environment variables
dotenv.config();

// Create the Anthropic client
const anthropic = new Anthropic({
  apiKey: process.env.ANTHROPIC_API_KEY,
});

// Import the Chain of Draft client
import chainOfDraftClient from './lib/chain-of-draft-client.js';

// Use the client
async function solveMathProblem() {
  const result = await chainOfDraftClient.solveWithReasoning({
    problem: "Solve: 247 + 394 = ?",
    domain: "math",
    max_words_per_step: 5
  });
  
  console.log(`Answer: ${result.final_answer}`);
  console.log(`Reasoning: ${result.reasoning_steps}`);
  console.log(`Tokens used: ${result.token_count}`);
}

solveMathProblem();

구현 세부 사항

서버는 Python과 JavaScript 구현 모두에서 사용 가능하며, 두 구현 모두 여러 가지 통합 구성 요소로 구성됩니다.

파이썬 구현

  1. AnalyticsService : 다양한 문제 도메인과 추론 접근 방식에 걸쳐 성과 측정 항목을 추적합니다.

  2. ComplexityEstimator : 문제를 분석하여 적절한 단어 제한을 결정합니다.

  3. ExampleDatabase : CoT 예제를 CoD 형식으로 변환하여 예제를 관리하고 검색합니다.

  4. FormatEnforcer : 추론 단계가 단어 제한을 준수하도록 보장합니다.

  5. ReasoningSelector : 문제 특성에 따라 CoD와 CoT 중에서 지능적으로 선택합니다.

JavaScript 구현

  1. analyticsDb : 성능 지표 추적을 위한 메모리 내 데이터베이스

  2. 복잡성 추정기 : 문제를 분석하여 복잡성과 적절한 단어 제한을 결정합니다.

  3. formatEnforcer : 추론 단계가 단어 제한을 준수하도록 보장합니다.

  4. reasoningSelector : 문제 특성과 과거 성과를 기반으로 CoD와 CoT 중에서 자동으로 선택합니다.

두 구현 모두 동일한 핵심 원칙을 따르고 동일한 MCP 도구를 제공하므로 대부분의 사용 사례에서 상호 교환이 가능합니다.

특허

이 프로젝트는 오픈 소스이며 MIT 라이선스에 따라 제공됩니다.

Available Tools

7 tools
analyze_problem_complexityC

Analyze the complexity of a problem

ParametersJSON Schema
NameRequiredDescriptionDefault
problemYesThe problem to analyze
domainNoProblem domain

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It doesn't disclose behavioral traits such as whether this is a read-only analysis, if it requires specific inputs beyond the schema, what the output format might be, or any rate limits. The description is too minimal to offer meaningful context beyond the basic action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no wasted words, making it front-loaded and easy to parse. However, it's so brief that it under-specifies the tool's purpose, slightly reducing its effectiveness despite the conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete. It doesn't explain what 'analyze' entails, what results to expect, or how it fits with sibling tools. For a tool with 2 parameters and no structured behavioral hints, more context is needed to guide an agent effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with clear documentation for both parameters ('problem' and 'domain'). The description doesn't add any meaning beyond what the schema provides, such as explaining how 'domain' influences the analysis. Since schema coverage is high, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Analyze the complexity of a problem' states a vague purpose with the verb 'analyze' and resource 'complexity of a problem', but it doesn't specify what complexity means (e.g., computational, conceptual, time) or how it differs from siblings like 'logic_solve' or 'math_solve'. It's not tautological but lacks specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description doesn't mention context, prerequisites, or exclusions, and with siblings like 'logic_solve' or 'code_solve' that might handle related tasks, there's no differentiation to help an agent choose appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chain_of_draft_solveC

Solve a reasoning problem using Chain of Draft approach

ParametersJSON Schema
NameRequiredDescriptionDefault
problemYesThe problem to solve
domainNoDomain for context (math, logic, code, common-sense, etc.)
max_words_per_stepNoMaximum words per reasoning step
approachNoForce 'CoD' or 'CoT' approach
enforce_formatNoWhether to enforce the word limit
adaptive_word_limitNoAdjust word limits based on complexity

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions 'Chain of Draft approach' but doesn't disclose behavioral traits such as how it handles reasoning steps, output format, error conditions, or computational requirements. For a tool with 6 parameters and no output schema, this lack of detail is a significant gap in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose. It's appropriately sized and front-loaded, with no wasted words. However, it could be more structured by briefly hinting at key parameters or outcomes to improve clarity without adding bulk.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, no annotations, no output schema), the description is incomplete. It doesn't explain what 'Chain of Draft' means, how the output is structured, or provide context for parameter usage. Without this, an AI agent might struggle to invoke the tool correctly or interpret results, especially compared to more specific siblings like 'math_solve'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema documents all parameters thoroughly. The description adds no meaning beyond the schema, as it doesn't explain parameter interactions (e.g., how 'max_words_per_step' relates to 'adaptive_word_limit') or provide examples. Baseline score of 3 is appropriate since the schema does the heavy lifting, but the description doesn't compensate with additional insights.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool 'Solve[s] a reasoning problem using Chain of Draft approach', which provides a verb ('solve') and resource ('reasoning problem'), but it's vague about what 'Chain of Draft' entails compared to alternatives like Chain of Thought (CoT). It doesn't distinguish from siblings like 'math_solve' or 'logic_solve', leaving ambiguity about when to use this over those specific tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives is provided. The description mentions 'Chain of Draft approach' but doesn't explain its advantages over CoT or when to choose it over sibling tools like 'code_solve' or 'analyze_problem_complexity'. Usage is implied by the tool name and approach, but no clear context or exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

code_solveC

Solve a coding problem using Chain of Draft reasoning

ParametersJSON Schema
NameRequiredDescriptionDefault
problemYesThe coding problem to solve
approachNoForce 'CoD' or 'CoT' approach
max_words_per_stepNoMaximum words per reasoning step

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'Chain of Draft reasoning' but does not explain what this entails, such as step-by-step reasoning, potential outputs, error handling, or computational limits. This leaves significant gaps in understanding how the tool behaves beyond its basic function.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose. It is appropriately sized and front-loaded, with no unnecessary words, though it could benefit from more detail to improve clarity and completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (solving coding problems with a specific reasoning approach), no annotations, and no output schema, the description is incomplete. It fails to explain what 'Chain of Draft reasoning' is, what the output looks like, or any behavioral traits, leaving significant gaps for effective tool use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents all parameters ('problem', 'approach', 'max_words_per_step') with descriptions. The description does not add any additional meaning or context beyond what the schema provides, such as examples or usage tips for parameters, resulting in a baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool 'Solve[s] a coding problem using Chain of Draft reasoning', which provides a verb ('Solve') and resource ('coding problem') but is vague about what 'Chain of Draft reasoning' entails. It distinguishes from some siblings like 'logic_solve' or 'math_solve' by specifying 'coding problem', but the distinction from 'chain_of_draft_solve' is unclear, making the purpose somewhat ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance is provided on when to use this tool versus alternatives. The description mentions 'Chain of Draft reasoning' but does not explain when this approach is preferred over other methods or tools like 'analyze_problem_complexity' or 'logic_solve'. This lack of context leaves usage unclear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_performance_statsC

Get performance statistics for CoD vs CoT approaches

ParametersJSON Schema
NameRequiredDescriptionDefault
domainNoFilter for specific domain

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool 'Get[s] performance statistics,' implying a read-only operation, but doesn't specify whether it requires authentication, has rate limits, returns real-time or historical data, or what format the statistics are in. For a tool with no annotations, this is a significant gap in behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence: 'Get performance statistics for CoD vs CoT approaches.' It's front-loaded with the core purpose, has zero wasted words, and is appropriately sized for the tool's apparent complexity. Every part of the sentence contributes to understanding the tool's function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (performance statistics comparison), lack of annotations, and no output schema, the description is incomplete. It doesn't explain what 'performance statistics' entail (e.g., metrics like accuracy, speed, cost), how CoD vs CoT are defined, or what the return values look like. For a tool that likely involves nuanced data analysis, this leaves too many gaps for effective agent use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with one parameter 'domain' documented as 'Filter for specific domain.' The description doesn't add any meaning beyond this, such as examples of domains or how filtering affects the results. Since the schema does the heavy lifting, the baseline score of 3 is appropriate, as the description neither compensates nor detracts from the schema's information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Get performance statistics for CoD vs CoT approaches.' It specifies the verb ('Get') and resource ('performance statistics'), and distinguishes the scope (CoD vs CoT approaches). However, it doesn't explicitly differentiate from sibling tools like 'get_token_reduction' or 'analyze_problem_complexity', which might also relate to performance metrics, so it's not a perfect 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any prerequisites, context, or exclusions, and with sibling tools like 'get_token_reduction' that might overlap in performance analysis, there's no explicit comparison or usage rules. This leaves the agent without clear direction on tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_token_reductionB

Get token reduction statistics for CoD vs CoT

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states what the tool does (get statistics) without revealing any behavioral traits such as whether it's read-only, requires authentication, has rate limits, or what the output format might be. This is a significant gap for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It is appropriately sized and front-loaded, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete. It doesn't explain what the statistics include (e.g., metrics, timeframes, or data sources), how the results are structured, or any behavioral context. For a tool that likely returns data, this leaves significant gaps in understanding its full functionality.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 0 parameters, and the schema description coverage is 100%, so there are no parameters to document. The description doesn't need to add parameter semantics, and it appropriately avoids mentioning any. A baseline of 4 is applied since no parameters exist, and the description doesn't introduce unnecessary complexity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose as retrieving 'token reduction statistics for CoD vs CoT' (Chain-of-Draft vs Chain-of-Thought), which is a specific verb+resource combination. However, it doesn't differentiate this tool from sibling tools like 'get_performance_stats' or 'analyze_problem_complexity', which might provide related metrics, so it falls short of a perfect score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any context, prerequisites, or exclusions, nor does it reference sibling tools that might offer overlapping or complementary functionality, leaving the agent with no usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

logic_solveC

Solve a logic problem using Chain of Draft reasoning

ParametersJSON Schema
NameRequiredDescriptionDefault
problemYesThe logic problem to solve
approachNoForce 'CoD' or 'CoT' approach
max_words_per_stepNoMaximum words per reasoning step

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'Chain of Draft reasoning' but doesn't explain what this entails, how it differs from other approaches, or any operational constraints like rate limits, error handling, or output format. The description is too vague to inform the agent about the tool's behavior beyond its basic purpose.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It is appropriately sized and front-loaded, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a logic-solving tool with no annotations and no output schema, the description is insufficient. It doesn't explain the 'Chain of Draft reasoning' method, how results are returned, or any behavioral traits. The agent lacks critical context to use this tool effectively compared to its siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with clear documentation for all three parameters. The description adds no additional semantic information about the parameters beyond what's in the schema. According to the rules, with high schema coverage (>80%), the baseline score is 3 even without param info in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Solve a logic problem using Chain of Draft reasoning'. It specifies the verb ('solve'), resource ('logic problem'), and method ('Chain of Draft reasoning'). However, it doesn't explicitly differentiate from sibling tools like 'chain_of_draft_solve' or 'math_solve', which appear related.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. With siblings like 'chain_of_draft_solve', 'code_solve', and 'math_solve', there's no indication of when this specific 'logic_solve' tool is appropriate, nor any mention of prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

math_solveC

Solve a math problem using Chain of Draft reasoning

ParametersJSON Schema
NameRequiredDescriptionDefault
problemYesThe math problem to solve
approachNoForce 'CoD' or 'CoT' approach
max_words_per_stepNoMaximum words per reasoning step

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'Chain of Draft reasoning' but does not explain what this means in practice, such as how it processes the problem, what output to expect, or any limitations (e.g., accuracy, computational constraints). This lack of detail makes it inadequate for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It is front-loaded and appropriately sized, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a math-solving tool with no annotations and no output schema, the description is incomplete. It lacks details on how the tool behaves, what the output looks like, or any error conditions. This makes it insufficient for an agent to understand the tool's full context and usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters ('problem', 'approach', 'max_words_per_step') with descriptions. The description does not add any meaning beyond this, such as clarifying the 'approach' parameter's 'CoD' or 'CoT' options or providing examples. Baseline 3 is appropriate when the schema handles parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool 'Solve[s] a math problem using Chain of Draft reasoning', which provides a verb ('Solve') and resource ('math problem') but is vague about what 'Chain of Draft reasoning' entails. It distinguishes from some siblings like 'code_solve' or 'logic_solve' by specifying 'math problem', but the distinction from 'chain_of_draft_solve' is unclear, as both mention 'Chain of Draft'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance is provided on when to use this tool versus alternatives. The description mentions 'Chain of Draft reasoning' but does not explain when this approach is preferred over other methods or tools like 'analyze_problem_complexity' or 'chain_of_draft_solve'. This leaves the agent without clear usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updates
    • First observedanalyze_problem_complexity
    • First observedchain_of_draft_solve
    • First observedcode_solve
    • First observedget_performance_stats
    • First observedget_token_reduction
    • First observedlogic_solve
    • First observedmath_solve

TDQS

B3.1/5.0

Scored across 7 tools

Disambiguation3/5

The tools have clear distinctions in problem domains (math, logic, code, general reasoning) but there is significant overlap between 'chain_of_draft_solve' and the domain-specific solvers (math_solve, logic_solve, code_solve), which could cause confusion about when to use the general versus specific versions. The analysis and statistics tools are clearly distinct from the solving tools.

Naming Consistency4/5

Most tools follow a consistent snake_case pattern with descriptive names (e.g., 'analyze_problem_complexity', 'get_performance_stats'), but there's a minor inconsistency with 'chain_of_draft_solve' using a longer prefix while others use simpler domain names. The verb usage is reasonably consistent with 'solve', 'analyze', and 'get' patterns.

Tool Count5/5

With 7 tools, this is well-scoped for a server focused on Chain of Draft problem-solving and analysis. The count covers solving across multiple domains plus performance analytics, without being overwhelming or too sparse for the apparent purpose.

Completeness4/5

The toolset provides good coverage for solving problems (math, logic, code, general) and analyzing performance/token usage, which aligns well with the CoD domain. A minor gap exists in not having tools for configuring or customizing the CoD approach (e.g., setting parameters), but core workflows are well supported.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers