Markdownify MCP Server
Markdownify MCP 서버

Markdownify는 다양한 파일 형식과 웹 콘텐츠를 마크다운 형식으로 변환하는 모델 컨텍스트 프로토콜(MCP) 서버입니다. PDF, 이미지, 오디오 파일, 웹 페이지 등을 읽기 쉽고 공유 가능한 마크다운 텍스트로 변환하는 도구 세트를 제공합니다.
기능
여러 파일 형식을 마크다운으로 변환:
PDF
이미지
오디오 (전사 포함)
DOCX
XLSX
PPTX
웹 콘텐츠를 마크다운으로 변환:
YouTube 동영상 자막
Bing 검색 결과
일반 웹 페이지
기존 마크다운 파일 검색
Related MCP server: Markdownify MCP Server
시작하기
이 저장소를 복제합니다.
의존성을 설치합니다:
bun installpreinstall단계에서.venv에 파이썬 가상 환경을 생성하고markitdown[all]을 설치합니다.프로젝트를 빌드합니다:
bun run build서버를 시작합니다:
bun start
개발
bun run dev를 사용하여 감시 모드에서 TypeScript 컴파일러를 시작합니다.src/server.ts를 수정하여 서버 동작을 사용자 정의합니다.src/tools.ts에서 도구를 추가하거나 수정합니다.
데스크톱 앱과 함께 사용하기
이 서버를 데스크톱 앱과 통합하려면 앱의 서버 구성에 다음을 추가하십시오:
{
"mcpServers": {
"markdownify": {
"command": "node",
"args": [
"{ABSOLUTE PATH TO FILE HERE}/dist/index.js"
]
}
}
}환경 변수
모든 경로는 기본적으로 적절한 값으로 설정되어 있습니다. 기본값이 설치 레이아웃에 맞지 않는 경우에만 재정의하십시오.
변수 | 기본값 | 목적 |
|
|
|
|
|
|
| 설정되지 않음 (제한 없음) | 서버가 읽을 수 있는 디렉토리의 경로 구분 기호로 구분된 목록(POSIX에서는 |
| 설정되지 않음 |
|
Docker와 함께 사용하기
빌드 및 실행:
docker build -t markdownify-mcp .
docker run --rm -i \
-v "$HOME/Documents:/data:ro" \
-e MD_ALLOWED_PATHS=/data \
markdownify-mcpDocker MCP 카탈로그(mcp/markdownify)에 대한 참고 사항:
서버가 읽기를 원하는 호스트 디렉토리를 컨테이너에 마운트한 다음, 도구에 컨테이너 경로를 전달하십시오(예:
/Users/you/Documents/foo.pdf가 아닌/data/foo.pdf).서버가 바인드 마운트와 일치하는 읽기 경계를 강제하도록
MD_ALLOWED_PATHS를 마운트된 디렉토리의 콜론 구분 목록으로 설정하십시오.게시된 Docker 이미지는
markitdown[pdf]만 설치합니다. 오디오 전사 및 이미지 OCR(audio-to-markdown,image-to-markdown)은[all]추가 기능이 필요하며 슬림 이미지에서는 실패합니다. 전체 기능 세트를 사용하려면 로컬 설치(bun install)를 사용하십시오.
사용 가능한 도구
youtube-to-markdown: YouTube 동영상을 마크다운으로 변환pdf-to-markdown: PDF 파일을 마크다운으로 변환bing-search-to-markdown: Bing 검색 결과를 마크다운으로 변환webpage-to-markdown: 웹 페이지를 마크다운으로 변환image-to-markdown: 이미지를 메타데이터와 함께 마크다운으로 변환audio-to-markdown: 오디오 파일을 전사와 함께 마크다운으로 변환docx-to-markdown: DOCX 파일을 마크다운으로 변환xlsx-to-markdown: XLSX 파일을 마크다운으로 변환pptx-to-markdown: PPTX 파일을 마크다운으로 변환get-markdown-file: 기존 마크다운 파일을 검색합니다. 파일 확장자는 *.md, *.markdown으로 끝나야 합니다.선택 사항:
MD_ALLOWED_PATHS를 설정하여 모든 파일 입력 도구를 디렉토리 목록으로 제한하십시오(예:MD_ALLOWED_PATHS=/data/in:/data/out bun start).
기여
기여를 환영합니다! 자유롭게 Pull Request를 제출해 주십시오.
라이선스
이 프로젝트는 MIT 라이선스에 따라 라이선스가 부여됩니다. 자세한 내용은 LICENSE 파일을 참조하십시오.
Available Tools
10 toolsaudio-to-markdownC
Convert an audio file to markdown, including transcription if possible
| Name | Required | Description | Default |
|---|---|---|---|
| filepath | Yes | Absolute path of the audio file to convert |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool converts audio to markdown with transcription, but lacks details on permissions needed, rate limits, error handling, or what 'if possible' implies (e.g., supported formats, transcription accuracy). For a tool with no annotations, this leaves significant gaps in understanding its operation and constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function without unnecessary words. It is front-loaded with the core action ('Convert an audio file to markdown') and adds a useful qualifier ('including transcription if possible'). Every part earns its place, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (audio processing and transcription) and lack of annotations and output schema, the description is incomplete. It doesn't explain return values (e.g., markdown content format), error cases, or behavioral details like supported audio formats. For a tool with no structured data to rely on, this leaves the agent under-informed about how to use it effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds no parameter-specific information beyond what the input schema provides. The schema has 100% coverage with a clear description for 'filepath' as the 'Absolute path of the audio file to convert.' Since the schema does the heavy lifting, the baseline score of 3 is appropriate, as the description doesn't compensate with additional semantics like file format requirements or path examples.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Convert an audio file to markdown, including transcription if possible.' It specifies the verb ('convert'), resource ('audio file'), and outcome ('to markdown'), distinguishing it from siblings that convert other file types (e.g., pdf-to-markdown). However, it doesn't explicitly differentiate from siblings beyond the resource type, such as by noting unique features like audio-specific transcription.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It mentions 'including transcription if possible,' which hints at a capability, but doesn't specify prerequisites (e.g., file format support), exclusions, or direct comparisons to siblings like youtube-to-markdown for audio from videos. Without explicit when/when-not instructions, it leaves usage ambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bing-search-to-markdownC
Convert a Bing search results page to markdown
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL of the Bing search results page |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the action ('Convert') but doesn't explain what the conversion entails (e.g., formatting, extraction of specific elements), potential errors (e.g., invalid URLs), or output characteristics (e.g., markdown structure). This leaves significant gaps in understanding the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste. It's front-loaded with the core action and resource, making it easy to parse quickly without unnecessary details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete. It doesn't address key aspects like what the markdown output includes (e.g., search results, metadata), error handling, or how it differs from sibling tools, leaving the agent with insufficient context for reliable use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, with the parameter 'url' clearly documented in the input schema. The description doesn't add any additional meaning beyond this, such as URL format requirements or examples, but the schema provides adequate baseline information, justifying a score of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Convert') and resource ('Bing search results page'), making the purpose understandable. However, it doesn't differentiate from sibling tools like 'webpage-to-markdown' or 'get-markdown-file', which could cause confusion about when to use this specific tool versus similar alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. With siblings like 'webpage-to-markdown' that might handle general webpages, the description lacks context on prerequisites (e.g., requiring a Bing search URL) or exclusions, leaving the agent to guess based on the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
docx-to-markdownC
Convert a DOCX file to markdown
| Name | Required | Description | Default |
|---|---|---|---|
| filepath | Yes | Absolute path of the DOCX file to convert |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the conversion action but doesn't describe what happens during conversion (e.g., formatting preservation, error handling, output location, or performance characteristics). For a file conversion tool with zero annotation coverage, this leaves significant gaps in understanding the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise—a single sentence with no wasted words. It is front-loaded with the core purpose, making it easy to scan and understand immediately. Every word earns its place by directly contributing to the tool's intent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete for a file conversion tool. It doesn't explain what the output is (e.g., markdown text, saved file), where it goes, or any limitations (e.g., file size, supported DOCX features). With siblings offering similar conversions, more context is needed to differentiate and use this tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, with the single parameter 'filepath' well-documented in the schema. The description doesn't add any parameter-specific details beyond what the schema provides, such as file format requirements or conversion options. With high schema coverage, the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('convert') and resource ('DOCX file to markdown'), making the purpose immediately understandable. It distinguishes from siblings by specifying the DOCX format, though it doesn't explicitly contrast with other conversion tools like pdf-to-markdown or pptx-to-markdown.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. With multiple sibling tools for converting different formats to markdown (e.g., pdf-to-markdown, pptx-to-markdown), there is no indication of when DOCX conversion is appropriate or what distinguishes it from other conversion methods.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get-markdown-fileC
Get a markdown file by absolute file path
| Name | Required | Description | Default |
|---|---|---|---|
| filepath | Yes | Absolute path to file of markdown'd text |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the tool retrieves a file but doesn't disclose behavioral traits such as error handling (e.g., if file doesn't exist), permissions needed, rate limits, or output format. This leaves significant gaps for a tool that likely involves file system access.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with zero waste. It's front-loaded with the core action and resource, making it efficient and easy to understand without unnecessary details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and a simple input schema, the description is incomplete. It doesn't cover key aspects like what happens on failure, the return format (e.g., raw markdown text or structured data), or how it differs from sibling tools, which are crucial for effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the parameter 'filepath' fully documented in the schema as 'Absolute path to file of markdown'd text'. The description adds no additional meaning beyond this, such as examples or constraints, so it meets the baseline score of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Get') and resource ('a markdown file'), specifying it retrieves content by absolute file path. However, it doesn't distinguish from sibling tools like 'audio-to-markdown' or 'webpage-to-markdown', which are conversion tools rather than retrieval tools, so it's not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The description implies usage for retrieving existing markdown files, but it doesn't mention prerequisites (e.g., file must exist), exclusions (e.g., not for creating or converting files), or compare to siblings like 'docx-to-markdown' for conversion tasks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
image-to-markdownC
Convert an image to markdown, including metadata and description
| Name | Required | Description | Default |
|---|---|---|---|
| filepath | Yes | Absolute path of the image file to convert |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions that the conversion includes 'metadata and description,' which adds some context about output content. However, it fails to disclose critical behavioral traits such as error handling (e.g., for invalid file paths), performance aspects (e.g., processing time or size limits), or output format details. For a tool with no annotations, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and front-loaded, consisting of a single sentence that directly states the tool's function. There is no wasted language or redundancy, making it efficient and easy to parse. Every word earns its place by conveying essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of file conversion and the lack of annotations and output schema, the description is incomplete. It does not explain what the markdown output includes beyond 'metadata and description,' such as formatting details or error cases. For a tool that performs a non-trivial operation with no structured output information, more context is needed to guide the agent effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with the 'filepath' parameter clearly documented as 'Absolute path of the image file to convert.' The description does not add any meaning beyond this, as it does not elaborate on parameter usage or constraints. According to the rules, with high schema coverage (>80%), the baseline score is 3, which is appropriate here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: converting an image to markdown, including metadata and description. It specifies the verb ('convert') and resource ('image'), making it easy to understand. However, it does not explicitly differentiate from sibling tools like 'pdf-to-markdown' or 'webpage-to-markdown', which perform similar conversions on different file types, so it lacks sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention any context, prerequisites, or exclusions, such as when to choose 'image-to-markdown' over other conversion tools like 'pdf-to-markdown' for different file formats. This leaves the agent without clear usage instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pdf-to-markdownC
Convert a PDF file to markdown
| Name | Required | Description | Default |
|---|---|---|---|
| filepath | Yes | Absolute path of the PDF file to convert |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the conversion action but doesn't mention potential limitations (e.g., formatting accuracy, large file handling, error conditions), output details, or side effects. For a tool with zero annotation coverage, this leaves significant gaps in understanding how it behaves.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero wasted words. It's front-loaded with the core action and resource, making it immediately understandable. This is an excellent example of conciseness in tool descriptions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete for a file conversion tool. It doesn't address output format details, error handling, or performance considerations, which are crucial for an agent to use it effectively. The simplicity of one parameter doesn't compensate for these omissions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the single parameter 'filepath' fully documented in the schema as 'Absolute path of the PDF file to convert'. The description adds no additional parameter semantics beyond implying PDF file input, so it meets the baseline for high schema coverage without compensating value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'convert' and the resource 'PDF file to markdown', making the purpose unambiguous. It doesn't explicitly distinguish from siblings like 'docx-to-markdown' or 'pptx-to-markdown', but the PDF specificity is inherent. This is clear but lacks explicit sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. With siblings like 'docx-to-markdown' and 'webpage-to-markdown', there's no indication of prerequisites, file format requirements, or comparative use cases. It's a basic statement of function without contextual usage advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pptx-to-markdownC
Convert a PPTX file to markdown
| Name | Required | Description | Default |
|---|---|---|---|
| filepath | Yes | Absolute path of the PPTX file to convert |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the conversion action but lacks behavioral details: it doesn't mention if the tool overwrites files, requires specific permissions, handles errors, or outputs to a specific location. For a file conversion tool with zero annotation coverage, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste: 'Convert a PPTX file to markdown'. It is front-loaded and appropriately sized, earning its place by clearly stating the core functionality without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of file conversion (which involves input/output handling and potential errors), no annotations, and no output schema, the description is incomplete. It doesn't explain what the tool returns (e.g., markdown text, file path), error conditions, or dependencies, leaving gaps for an AI agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with the 'filepath' parameter well-documented as 'Absolute path of the PPTX file to convert'. The description adds no additional parameter semantics beyond what the schema provides, so it meets the baseline of 3 for high schema coverage without extra value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Convert a PPTX file to markdown' clearly states the verb 'Convert' and the resource 'PPTX file', making the purpose immediately understandable. It distinguishes from siblings by specifying the PPTX format, though it doesn't explicitly differentiate from similar conversion tools like 'docx-to-markdown' beyond the file type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. With siblings like 'docx-to-markdown' and 'pdf-to-markdown', it doesn't specify scenarios where PPTX conversion is preferred or any prerequisites (e.g., file format compatibility). Usage is implied by the tool name but not articulated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
webpage-to-markdownC
Convert a webpage to markdown
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL of the webpage to convert |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the conversion action but lacks details on traits like error handling (e.g., invalid URLs, network issues), performance (e.g., timeouts, size limits), or output specifics (e.g., markdown format quality, included elements). This is a significant gap for a tool with no structured safety hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise—a single, direct sentence with zero wasted words. It's front-loaded with the core action, making it easy to parse quickly. Every word earns its place by conveying the essential purpose without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (conversion operation), lack of annotations, and no output schema, the description is incomplete. It doesn't address behavioral aspects like what happens on failure, what the markdown output includes, or limitations (e.g., dynamic content handling). For a tool with no structured safety or output info, more context is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, with the single parameter 'url' fully documented in the schema. The description adds no additional meaning beyond implying the URL is for a webpage, which is already clear from the schema. This meets the baseline score when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Convert') and resource ('a webpage to markdown'), making it immediately understandable. However, it doesn't distinguish this tool from its siblings (e.g., 'pdf-to-markdown', 'docx-to-markdown'), which all convert different content types to markdown, so it doesn't reach the highest score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any prerequisites (e.g., URL accessibility), exclusions (e.g., unsupported webpage types), or comparisons to sibling tools like 'bing-search-to-markdown' for web content. This leaves the agent with minimal context for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xlsx-to-markdownC
Convert an XLSX file to markdown
| Name | Required | Description | Default |
|---|---|---|---|
| filepath | Yes | Absolute path of the XLSX file to convert |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool converts files but doesn't explain how the conversion works, what happens to the original file, or any limitations like file size or formatting issues. This leaves key behavioral traits unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise—just one sentence—and front-loaded with the core action. There's no wasted language, making it easy to parse quickly. This efficiency is ideal for a simple tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete. It doesn't cover behavioral aspects like conversion quality or error handling, and with no output schema, it fails to explain what the markdown output looks like. For a conversion tool, this leaves significant gaps in understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with the 'filepath' parameter clearly documented. The description doesn't add any extra meaning beyond the schema, such as examples or constraints, but since the schema is comprehensive, a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: converting XLSX files to markdown format. It specifies the verb 'convert' and the resource 'XLSX file', making it easy to understand. However, it doesn't explicitly differentiate from sibling tools like 'docx-to-markdown' or 'pdf-to-markdown' beyond the file type, which keeps it from a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any prerequisites, such as file format requirements or when other conversion tools might be more appropriate. This lack of context makes it harder for an AI agent to choose correctly among the many sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube-to-markdownC
Convert a YouTube video to markdown, including transcript if available
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL of the YouTube video |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but offers minimal behavioral insight. It mentions transcript inclusion 'if available', hinting at conditional behavior, but lacks details on error handling, rate limits, authentication needs, or output format beyond markdown.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste. It's front-loaded with the core purpose and includes a useful qualifier about transcripts, making it appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description is incomplete. It doesn't explain what the markdown output contains (e.g., structure, metadata), potential errors, or dependencies like internet access, leaving significant gaps for agent usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the 'url' parameter fully. The description adds no additional parameter semantics beyond implying the URL must be for a YouTube video, which is minimal value over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'convert' and the resource 'YouTube video to markdown', specifying it includes transcript if available. It distinguishes from siblings by focusing on YouTube specifically, though it doesn't explicitly contrast with other conversion tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'audio-to-markdown' or 'webpage-to-markdown'. The description implies usage for YouTube videos but doesn't mention prerequisites, limitations, or comparative advantages.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
- First observed
audio-to-markdown - First observed
bing-search-to-markdown - First observed
docx-to-markdown - First observed
get-markdown-file - First observed
image-to-markdown - First observed
pdf-to-markdown - First observed
pptx-to-markdown - First observed
webpage-to-markdown - First observed
xlsx-to-markdown - First observed
youtube-to-markdown
TDQS
Scored across 10 tools
Each tool has a clearly distinct purpose focused on converting a specific input format (audio, Bing search, DOCX, etc.) to markdown. The descriptions specify unique source types, making it impossible to confuse which tool to use for a given conversion task.
All tools follow a consistent pattern of 'source-to-markdown' (e.g., audio-to-markdown, docx-to-markdown). This uniform naming convention makes it easy to predict tool names and understand their functions at a glance.
With 10 tools, the server covers a comprehensive range of common input formats (audio, documents, web content, spreadsheets, presentations, etc.) for markdown conversion. This count is well-scoped for the domain, providing thorough coverage without being overwhelming.
The tool set offers complete coverage for converting various media and file types to markdown, including audio, images, documents, webpages, and videos. There are no obvious gaps; each tool handles a distinct input format, ensuring agents can convert any supported source without dead ends.
Maintenance
Related MCP Connectors
Convert documents and web pages to clean Markdown: PDF, DOCX, XLSX, EPUB, scanned files, any URL.
Convert PDF, DOCX, HTML, and URLs to clean, LLM-ready markdown with tables preserved
Convert PDF, Word, Excel and scanned documents to Markdown, tables and RAG chunks. OCR.
Convert PDFs/images to Word, Excel, and Markdown. Split, merge, and watermark PDF files.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceConverts various file types (documents, images, audio, web content) to markdown format without requiring Docker, supporting PDF, Word, Excel, PowerPoint, images, audio files, web URLs, and more.308 npm14MIT
- AlicenseAqualityDmaintenanceConverts various file types (PDF, images, audio, DOCX, XLSX, PPTX) and web content (YouTube videos, web pages, Bing search results) into Markdown format for easy reading and sharing.10336 npmMIT
- AlicenseNot gradedqualityCmaintenanceConverts documents, webpages, and media files into markdown for AI assistants using Microsoft's MarkItDown and Crawl4AI. It enables tools to read PDFs, Office files, and JavaScript-rendered websites with support for OCR and image extraction.3MIT
- AlicenseNot gradedqualityDmaintenanceConverts files (PDF, images, audio, DOCX, XLSX, PPTX) and web content (YouTube transcripts, Bing search, general pages) to Markdown via the Model Context Protocol.Apache 2.0