calibre-mcp
calibre-mcp
로컬 Calibre 전자책 라이브러리에 대한 읽기/쓰기 액세스를 AI 어시스턴트(Claude Code, Claude Desktop, Cursor 등)에 제공하는 MCP 서버입니다.
라이브러리 검색, 최근 추가 항목 찾아보기, 자동 중복 감지와 함께 새 전자책 가져오기, "드롭 폴더" 받은 편지함 관리까지 — 모두 채팅으로 처리하세요.
일반적인 설정에서 구성 불필요: calibredb 실행 파일과 Calibre GUI가 사용하는 라이브러리는 macOS, Windows, Linux에서 자동으로 감지됩니다.
로케일 독립적: 모든 읽기는
calibredb --for-machine(JSON)을 거치므로 결과가 Calibre 인터페이스 언어에 의존하지 않습니다.기본적으로 안전: 중복 감지는 Calibre 자체에 의존합니다. 제목+저자가 이미 라이브러리에 있는 책은 건너뛰고 두 번 가져오지 않습니다.
도구
도구 | 설명 |
| 라이브러리 경로, 책 수, calibre 버전 |
| 전체 calibre 쿼리 언어 검색( |
| 라이브러리 찾아보기, 기본값은 최근 추가된 순 |
| 중복 감지와 함께 전자책 파일 또는 전체 디렉터리 가져오기 |
| 받은 편지함 폴더의 모든 항목을 가져와 정리 |
Related MCP server: Calibre MCP Server
요구 사항
Calibre(
calibredb를 제공하는 모든 버전, 9.x로 테스트됨)Python ≥ 3.10
설치
pipx(권장)
pipx install calibre-mcpuv
uv tool install calibre-mcppip
pip install calibre-mcp소스에서
git clone https://github.com/ZehuiTIAN/calibre-mcp.git
cd calibre-mcp
pip install -e ".[dev]" # dev extras add pytest + ruff구성
일반적인 설정에서는 모든 것이 선택 사항입니다. 각 값은 환경 변수로 고정할 수 있습니다.
변수 | 의미 | 기본값 |
| 사용할 라이브러리 디렉터리 | Calibre GUI가 사용하는 라이브러리( |
|
|
|
|
| 비활성화(사용 시 도구 오류 발생) |
| calibredb 호출이 중단되기까지의 시간(초) |
|
Claude Code에 등록
claude mcp add calibre-mcp -- calibre-mcp또는 ~/.claude.json → mcpServers(사용자 범위)에 추가:
"calibre_mcp": {
"type": "stdio",
"command": "calibre-mcp",
"args": [],
"env": {
"CALIBRE_LIBRARY_PATH": "/path/to/your/library",
"CALIBRE_INBOX_DIR": "/path/to/your/inbox"
}
}Claude Desktop / Cursor에 등록
claude_desktop_config.json(Claude Desktop) 또는 Cursor의 MCP 설정에 있는 MCP
서버 섹션에 동일한 JSON 블록을 추가하세요.
받은 편지함 워크플로
드롭 폴더는 손으로 다운로드한 책에 대해 "찾기 → 놓기 → 보관" 루프를 만들어 줍니다.
CALIBRE_INBOX_DIR을 폴더(예:~/Downloads/Calibre-Inbox)로 지정합니다.해당 폴더에 전자책을 다운로드합니다.
어시스턴트에게
import_inbox를 실행하도록 요청합니다.
모든 파일이 라이브러리에 추가됩니다. 성공적으로 추가된 파일과 중복 파일은
imported/YYYY-MM/ 아래에, 실패한 파일은 failed/ 아래에 정리됩니다. 받은
편지함 자체는 깨끗하게 유지되며 어떤 것도 덮어쓰지 않습니다(이름 충돌 시
-1, -2, ... 접미사가 붙습니다).
지원 확장자: .epub .pdf .mobi .azw .azw3 .djvu .cbz .cbr .fb2 .txt .md .docx .rtf .lit .prc .pdb .chm .htmlz.
주의 사항
데이터베이스 잠금: Calibre GUI가 라이브러리를 열어 둔 동안에는 쓰기 작업(
add_books,import_inbox)이 데이터베이스 잠금 오류로 실패할 수 있습니다. 읽기 전용 도구는 계속 작동합니다. 가져오기 전에 Calibre를 닫거나(또는 다른 라이브러리로 전환) 하세요.대용량 라이브러리:
add_books/import_inbox는 각 추가 전에 전체 책 ID 집합의 스냅샷을 만듭니다. 실제로는 빠르지만 가져오기당 책 수에 비례하는 O(책 수)입니다.이 도구는 여러분의 라이브러리만 정리합니다. 검색이나 다운로드 기능은 포함되어 있지 않습니다. 보유할 권리가 있는 콘텐츠에 사용하세요.
개발
pip install -e ".[dev]"
ruff check .
pytest -m "not integration" # unit tests, run anywhere
pytest # + integration tests (need local calibre)통합 테스트는 /tmp 아래에 임시 라이브러리를 만들고 실제 라이브러리는
건드리지 않습니다.
라이선스
MIT — LICENSE를 참조하세요.
중국어 요약
로컬 Calibre 라이브러리를 Claude/Cursor 등의 AI 어시스턴트에 연결하는 MCP
서버: 다 가지 도구(library_info / search_books / list_books /
add_books / import_inbox), calibredb와 라이브러리 위치 자동 감지, 전
플랫폼(Windows/macOS/Linux)에서 바로 사용 가능.
pipx install calibre-mcp
claude mcp add calibre-mcp -- calibre-mcp가져올 때 자동 중복 검사(제목+저자가 같은 책은 중복 가져오지 않음);
CALIBRE_INBOX_DIR 받은 편지함 디렉터리와 함께 사용하면 다운로드한 책을
넣고 "받은 편지함 가져오기"라고 말하면 아카이브되며, 처리된 파일은 자동으로
imported/ 또는 failed/로 분류됩니다.
참고: Calibre 데스크톱 프로그램이 같은 라이브러리를 열어 두면 데이터베이스 잠금으로 인해 쓰기 작업이 실패할 수 있으므로, 가져오기 전에 Calibre를 먼저 닫으세요.
Available Tools
11 toolsadd_booksA
Import ebook files (or directories of ebooks) into the library.
Books whose title + author are already in the library are skipped by calibre and reported as "duplicate". The Calibre GUI must not hold the library open while adding.
Args: paths: Absolute paths to ebook files or folders containing ebooks.
| Name | Required | Description | Default |
|---|---|---|---|
| paths | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does reveal two useful traits: books with matching title+author are skipped and reported as 'duplicate', and the Calibre GUI must not hold the library open. However, it does not mention whether files are copied or moved, whether directory traversal is recursive, what happens on partial failure, or any other side effects, leaving notable behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a one-line purpose statement, two concise behavioral caveats, and an Args section with a single parameter definition. No sentence is wasted, and the most important usage constraints are front-loaded before the parameter details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-parameter import tool with an output schema present, the description covers the essential prerequisites: absolute path semantics, file-or-directory inputs, duplicate handling, and the GUI-lock constraint. It does not mention whether folders are processed recursively or which ebook formats are supported, but these are minor gaps given the tool's simplicity and the availability of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only defines 'paths' as an array of strings with no description (0% schema coverage). The description fully compensates by specifying that paths must be 'Absolute paths' and can point to 'ebook files or folders containing ebooks.' This directly adds crucial meaning that the schema lacks, making the parameter semantics clear and unambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb, 'Import', and a clear resource, 'ebook files (or directories of ebooks) into the library.' It clearly differentiates from most siblings like search_books, list_books, and read_book. However, it does not explicitly distinguish itself from the sibling 'import_inbox', which may represent an overlapping alternative, so the differentiation is not complete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear operational context, explicitly stating that the Calibre GUI must not hold the library open while adding. It also explains the duplicate-skipping behavior, which helps an agent anticipate outcomes. However, it does not mention any alternative tools or conditions under which a sibling like import_inbox would be more appropriate, so exclusions are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_indexA
Index book contents for full-text search (one-time, cached).
Converts each book's text with calibre's converter and stores it in a local SQLite FTS5 index kept outside the library. Only books that are not indexed yet are processed; an initial run over a whole library can take several minutes.
Args: book_ids: Optional list of book ids to index; default: all books that are not indexed yet. limit: Optional cap on how many books to index in this call.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| book_ids | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure, and it delivers: it explains that the tool converts book text via calibre's converter, stores data in a local SQLite FTS5 index outside the library, only processes unindexed books, and may take several minutes initially. This gives the agent a strong understanding of side effects, persistence, and cost.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a concise summary up front followed by an Args section. Every sentence adds value: the operation, the caching behavior, the storage location, the idempotency, the time cost, and parameter semantics. There is no filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with only two optional parameters and an output schema, the description covers the important operational details: what gets built, where it is stored, which books are eligible, and the potential long runtime. The agent can confidently decide when to call it and what to expect, even without annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides only titles and defaults with 0% coverage, so the description fully compensates by explaining both parameters: book_ids defaults to all unindexed books, and limit caps how many books are processed in a single call. This is exactly the semantic information an agent needs to invoke the tool correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Index book contents for full-text search', which is a specific verb and resource that clearly distinguishes build_index from sibling search tools like search_books and search_in_book. The parenthetical 'one-time, cached' further clarifies its role as a maintenance/indexing operation rather than a query operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly establishes when this tool is relevant: it is a one-time cached indexing step that processes only unindexed books, and warns that an initial full-library run can take minutes. It does not explicitly name alternatives or state when not to use it, but the context is clear enough for an agent to recognize this as the setup step for full-text search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
detect_scanned_booksA
Find library books whose PDF has no usable text layer (scanned pages).
Such books stay unsearchable until OCR'd with ocr_book. Requires the
ocr extra (pip install 'calibre-mcp[ocr]'); no API key needed.
Args: limit: Maximum number of scanned books to report (1-200).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It usefully discloses the ocr-extra requirement and the no-API-key fact, and the verb 'Find' implies a non-mutating scan of PDFs. It does not explicitly state side-effect-free behavior or performance costs, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences deliver purpose, follow-up context, prerequisites, and parameter documentation without filler. The purpose is front-loaded and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a single-optional-parameter tool with an output schema, so return-value documentation is not needed. The description covers the task, prerequisite, parameter range, and next-step tool, making it complete for an agent to select and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description fully documents the only parameter: 'limit: Maximum number of scanned books to report (1-200).' This adds meaning and a validation range that the input schema completely lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Find library books whose PDF has no usable text layer (scanned pages).' This clearly distinguishes the tool from siblings like list_books and search_books, and connects it to the OCR workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states when this tool matters ('Such books stay unsearchable until OCR'd') and names the follow-up tool ocr_book, plus the installation prerequisite. It lacks an explicit when-not-to-use statement, but the context is clear enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_inboxA
Import every ebook file in the configured inbox folder into the library.
Requires the CALIBRE_INBOX_DIR environment variable. Successfully added
files and duplicates are moved to <inbox>/imported/YYYY-MM/; files that
fail to import are moved to <inbox>/failed/ for inspection.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the behavioral burden and does so well. It discloses the environment variable requirement and the exact file disposition side effects: successful files and duplicates move to `<inbox>/imported/YYYY-MM/`, while failed files move to `<inbox>/failed/` for inspection. This is strong transparency for a tool that mutates the filesystem.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the primary purpose, followed by prerequisites and side effects. Every sentence earns its place, with no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema present, the description is complete: it states what is imported, what configuration is required, and what happens to files in each outcome case. An agent has enough information to invoke the tool correctly and predict its effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no input parameters, so the schema requires no explanation. The description still adds relevant operational context by referencing the CALIBRE_INBOX_DIR environment variable, which effectively serves as the tool's input. Baseline 4 is appropriate for a zero-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Import every ebook file in the configured inbox folder into the library.' It clearly distinguishes this from siblings like add_books by emphasizing batch import from a configured inbox folder rather than individual file addition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is clear: this tool processes all ebook files in the inbox folder, and the CALIBRE_INBOX_DIR prerequisite is stated. However, it does not explicitly name alternatives or state when not to use it relative to add_books or other sibling tools, so it stops just short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
library_infoA
Show the active library: its path on disk, number of books, calibre version.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. The verb 'Show' and the listed output items convey a read-only informational operation. It does not explicitly state that nothing is modified, but the phrasing is unambiguous and sufficient for a stateless info tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that names the action, target, and expected output without any filler. Every word contributes meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter informational tool with an output schema present, the description fully covers what the agent needs to know: what the tool does and what information it returns. No additional context is required.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description focuses on output rather than parameters, which is appropriate since there are none to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Show') and resource ('the active library'), and enumerates exactly what is displayed: path on disk, number of books, and calibre version. This clearly distinguishes it from siblings like search_books or add_books, which perform different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose implies when to use it (whenever library-level metadata is needed), but it does not explicitly state alternatives or when not to use it. The sibling tools are functionally distinct enough that misrouting is unlikely, but there is no direct usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_booksA
Browse books in the library; defaults to most recently added first.
Args: limit: Maximum number of results (1-200, default 20). sort_by: calibre field to sort by, e.g. id, title, authors, timestamp. order: "desc" (default) or "asc". search: Optional calibre search expression to filter the listing.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| order | No | desc | |
| search | No | ||
| sort_by | No | id |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosing behavioral traits. It conveys listing behavior, default sort order, and filtering support, which is good for a read-oriented tool. However, it does not explicitly state that the operation is read-only/idempotent or describe any limitations beyond parameter ranges.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured. The opener is one sentence, followed by a clear Args list where every line adds meaningful detail without redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only listing tool with an output schema and four optional parameters, the description covers the key invocation details: defaults, sorting, and filtering. The main omission is guidance about when to use this tool instead of the search-related siblings, but that gap is largely a usage-guidelines issue rather than a completeness failure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description thoroughly documents all four parameters with defaults, allowed values, and examples (e.g., limit range 1-200, sort_by fields, order values, calibre search expression). This fully compensates for the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear operation ('Browse books in the library') and gives the resource and default ordering. It differentiates itself from sibling tools like search_books mostly through the verb 'browse', but it does not explicitly contrast with them, and the 'search' parameter introduces some overlap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives such as search_books or search_in_book. The description only implies a listing/filtering use case but never states exclusions or the preferred tool for complex search scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ocr_bookA
OCR a scanned PDF via the configured cloud provider and index the text.
Detection → page rendering → provider transcription → cache → full-text index → (by default) pandoc re-typeset EPUB attached to the book's existing record. Books with a usable text layer are left untouched.
Configuration: CALIBRE_OCR_PROVIDER (default "anthropic"), CALIBRE_OCR_API_KEY, CALIBRE_OCR_MODEL, CALIBRE_OCR_BASE_URL, CALIBRE_OCR_MAX_PAGES.
Args: book_id: Book whose PDF format should be OCR'd. typeset: Also build a re-typeset EPUB with pandoc (default True). import_format: Attach the EPUB to the book record (default True).
| Name | Required | Description | Default |
|---|---|---|---|
| book_id | Yes | ||
| typeset | No | ||
| import_format | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so admirably. It discloses the complete processing pipeline (detection, rendering, provider transcription, cache, full-text index, optional pandoc re-typeset EPUB attachment), the side effect on the book record, and the no-op condition for books with a text layer. It also surfaces relevant configuration variables, giving the agent insight into provider and model dependencies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured with a lead purpose sentence, a compact pipeline outline, a configuration list, and an Args section. Each section earns its place and the main action is front-loaded. The configuration block is somewhat lengthy for an agent that may not need these details, but it remains concise enough given the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter tool with no schema descriptions, the description covers the parameters, the operation pipeline, side effects, and the no-op condition. An output schema exists, so return values do not need explanation. The only notable gap is explicit routing to batch alternatives, but that is a usage-guidance nuance rather than a completeness blocker. Overall, the definition is sufficiently complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does. Each parameter is explicitly explained: book_id identifies the PDF to OCR, typeset controls pandoc re-typesetting, and import_format controls whether the EPUB is attached. Defaults are also noted, fully bridging the schema's lack of descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'OCR a scanned PDF via the configured cloud provider and index the text.' This clearly identifies the operation and distinguishes it from siblings like detect_scanned_books, which only detects, and ocr_scanned_books, which implies batch processing. The singular 'book_id' argument further anchors it as a per-book tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes clear context: the tool targets a single scanned PDF book and leaves books with a usable text layer untouched. However, it never explicitly names alternatives like ocr_scanned_books for batch OCR or build_index for indexing-only operations. The usage guidance is implied by the singular argument and pipeline, but not stated as when-to-use versus alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ocr_scanned_booksA
OCR all scanned-PDF books in the library, in resumable batches.
Detects scanned books, skips any that already have OCR output cached,
and runs the full ocr_book pipeline (OCR → index → optional re-typeset
EPUB) on up to limit remaining books. Repeat the call to work through
the whole library; per-batch caching makes interruptions safe.
Args: limit: Maximum books to OCR in this call (1-100, default 10). typeset: Also build re-typeset EPUBs with pandoc (default True). import_format: Attach the EPUBs to book records (default True).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| typeset | No | ||
| import_format | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and does it well: it skips books with cached OCR output, runs the full pipeline, attaches EPUBs when import_format is set, and is resumable. It does omit failure-mode or side-effect detail, but for a batch tool this is substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description opens with a one-sentence purpose, follows with a compact behavior paragraph, and ends with a terse Args list. Every sentence adds necessary information and there is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter batch tool with an output schema present, the description covers the what, the how (repeat calls), the safety of interruption, and detailed parameter semantics. The only minor omission is an explicit comparison to the single-book sibling, but nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description fully compensates by documenting all three parameters with meaningful behavior, defaults, and constraints. For example, 'limit' gets a range and default, 'typeset' explains the pandoc EPUB step, and 'import_format' clarifies attaching to book records.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('OCR'), a clear resource ('all scanned-PDF books in the library'), and a key mode ('resumable batches'). It also distinguishes itself from siblings by mentioning it runs the full ocr_book pipeline rather than just detection or a single book.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the intended invocation pattern: call repeatedly to work through the library, and per-batch caching makes interruptions safe. It does not explicitly contrast with single-book ocr_book or detect_scanned_books, but the batch scope and repeat-call guidance are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_bookA
Read a page of a book's plain text, converted on demand by calibre.
The first call converts the book (a copy in a cache directory — the
library is never modified); later calls reuse the cache. Returns the
chunk plus next_offset for paging and total_chars for context.
Args: book_id: The book's id (find it with search_books). offset: Character offset to start reading from (0 = beginning). limit: Maximum characters to return (1-40000, default 12000).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| book_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses on-demand conversion by calibre, cache reuse across calls, and explicitly states the library is never modified. It also mentions the return structure (chunk, next_offset, total_chars). This is strong behavioral context, though it omits potential error conditions or performance implications of first-time conversion.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured: a one-sentence purpose, a short behavioral note, and a bullet-point Args list. Every sentence earns its place, and the most important information is front-loaded. No filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity (one required integer parameter and two optional integers) and the existence of an output schema, the description is complete. It covers purpose, conversion/caching behavior, parameter semantics, and the key return fields needed for paging. There is nothing an agent needs to invoke and page through the book that is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully compensate, and it does. The Args block explains book_id ('find it with search_books'), offset as a character offset with the 0 meaning, and limit with range (1-40000) and default (12000). This adds substantive meaning beyond the raw schema properties.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Read a page of a book's plain text, converted on demand by calibre.' This clearly identifies what the tool does. It does not explicitly contrast with siblings like search_in_book, but the verb 'read' plus 'plain text' strongly implies the distinction, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: whenever you need to read a book's text content, with paging via offset/limit. It also tells the user to find book_id via search_books. However, it does not explicitly state when not to use it or mention alternatives such as search_in_book for searching within the text, so the guidance is present but not thorough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_booksA
Search the library with calibre's query language.
Args:
query: A calibre search expression, e.g. author:asimov,
title:"i robot", tags:python. An empty string matches everything.
limit: Maximum number of results (1-200, default 20).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses genuinely useful traits: the accepted query language, the empty-string-matches-everything edge case, and the limit bounds (1-200). It does not explicitly assert read-only safety, but 'Search' makes that evident.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line summary followed by a compact Args block. No filler sentences; every element carries information and the examples are illustrative without bloat. Clean and scannable, though the docstring format is a conventional choice rather than exceptional craft.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with an output schema present, the description covers query syntax, the empty-string edge case, and result limits — while the output schema handles return-value expectations. The only omissions are minor (behavior on malformed queries, result ordering), which do not undermine correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate — and it does so thoroughly. Each parameter gains real meaning: query is explained with three calibre expressions plus empty-string behavior, and limit is given a range and default beyond the schema's bare type and default. This exceeds the compensation baseline for low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Search'), resource ('the library'), and distinctive mechanism ('calibre's query language'). The library-level scope implicitly differentiates it from sibling search_in_book, and the query-language detail adds precision. However, it does not explicitly name sibling tools or draw the contrast, so it just misses the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete how-to guidance — three worked query examples and the note that an empty string matches everything — which implies when the tool applies. But it never states when to prefer this over list_books or search_in_book, and offers no exclusions or conditions. Usage is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_in_bookA
Full-text search inside book contents, with match snippets for quoting.
Searches the local text index; books must be indexed first with build_index (a one-time, cached step). Matching is substring-based, so Chinese and English queries both work without word segmentation.
Args: query: Text to find inside book contents. book_id: Optional book id to restrict the search to a single book. limit: Maximum number of matches to return (1-100, default 20).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| book_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It reveals that matching is substring-based, works for both Chinese and English without word segmentation, queries a local text index, and returns snippets for quoting. This is meaningful context beyond the schema, though it does not state whether the operation is read-only or describe error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly structured: a purpose line, prerequisite and matching semantics, then a clean Args list. Every sentence adds distinct value and there is no filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter search tool with an output schema, the description covers prerequisites, search semantics, param semantics, and output behavior (snippets for quoting). The output schema exists, so the description does not need to document return fields. Nothing essential is missing for an agent to invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description fully compensates by documenting all three parameters: query as the text to find, book_id as an optional single-book restriction, and limit with its range and default. It even adds the 1-100 constraint not present as an explicit schema boundary.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: "Full-text search inside book contents, with match snippets for quoting." This clearly differentiates the tool from siblings like search_books (metadata search) and read_book (reading content), so an agent can select it correctly without inspecting sibling schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear precondition: "books must be indexed first with build_index (a one-time, cached step)." This tells the agent when the tool is usable and names the prerequisite tool. It does not explicitly contrast with search_books or explain when not to use it, so it stops short of a perfect 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v0.3.0- First observed
add_books - First observed
build_index - First observed
detect_scanned_books - First observed
import_inbox - First observed
library_info - First observed
list_books - First observed
ocr_book - First observed
ocr_scanned_books - First observed
read_book - First observed
search_books - First observed
search_in_book
TDQS
Scored across 11 tools
Most tools target distinct actions: metadata search, browsing, full-text search, reading, and OCR are clearly separated. Minor overlap exists between list_books (which accepts a search filter) and search_books, and between ocr_scanned_books and ocr_book, but descriptions clarify the batch-vs-single distinction.
The majority of tools follow a clear verb_noun pattern like search_books, add_books, build_index, and read_book. A few deviations stand out: library_info uses noun_info, search_in_book uses a preposition, and ocr_book/ocr_scanned_books use an acronym as the verb, but the overall style remains predictable.
Eleven tools is a well-scoped size for a Calibre-focused server covering library browsing, metadata search, book import, reading, full-text search, and OCR. Each tool addresses a distinct workflow without feeling bloated or redundant.
The toolset covers the core workflows thoroughly: discovering, searching, adding, reading, indexing, and OCR-ing books. Obvious gaps like metadata editing, book deletion, or retrieving original format files are absent, but they appear to be outside the server's intended reading/OCR-oriented scope.
Maintenance
Related MCP Connectors
- LeafOAuthapp.readwithleaf
AI assistant integration for Leaf — track books, log reading sessions, and manage your library.
Academic literature search, retrieval, and private library management on top of OpenAlex.
The media memory layer for AI agents and their humans. Your AI client gets 29 tools to search your collection, add items, update ratings, preview music, and find patterns across everything you've read, watched, and listened to.
Federated search of books and papers, BibTeX/RIS citations, open-access retrieval and reading.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceProvides AI agents with research capabilities for local Calibre e-book libraries, including fulltext search across titles, ISBNs, and comments, plus structured excerpt retrieval from books.2GPL 3.0
- AlicenseBqualityCmaintenanceConnects AI agents to Calibre ebook libraries for searching, reading, and managing digital collections. It supports metadata updates, format conversion, and full-text content searches while providing granular permission controls for library access.723MIT
- AlicenseNot gradedqualityDmaintenanceEnables searching, reading, and managing a Calibre ebook library through natural language, with features like metadata search, full-text search, content extraction, and library management.113 npmApache 2.0
- AlicenseNot gradedqualityBmaintenanceBridges your Calibre e-book library with AI assistants via the Model Context Protocol, enabling natural-language library management, semantic search, RAG, and agentic workflows.45MIT