RAGFlow Claude MCP Server
RAGFlow Claude MCP 服务器
这是一个小型模型上下文协议 (MCP) 服务器,用于将 Claude Desktop(以及其他 MCP 客户端)连接到 RAGFlow 实例。它将 RAGFlow REST API 公开为一组工具,以便 LLM 可以查询知识库并将文档块提取到其上下文中。
这是我为自己的研发工作编写的个人使用软件。它并非没有 bug,代码也不够优雅。但它能满足我的需求。
功能
直接检索:从 RAGFlow 的
/retrieval端点提取带有相似度分数的原始文档块。多知识库搜索:单个查询可以同时命中多个知识库。
DSPy 查询深化:可选的迭代查询优化(使用 LLM 分析中间结果并重写查询)。
~~重排序 (Reranking)~~ — 目前在 RAGFlow 端存在问题,请参阅 已知问题。
可调结果控制:
page_size、similarity_threshold、top_k、分页。文档过滤:将结果限制在数据集内的一个文档中(模糊名称匹配)。
按名称(不区分大小写,模糊匹配)而不是 ID 查找数据集。
当您的 RAGFlow 位于 Cloudflare Zero Trust 之后时,支持其身份验证。
Related MCP server: RAGBrain MCP
安装
克隆:
git clone https://github.com/norandom/ragflow-claude-desktop-local-mcp cd ragflow-claude-desktop-local-mcp安装:
# On macOS, install DSPy first to dodge build issues: pip install git+https://github.com/stanfordnlp/dspy.git uv install配置:复制示例并填写您的 RAGFlow 详细信息。
cp config.json.sample config.json键:
RAGFLOW_BASE_URL:例如http://your-ragflow-server:9380RAGFLOW_API_KEY:您的 RAGFlow API 密钥RAGFLOW_DEFAULT_RERANK:重排序模型(默认rerank-multilingual-v3.0)CF_ACCESS_CLIENT_ID(可选):Cloudflare Zero Trust 服务令牌 IDCF_ACCESS_CLIENT_SECRET(可选):Cloudflare Zero Trust 服务令牌密钥DSPY_MODEL:DSPy LM(默认openai/gpt-4o-mini)OPENAI_API_KEY:DSPy 深化所需
Cloudflare Zero Trust
如果您的 RAGFlow 位于 Cloudflare Zero Trust 之后,请从仪表板获取服务令牌并将其添加到 config.json 中:
{
"CF_ACCESS_CLIENT_ID": "your-client-id.access",
"CF_ACCESS_CLIENT_SECRET": "your-client-secret"
}当两者都设置时,每个 API 请求都会携带 CF-Access-Client-Id 和 CF-Access-Client-Secret 标头。无需更改代码。
Claude Desktop 配置
{
"mcpServers": {
"ragflow": {
"command": "uv",
"args": [
"run",
"--directory",
"/path/to/ragflow-claude-desktop-local-mcp",
"ragflow-claude-mcp"
]
}
}
}工具
ragflow_retrieval_by_name(我最常用的工具)
按名称检索一个或多个数据集中的块。返回带有相似度分数的原始块。
参数:
dataset_names(必需)— 列表,例如["BASF", "Quant Literature"]query(必需)document_name(可选)— 限制为一个文档;模糊匹配top_k(可选,默认 1024)— 向量候选数similarity_threshold(可选,默认 0.2)— 0.0–1.0page(可选,默认 1)page_size(可选,默认 10)use_rerank(可选,默认 false)— 目前在上游已损坏,请参阅已知问题deepening_level(可选,默认 0)— DSPy 优化,0–3
ragflow_retrieval
形状相同,但使用 dataset_ids: List[str] 而不是名称。
多知识库搜索
您可以在一次调用中搜索多个知识库。请确保它们共享一个嵌入模型 — 混合不兼容的嵌入会降低相关性分数。
Use ragflow_retrieval_by_name with dataset_names ["Finance Reports", "Legal Documents"] and query "Summarize the key financial risks and compliance requirements for new market entry."ragflow_list_datasets
列出 RAGFlow 实例上的所有知识库。无参数。在内部遍历所有页面。
ragflow_list_documents
列出数据集中的文档。遍历所有页面。
dataset_id(必需)
ragflow_get_chunks
返回一个文档的块(带有引用)。
dataset_id(必需)document_id(必需)
ragflow_list_sessions
显示每个数据集的活动聊天会话。无参数。
ragflow_list_documents_by_name
按名称查找并列出数据集中的文档。
dataset_name(必需)
ragflow_reset_session
删除数据集的聊天会话。
dataset_id(必需)
调整检索
检索工具提供三个调节旋钮:
page_size— 每页块数(默认 10)。similarity_threshold— 丢弃低于此分数的块(默认 0.2)。top_k— 过滤前向量搜索的池大小(默认 1024)。
对我有效的一些起始设置:
更广泛的召回:
page_size=15,similarity_threshold=0.15。更高的精度:
page_size=5,similarity_threshold=0.4。深度研究:
page_size=20,similarity_threshold=0.1,deepening_level=1。困难查询:
deepening_level=2。速度:保持
deepening_level=0并跳过重排序。
示例
按名称进行基本检索:
Use ragflow_retrieval_by_name with dataset_names ["BASF"] and query "What is BASF's latest income statement? Revenue, operating income, net income, and other key figures."限制为一个文档:
Use ragflow_retrieval_by_name with dataset_names ["BASF"], document_name "annual_report_2023", and query "What were the key financial highlights for 2023?"文档名称进行模糊匹配 — "annual" 将匹配 annual_report_2023.pdf 和 annual_report_2024.pdf。当有多个匹配项时,服务器会选择最近的一个,并在响应元数据中列出其他选项。
针对棘手查询的 DSPy 深化:
Use ragflow_retrieval_by_name with dataset_names ["Quant Literature"], query "what is a volatility clock", deepening_level 2.多页:
Use ragflow_retrieval_by_name with dataset_names ["BASF"], query "BASF business segments", page_size 10, page 2.列出可用内容:
Use ragflow_list_datasets.Use ragflow_list_documents_by_name with dataset_name "BASF".提取特定块:
Use ragflow_get_chunks with dataset_id "43066ee0599411f089787a39c10de57b" and document_id "d74a1c105a3311f09fc94a0fcd8b7722".更大的提示词
一些关于我如何从 Claude Desktop 驱动它的示例。
财务深度分析:
Help me analyse BASF's recent financials.
1. Use ragflow_retrieval_by_name to search ["BASF"] for the latest income statement
(revenue, operating income, net income). Use page_size 15,
similarity_threshold 0.15, deepening_level 1.
2. Then run ragflow_retrieval_by_name again for the cash flow statement,
page_size 10, similarity_threshold 0.2.
3. Finally look for year-over-year changes with page_size 12,
similarity_threshold 0.18.多语言研究:
Use ragflow_retrieval_by_name with dataset_names ["BASF"],
query "Was sind die wichtigsten Geschäftsbereiche von BASF?",
deepening_level 2.DSPy 会检测查询语言并相应地进行优化。我已将其用于德语、英语和混合语言查询。只要底层文档包含这些语言的内容,它就能正常工作。
文档过滤研究:
1. Use ragflow_list_documents_by_name with dataset_name "BASF" to see what's in there.
2. Use ragflow_retrieval_by_name with dataset_names ["BASF"],
document_name "sustainability_report", query "carbon neutrality goals",
page_size 15, deepening_level 1.
3. Follow up with document_name "annual_report_2023" and
query "environmental investments".跨知识库查询:
Use ragflow_retrieval_by_name with dataset_names ["BASF", "Industry Reports"],
query "chemical industry sustainability benchmarks",
page_size 12, deepening_level 1.DSPy 深化是如何工作的
deepening_level 在检索之上运行一个 LLM 驱动的优化循环:
0:无深化(默认)。
1:一次优化传递。
2:两次带有差距分析的传递。
3:三次以上传递加上结果合并。
每次传递:执行搜索,总结前几个结果,询问 LLM 缺少什么,生成新查询,运行该查询。响应元数据包括原始查询、每个优化后的查询以及每一步的推理。
DSPy 需要:
DSPY_MODEL—openai/gpt-4o-mini效果很好OPENAI_API_KEY
重排序(目前已损坏)
当工作正常时,重排序会用重排序模型的分数替换向量余弦分数(根据我的经验,相关性通常提高 10-30%)。RAGFlow 目前有一个已知 bug,即 use_rerank=true 会产生:
UnsupportedProtocol: Request URL is missing an 'http://' or 'https://' protocol
因此,在上游问题修复之前,请保持 use_rerank=false。标准向量检索可以正常工作。
数据集查找是如何工作的
不区分大小写的名称匹配。
部分名称的模糊匹配。
数据集会被缓存以进行名称查找;缓存未命中会触发刷新。
如果查找失败,错误信息会包含可用的数据集名称,以便您知道实际存在的内容。
文档匹配
当您传递 document_name 时:
精确匹配优先,然后是“以...开头”,然后是“包含”,最后是部分匹配。
在平局的情况下,最近更新的文档获胜。
名称中包含
2024、2023、latest、current或new的文档会获得少量分数奖励。所有匹配项都会返回在响应元数据中,以便您可以使用更具体的名称重新发出请求。
错误处理
针对以下情况提供合理的错误消息:API 错误、缺少数据集、无法访问 RAGFlow、会话中断、输入无效和配置问题。敏感值在日志中会被屏蔽。
环境变量
RAGFLOW_BASE_URL— 覆盖配置文件。代码中的默认值:http://192.168.122.93:9380(这是我的本地实例)。RAGFLOW_API_KEY— 必需。
开发
直接运行服务器:
uv run ragflow-claude-mcp它像 MCP 服务器一样在 stdio 上监听。
开发依赖:
uv install --extra dev这会安装 pytest + asyncio/mock/cov 插件。
测试:
uv run pytest
uv run pytest --cov=src --cov-report=html --cov-report=term
uv run pytest tests/test_server.py
uv run pytest -v覆盖率约为 44%,22/23 个测试通过(一个因间歇性 CI 故障被跳过)。测试涵盖服务器初始化、RAGFlow API 集成、DSPy 深化、OpenAI/OpenRouter 配置分支和配置加载。
实现说明
检索 API 是服务器实际依赖的唯一 RAGFlow 表面。没有助手/聊天依赖,没有服务器端提示词配置 — 只是返回块。更容易推理,更容易调试。
故障排除
“Dataset not found”:运行
ragflow_list_datasets查看实际存在的内容。连接错误:仔细检查
RAGFLOW_BASE_URL和RAGFLOW_API_KEY。服务器无法启动:
uv install是否真的完成了?需要原始块:使用
ragflow_retrieval_by_name/ragflow_retrieval。会话卡住:运行
ragflow_list_sessions然后ragflow_reset_session。Cloudflare 403:确认
CF_ACCESS_CLIENT_ID/CF_ACCESS_CLIENT_SECRET与 Zero Trust 应用上的活动服务令牌匹配。
已知问题
重排序在上游已损坏
use_rerank=true 会报错 UnsupportedProtocol: Request URL is missing an 'http://' or 'https://' protocol。这是 RAGFlow 端的缺陷。解决方法:保持关闭。我正在关注 RAGFlow 仓库以获取修复。
贡献
仅限 PR — main 分支受保护。提交必须经过 SSH 签名。
Fork。
git checkout -b feature/your-thing。进行更改,编写清晰的提交信息。
推送到您的 Fork。
针对
main打开 PR。
PR 会自动运行 TruffleHog — 不要包含密钥、令牌或机密。有关详细信息,请参阅 CONTRIBUTING.md。
Available Tools
8 toolsragflow_get_chunksC
Get chunks with references from a specific document
| Name | Required | Description | Default |
|---|---|---|---|
| dataset_id | Yes | ID of the dataset | |
| document_id | Yes | ID of the document to get chunks from |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, but it only states a simple data retrieval. It omits important traits like pagination, rate limits, authentication, or potential side effects, leaving the agent under-informed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words, but it is overly brief and lacks important details. Conciseness is not valuable at the expense of completeness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description should explain what 'chunks with references' means and the format of the return value. It does not, leaving the agent with insufficient context for a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters. The description does not add meaning beyond what the schema provides, earning a baseline score of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get') and the resource ('chunks with references from a specific document'), effectively distinguishing it from sibling tools like listing datasets or retrieval. However, 'references' could be more explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as retrieval tools. There is no mention of prerequisites, context, or situations where this tool is inappropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ragflow_list_datasetsA
List all available datasets/knowledge bases in RAGFlow
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It states 'list all available' but omits details like pagination, ordering, or side effects. Adequate but minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with action. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While sufficient for a zero-parameter listing tool, the lack of output schema leaves the agent uninformed about the response structure, which could be improved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist, and schema coverage is 100%. Baseline 4 applies as the description adds no parameter info, which is acceptable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List'), the resource ('all available datasets/knowledge bases'), and distinguishes it from siblings which deal with chunks, documents, and sessions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus siblings. The description only states what it does, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ragflow_list_documentsC
List documents in a specific dataset
| Name | Required | Description | Default |
|---|---|---|---|
| dataset_id | Yes | ID of the dataset to list documents from |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses no behavioral traits such as read-only nature, pagination, error handling, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and front-loaded, stating the core purpose in a single phrase with no extraneous content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description fails to cover return format, pagination, or error conditions, even for a simple list tool it feels incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a description for the single parameter. The tool description adds no additional meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List', resource 'documents', and context 'in a specific dataset'. It distinguishes from siblings such as ragflow_list_datasets (lists datasets) and ragflow_get_chunks (gets chunks).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use or not use this tool versus alternatives. The description only states the basic action without any contextual hints or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ragflow_list_documents_by_nameC
List documents in a dataset by dataset name
| Name | Required | Description | Default |
|---|---|---|---|
| dataset_name | Yes | Name of the dataset/knowledge base to list documents from |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description is the sole source for behavioral clues. It implies a read operation but does not disclose details such as pagination, authentication requirements, rate limits, or what the response looks like. Minimal transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with key action and resource. Efficient but could benefit from additional context without being verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description should hint at what the returned list contains (e.g., document names, IDs, metadata). It only states what it does, not what the agent gets back. Missing return value details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description only restates the parameter's purpose ('by dataset name') which is already described in the schema. Adds no extra meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action (List), resource (documents), and filter (by dataset name). It is specific and suggests the tool's scope, but does not explicitly differentiate from the sibling tool 'ragflow_list_documents' which likely lists documents without a dataset name filter.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus the sibling 'ragflow_list_documents', which might list all documents or use different criteria. The description does not mention alternatives or conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ragflow_list_sessionsB
List active chat sessions for all datasets
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It only says 'List active chat sessions' but does not explain what 'active' means, any side effects, or limitations. Minimal behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, direct sentence with no wasted words. It is front-loaded with the key action and resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no parameters, the description lacks details on output format, pagination, or what constitutes an active session. Without output schema or annotations, the description is insufficient for complete understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema is empty (0 parameters), so schema coverage is 100%. The description adds meaning by specifying the resource and scope, which is beyond the empty schema. Baseline 3, but the context provided justifies a higher score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List' and the resource 'active chat sessions' with scope 'for all datasets', distinguishing it from sibling tools like ragflow_list_datasets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool over alternatives like ragflow_list_datasets or ragflow_reset_session. The description only states what it does without usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ragflow_reset_sessionB
Reset/clear the chat session for a specific dataset
| Name | Required | Description | Default |
|---|---|---|---|
| dataset_id | Yes | ID of the dataset to reset session for |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided. Description merely states the action without disclosing side effects (e.g., whether session history is deleted permanently, if it affects other datasets, or if confirmation is required).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, 10 words, no redundancy. Front-loaded with verb and resource. Efficiently communicates the core function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple reset action with one parameter and no output schema, but lacks behavioral details that would help the agent understand consequences. Could mention that the session is cleared without confirmation or return value.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for one parameter. Description mirrors the schema's description ('ID of the dataset to reset session for') without adding new meaning or constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action (reset/clear) and the resource (chat session for a specific dataset). It is distinct from sibling tools which are for listing or retrieval, not mutation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. Does not mention prerequisites, conditions, or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ragflow_retrievalB
Retrieve document chunks directly from RAGFlow datasets using the retrieval API. Returns raw chunks with similarity scores.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number for pagination. Defaults to 1. | |
| query | Yes | Search query or question | |
| top_k | No | Number of chunks for vector cosine computation. Defaults to 1024. | |
| page_size | No | Number of chunks per page. Defaults to 10. | |
| use_rerank | No | Whether to enable reranking for better result quality. Default: false (uses vector similarity only). | |
| dataset_ids | Yes | List of IDs of the datasets/knowledge bases to search | |
| document_name | No | Optional document name to filter results to specific document | |
| deepening_level | No | Level of DSPy query refinement (0-3). 0=none, 1=basic refinement, 2=gap analysis, 3=full optimization. Default: 0 | |
| similarity_threshold | No | Minimum similarity score for chunks (0.0 to 1.0). Defaults to 0.2. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must disclose behavioral traits like whether the tool is read-only, permission requirements, or pagination behavior. It only says 'Returns raw chunks' and does not address these aspects, leaving the agent with incomplete understanding of its side effects or constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is succinct: two sentences that convey the core function and output without extraneous words. It is front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 9 parameters and no output schema, the description should provide more context on how to use parameters like deepening_level or use_rerank, and what the returned chunks contain. It states 'raw chunks with similarity scores' but lacks detail on the structure of the response, which is necessary for an agent to process the output correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description does not add meaning beyond what the parameter descriptions already provide (e.g., page, top_k). It mentions 'similarity scores' but does not clarify how parameters like similarity_threshold relate to the output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Retrieve' and the resource 'document chunks' from RAGFlow datasets, and specifies the output as 'raw chunks with similarity scores'. However, it does not explicitly differentiate from sibling tools like ragflow_retrieval_by_name, which likely performs a similar function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as ragflow_get_chunks or ragflow_retrieval_by_name. It merely states what the tool does, without context on prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ragflow_retrieval_by_nameB
Retrieve document chunks by dataset names using the retrieval API. Returns raw chunks with similarity scores.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number for pagination. Defaults to 1. | |
| query | Yes | Search query or question | |
| top_k | No | Number of chunks for vector cosine computation. Defaults to 1024. | |
| page_size | No | Number of chunks per page. Defaults to 10. | |
| use_rerank | No | Whether to enable reranking for better result quality. Default: false (uses vector similarity only). | |
| dataset_names | Yes | List of names of the datasets/knowledge bases to search (e.g., ['BASF', 'Legal']) | |
| document_name | No | Optional document name to filter results to specific document | |
| deepening_level | No | Level of DSPy query refinement (0-3). 0=none, 1=basic refinement, 2=gap analysis, 3=full optimization. Default: 0 | |
| similarity_threshold | No | Minimum similarity score for chunks (0.0 to 1.0). Defaults to 0.2. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It mentions return type (raw chunks with similarity scores) but lacks information on side effects, permissions, rate limits, or destructive potential. 'Retrieve' implies read-only but is not explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loading the purpose. It is efficient but could be slightly more structured without adding verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 9 parameters and no output schema, the description is sparse. It omits details on pagination, reranking, deepening_level, and similarity_threshold behavior, leaving the agent to rely solely on the schema for context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds minimal meaning beyond the schema, only briefly noting retrieval by dataset names and return format. No parameter interaction hints are provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (retrieve), resource (document chunks), and distinguishing parameter (by dataset names). It differentiates from siblings like ragflow_retrieval which likely uses different criteria.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives. The description implies usage with dataset names but does not mention exclusions or compare to ragflow_retrieval or other search methods.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
ragflow_get_chunks - First observed
ragflow_list_datasets - First observed
ragflow_list_documents - First observed
ragflow_list_documents_by_name - First observed
ragflow_list_sessions - First observed
ragflow_reset_session - First observed
ragflow_retrieval - First observed
ragflow_retrieval_by_name
TDQS
Scored across 8 tools
Most tools have distinct purposes, but ragflow_list_documents and ragflow_retrieval each have an alternative by-name variant, which could cause confusion if descriptions are not heeded. However, descriptions clarify the difference between ID-based and name-based operations, keeping overlap minimal.
All tools follow a consistent verb_noun pattern with snake_case and the 'ragflow_' prefix. Variations like '_by_name' are systematic and predictable, enhancing readability for agents.
With 8 tools, the set is well-scoped for a knowledge base retrieval server. Each tool serves a clear function, and the count is neither too sparse nor overwhelming for the intended purpose.
The tool surface covers listing datasets, listing documents, retrieving chunks, and managing chat sessions. It lacks create/update/delete operations, but given the likely read-heavy focus of the server, these gaps are acceptable and do not impede the primary retrieval workflow.
Maintenance
Related MCP Connectors
Connect your team's living knowledge base — docs, data, issues, CRM — to Claude and ChatGPT.
Cloud or self-hosted knowledge for AI agents: hybrid search, reranking, GraphRAG, scoped MCP tools.
Ingest, manage, and retrieve documents for RAG-powered AI applications
Search your knowledge bases from any AI assistant using hybrid RAG.
Related MCP Servers
- FlicenseAqualityDmaintenanceIntegrates R2R (Retrieval-Augmented Generation) with Claude Desktop, enabling semantic search across knowledge bases and RAG-based question answering with support for vector, graph, web, and document search.2-
- AlicenseAqualityFmaintenanceConnects Claude Desktop to a RAGBrain knowledge base to enable semantic search, document retrieval, and namespace management. It allows users to browse collections, discover documents by topic, and access full text content through natural language.5MIT
- AlicenseAqualityDmaintenanceProvides semantic search capabilities by connecting Claude Desktop to a Cloudflare Workers backend powered by Vectorize. It enables natural language querying of knowledge bases using vector similarity and edge-based embedding generation.2MIT
- AlicenseNot gradedqualityDmaintenanceEnables semantic retrieval and knowledge base management through the RAGFlow API, including dataset, document, chunk, chat, and graph operations.5MIT