RAGFlow Claude MCP Server
RAGFlow Claude MCP サーバー
RAGFlow インスタンスを Claude Desktop (およびその他の MCP クライアント) に接続するための小さな Model Context Protocol (MCP) サーバーです。RAGFlow REST API をいくつかのツールとして公開し、LLM がナレッジベースをクエリしてドキュメントのチャンクをコンテキストに取り込めるようにします。
これは私自身の研究開発のために書いた個人用ソフトウェアです。バグがないわけではなく、コードも綺麗ではありません。私が必要とする用途には機能しています。
機能
直接取得: RAGFlow の
/retrievalエンドポイントから類似度スコア付きの生のドキュメントチャンクを取得します。マルチ KB 検索: 1 つのクエリで複数のナレッジベースを同時に検索できます。
DSPy クエリ深化: オプションの反復的なクエリ洗練(LLM を使用して中間結果を分析し、クエリを書き換えます)。
~~再ランキング~~ — 現在 RAGFlow 側で壊れています。既知の問題 を参照してください。
調整可能な結果制御:
page_size、similarity_threshold、top_k、ページネーション。ドキュメントフィルタ: データセット内の 1 つのドキュメントに結果を制限します(あいまいな名前マッチング)。
ID ではなく名前によるデータセット検索(大文字小文字を区別しない、あいまい検索)。
RAGFlow が Cloudflare Zero Trust の背後にある場合の認証。
Related MCP server: RAGBrain MCP
インストール
クローン:
git clone https://github.com/norandom/ragflow-claude-desktop-local-mcp cd ragflow-claude-desktop-local-mcpインストール:
# On macOS, install DSPy first to dodge build issues: pip install git+https://github.com/stanfordnlp/dspy.git uv install設定: サンプルをコピーし、RAGFlow の詳細を入力します。
cp config.json.sample config.jsonキー:
RAGFLOW_BASE_URL: 例:http://your-ragflow-server:9380RAGFLOW_API_KEY: RAGFlow API キーRAGFLOW_DEFAULT_RERANK: 再ランキングモデル(デフォルトrerank-multilingual-v3.0)CF_ACCESS_CLIENT_ID(オプション): Cloudflare Zero Trust サービス・トークン IDCF_ACCESS_CLIENT_SECRET(オプション): Cloudflare Zero Trust サービス・トークン・シークレットDSPY_MODEL: DSPy LM(デフォルトopenai/gpt-4o-mini)OPENAI_API_KEY: DSPy の深化に必要
Cloudflare Zero Trust
RAGFlow が Cloudflare Zero Trust の背後にある場合は、ダッシュボードからサービス・トークンを取得し、config.json に追加します:
{
"CF_ACCESS_CLIENT_ID": "your-client-id.access",
"CF_ACCESS_CLIENT_SECRET": "your-client-secret"
}両方が設定されると、すべての API リクエストが CF-Access-Client-Id および CF-Access-Client-Secret ヘッダーと共に送信されます。コードの変更は不要です。
Claude Desktop 設定
{
"mcpServers": {
"ragflow": {
"command": "uv",
"args": [
"run",
"--directory",
"/path/to/ragflow-claude-desktop-local-mcp",
"ragflow-claude-mcp"
]
}
}
}ツール
ragflow_retrieval_by_name (私が最もよく使うもの)
名前で 1 つ以上のデータセットにわたってチャンクを取得します。類似度スコア付きの生のチャンクを返します。
パラメータ:
dataset_names(必須) — リスト、例:["BASF", "Quant Literature"]query(必須)document_name(オプション) — 1 つのドキュメントに制限。あいまい一致top_k(オプション、デフォルト 1024) — ベクトル候補similarity_threshold(オプション、デフォルト 0.2) — 0.0–1.0page(オプション、デフォルト 1)page_size(オプション、デフォルト 10)use_rerank(オプション、デフォルト false) — 現在アップストリームで壊れています。既知の問題を参照deepening_level(オプション、デフォルト 0) — DSPy 洗練、0–3
ragflow_retrieval
形状は同じですが、名前の代わりに dataset_ids: List[str] を受け取ります。
マルチ KB 検索
1 回の呼び出しで複数のナレッジベースを検索できます。それらが同じ埋め込みモデルを共有していることを確認してください。互換性のない埋め込みを混在させると、関連性スコアが低下します。
Use ragflow_retrieval_by_name with dataset_names ["Finance Reports", "Legal Documents"] and query "Summarize the key financial risks and compliance requirements for new market entry."ragflow_list_datasets
RAGFlow インスタンス上のすべてのナレッジベースをリストします。パラメータなし。内部的にすべてのページを走査します。
ragflow_list_documents
データセット内のドキュメントをリストします。すべてのページを走査します。
dataset_id(必須)
ragflow_get_chunks
1 つのドキュメントのチャンク(参照付き)を返します。
dataset_id(必須)document_id(必須)
ragflow_list_sessions
データセットごとのアクティブなチャットセッションを表示します。パラメータなし。
ragflow_list_documents_by_name
名前で検索されたデータセット内のドキュメントをリストします。
dataset_name(必須)
ragflow_reset_session
データセットのチャットセッションを破棄します。
dataset_id(必須)
取得の調整
取得ツールには 3 つのノブがあります:
page_size— ページあたりのチャンク数(デフォルト 10)。similarity_threshold— このスコアを下回るチャンクを除外(デフォルト 0.2)。top_k— フィルタリング前のベクトル検索のプールサイズ(デフォルト 1024)。
私が使用している開始点:
より広いリコール:
page_size=15,similarity_threshold=0.15。厳密な精度:
page_size=5,similarity_threshold=0.4。詳細な調査:
page_size=20,similarity_threshold=0.1,deepening_level=1。難しいクエリ:
deepening_level=2。速度:
deepening_level=0に保ち、再ランキングをスキップ。
例
名前による基本的な取得:
Use ragflow_retrieval_by_name with dataset_names ["BASF"] and query "What is BASF's latest income statement? Revenue, operating income, net income, and other key figures."1 つのドキュメントに制限:
Use ragflow_retrieval_by_name with dataset_names ["BASF"], document_name "annual_report_2023", and query "What were the key financial highlights for 2023?"ドキュメント名はあいまい一致します — "annual" は annual_report_2023.pdf と annual_report_2024.pdf にヒットします。複数一致する場合、サーバーは最新のものを選択し、応答メタデータに代替案をリストします。
難しいクエリに対する DSPy の深化:
Use ragflow_retrieval_by_name with dataset_names ["Quant Literature"], query "what is a volatility clock", deepening_level 2.複数ページ:
Use ragflow_retrieval_by_name with dataset_names ["BASF"], query "BASF business segments", page_size 10, page 2.利用可能なものをリスト:
Use ragflow_list_datasets.Use ragflow_list_documents_by_name with dataset_name "BASF".特定のチャンクを取得:
Use ragflow_get_chunks with dataset_id "43066ee0599411f089787a39c10de57b" and document_id "d74a1c105a3311f09fc94a0fcd8b7722".より大きなプロンプト
Claude Desktop からこれをどのように操作するかの例。
財務の詳細調査:
Help me analyse BASF's recent financials.
1. Use ragflow_retrieval_by_name to search ["BASF"] for the latest income statement
(revenue, operating income, net income). Use page_size 15,
similarity_threshold 0.15, deepening_level 1.
2. Then run ragflow_retrieval_by_name again for the cash flow statement,
page_size 10, similarity_threshold 0.2.
3. Finally look for year-over-year changes with page_size 12,
similarity_threshold 0.18.多言語調査:
Use ragflow_retrieval_by_name with dataset_names ["BASF"],
query "Was sind die wichtigsten Geschäftsbereiche von BASF?",
deepening_level 2.DSPy はクエリ言語を検出し、それに応じて洗練します。ドイツ語、英語、混合言語のクエリで使用しました。基盤となるドキュメントにそれらの言語のコンテンツが含まれている限り機能します。
ドキュメントフィルタリングされた調査:
1. Use ragflow_list_documents_by_name with dataset_name "BASF" to see what's in there.
2. Use ragflow_retrieval_by_name with dataset_names ["BASF"],
document_name "sustainability_report", query "carbon neutrality goals",
page_size 15, deepening_level 1.
3. Follow up with document_name "annual_report_2023" and
query "environmental investments".KB 横断クエリ:
Use ragflow_retrieval_by_name with dataset_names ["BASF", "Industry Reports"],
query "chemical industry sustainability benchmarks",
page_size 12, deepening_level 1.DSPy の深化の仕組み
deepening_level は、取得の上に LLM 主導の洗練ループを実行します:
0: 深化なし(デフォルト)。
1: 1 回の洗練パス。
2: ギャップ分析を伴う 2 回のパス。
3: 3 回以上のパスと結果の統合。
各パス: 検索を実行し、上位結果を要約し、LLM に何が欠けているかを尋ね、新しいクエリを生成し、それを実行します。応答メタデータには、元のクエリ、すべての洗練されたクエリ、および各ステップでの推論が含まれます。
DSPy に必要なもの:
DSPY_MODEL—openai/gpt-4o-miniが正常に動作しますOPENAI_API_KEY
再ランキング(現在壊れています)
動作している場合、再ランキングはベクトルコサインスコアを再ランキングモデルのスコアに置き換えます(私の経験では、関連性が通常 10~30% 向上します)。RAGFlow には現在、use_rerank=true が以下を生成するという既知のバグがあります:
UnsupportedProtocol: Request URL is missing an 'http://' or 'https://' protocol
そのため、アップストリームの問題が修正されるまで use_rerank=false のままにしてください。標準のベクトル取得は通常通り機能します。
データセット検索の仕組み
大文字小文字を区別しない名前マッチング。
部分名に対するあいまい一致。
データセットは名前検索のためにキャッシュされます。キャッシュミスは更新をトリガーします。
検索が失敗した場合、エラーには利用可能なデータセット名が含まれるため、実際に何が存在したかを知ることができます。
ドキュメントマッチング
document_name を渡す場合:
完全一致が優先され、次に「で始まる」、次に「含む」、次に部分一致となります。
同点の場合、最近更新されたドキュメントが優先されます。
2024、2023、latest、current、またはnewを含む名前には小さなスコアボーナスが与えられます。すべての一致は応答メタデータで返されるため、より具体的な名前で再発行できます。
エラー処理
API エラー、データセットの欠落、到達不能な RAGFlow、壊れたセッション、無効な入力、設定の問題に対する妥当なエラーメッセージ。機密値はログで編集されます。
環境変数
RAGFLOW_BASE_URL— 設定ファイルを上書きします。コード内のデフォルト:http://192.168.122.93:9380(私のローカルインスタンス)。RAGFLOW_API_KEY— 必須。
開発
サーバーを直接実行:
uv run ragflow-claude-mcpMCP サーバーと同様に、stdio でリッスンします。
開発依存関係:
uv install --extra devこれで pytest + asyncio/mock/cov プラグインがインストールされます。
テスト:
uv run pytest
uv run pytest --cov=src --cov-report=html --cov-report=term
uv run pytest tests/test_server.py
uv run pytest -vカバレッジは約 44% で、22/23 のテストが合格しています(1 つは断続的な CI の不安定さのためスキップされています)。テストはサーバーの初期化、RAGFlow API 統合、DSPy の深化、OpenAI/OpenRouter 設定ブランチ、および設定の読み込みをカバーしています。
実装ノート
取得 API は、サーバーが実際に依存している唯一の RAGFlow サーフェスです。アシスタント/チャットの依存関係はなく、サーバー側のプロンプト設定もありません。チャンクが返されるだけです。推論しやすく、デバッグも容易です。
トラブルシューティング
「Dataset not found」:
ragflow_list_datasetsを実行して、実際に何があるかを確認してください。接続エラー:
RAGFLOW_BASE_URLとRAGFLOW_API_KEYを再確認してください。サーバーが起動しない:
uv installは完了しましたか?生のチャンクが必要: それは
ragflow_retrieval_by_name/ragflow_retrievalです。セッションがスタックしている:
ragflow_list_sessionsを実行してからragflow_reset_sessionを実行してください。Cloudflare 403:
CF_ACCESS_CLIENT_ID/CF_ACCESS_CLIENT_SECRETが Zero Trust アプリのアクティブなサービス・トークンと一致していることを確認してください。
既知の問題
再ランキングがアップストリームで壊れています
use_rerank=true は UnsupportedProtocol: Request URL is missing an 'http://' or 'https://' protocol というエラーになります。これは RAGFlow 側の欠陥です。回避策: オフのままにしてください。修正のために RAGFlow リポジトリを監視しています。
貢献
PR のみ — main は保護されています。コミットは SSH 署名されている必要があります。
フォークします。
git checkout -b feature/your-thing。変更を行い、明確なコミットメッセージを書きます。
フォークにプッシュします。
mainに対して PR を開きます。
PR は自動的に TruffleHog を実行します。キー、トークン、シークレットを含めないでください。詳細については CONTRIBUTING.md を参照してください。
Available Tools
8 toolsragflow_get_chunksC
Get chunks with references from a specific document
| Name | Required | Description | Default |
|---|---|---|---|
| dataset_id | Yes | ID of the dataset | |
| document_id | Yes | ID of the document to get chunks from |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, but it only states a simple data retrieval. It omits important traits like pagination, rate limits, authentication, or potential side effects, leaving the agent under-informed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words, but it is overly brief and lacks important details. Conciseness is not valuable at the expense of completeness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description should explain what 'chunks with references' means and the format of the return value. It does not, leaving the agent with insufficient context for a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters. The description does not add meaning beyond what the schema provides, earning a baseline score of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get') and the resource ('chunks with references from a specific document'), effectively distinguishing it from sibling tools like listing datasets or retrieval. However, 'references' could be more explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as retrieval tools. There is no mention of prerequisites, context, or situations where this tool is inappropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ragflow_list_datasetsA
List all available datasets/knowledge bases in RAGFlow
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It states 'list all available' but omits details like pagination, ordering, or side effects. Adequate but minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with action. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While sufficient for a zero-parameter listing tool, the lack of output schema leaves the agent uninformed about the response structure, which could be improved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist, and schema coverage is 100%. Baseline 4 applies as the description adds no parameter info, which is acceptable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List'), the resource ('all available datasets/knowledge bases'), and distinguishes it from siblings which deal with chunks, documents, and sessions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus siblings. The description only states what it does, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ragflow_list_documentsC
List documents in a specific dataset
| Name | Required | Description | Default |
|---|---|---|---|
| dataset_id | Yes | ID of the dataset to list documents from |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses no behavioral traits such as read-only nature, pagination, error handling, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and front-loaded, stating the core purpose in a single phrase with no extraneous content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description fails to cover return format, pagination, or error conditions, even for a simple list tool it feels incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a description for the single parameter. The tool description adds no additional meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List', resource 'documents', and context 'in a specific dataset'. It distinguishes from siblings such as ragflow_list_datasets (lists datasets) and ragflow_get_chunks (gets chunks).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use or not use this tool versus alternatives. The description only states the basic action without any contextual hints or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ragflow_list_documents_by_nameC
List documents in a dataset by dataset name
| Name | Required | Description | Default |
|---|---|---|---|
| dataset_name | Yes | Name of the dataset/knowledge base to list documents from |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description is the sole source for behavioral clues. It implies a read operation but does not disclose details such as pagination, authentication requirements, rate limits, or what the response looks like. Minimal transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with key action and resource. Efficient but could benefit from additional context without being verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description should hint at what the returned list contains (e.g., document names, IDs, metadata). It only states what it does, not what the agent gets back. Missing return value details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description only restates the parameter's purpose ('by dataset name') which is already described in the schema. Adds no extra meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action (List), resource (documents), and filter (by dataset name). It is specific and suggests the tool's scope, but does not explicitly differentiate from the sibling tool 'ragflow_list_documents' which likely lists documents without a dataset name filter.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus the sibling 'ragflow_list_documents', which might list all documents or use different criteria. The description does not mention alternatives or conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ragflow_list_sessionsB
List active chat sessions for all datasets
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It only says 'List active chat sessions' but does not explain what 'active' means, any side effects, or limitations. Minimal behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, direct sentence with no wasted words. It is front-loaded with the key action and resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no parameters, the description lacks details on output format, pagination, or what constitutes an active session. Without output schema or annotations, the description is insufficient for complete understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema is empty (0 parameters), so schema coverage is 100%. The description adds meaning by specifying the resource and scope, which is beyond the empty schema. Baseline 3, but the context provided justifies a higher score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List' and the resource 'active chat sessions' with scope 'for all datasets', distinguishing it from sibling tools like ragflow_list_datasets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool over alternatives like ragflow_list_datasets or ragflow_reset_session. The description only states what it does without usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ragflow_reset_sessionB
Reset/clear the chat session for a specific dataset
| Name | Required | Description | Default |
|---|---|---|---|
| dataset_id | Yes | ID of the dataset to reset session for |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided. Description merely states the action without disclosing side effects (e.g., whether session history is deleted permanently, if it affects other datasets, or if confirmation is required).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, 10 words, no redundancy. Front-loaded with verb and resource. Efficiently communicates the core function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple reset action with one parameter and no output schema, but lacks behavioral details that would help the agent understand consequences. Could mention that the session is cleared without confirmation or return value.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for one parameter. Description mirrors the schema's description ('ID of the dataset to reset session for') without adding new meaning or constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action (reset/clear) and the resource (chat session for a specific dataset). It is distinct from sibling tools which are for listing or retrieval, not mutation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. Does not mention prerequisites, conditions, or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ragflow_retrievalB
Retrieve document chunks directly from RAGFlow datasets using the retrieval API. Returns raw chunks with similarity scores.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number for pagination. Defaults to 1. | |
| query | Yes | Search query or question | |
| top_k | No | Number of chunks for vector cosine computation. Defaults to 1024. | |
| page_size | No | Number of chunks per page. Defaults to 10. | |
| use_rerank | No | Whether to enable reranking for better result quality. Default: false (uses vector similarity only). | |
| dataset_ids | Yes | List of IDs of the datasets/knowledge bases to search | |
| document_name | No | Optional document name to filter results to specific document | |
| deepening_level | No | Level of DSPy query refinement (0-3). 0=none, 1=basic refinement, 2=gap analysis, 3=full optimization. Default: 0 | |
| similarity_threshold | No | Minimum similarity score for chunks (0.0 to 1.0). Defaults to 0.2. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must disclose behavioral traits like whether the tool is read-only, permission requirements, or pagination behavior. It only says 'Returns raw chunks' and does not address these aspects, leaving the agent with incomplete understanding of its side effects or constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is succinct: two sentences that convey the core function and output without extraneous words. It is front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 9 parameters and no output schema, the description should provide more context on how to use parameters like deepening_level or use_rerank, and what the returned chunks contain. It states 'raw chunks with similarity scores' but lacks detail on the structure of the response, which is necessary for an agent to process the output correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description does not add meaning beyond what the parameter descriptions already provide (e.g., page, top_k). It mentions 'similarity scores' but does not clarify how parameters like similarity_threshold relate to the output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Retrieve' and the resource 'document chunks' from RAGFlow datasets, and specifies the output as 'raw chunks with similarity scores'. However, it does not explicitly differentiate from sibling tools like ragflow_retrieval_by_name, which likely performs a similar function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as ragflow_get_chunks or ragflow_retrieval_by_name. It merely states what the tool does, without context on prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ragflow_retrieval_by_nameB
Retrieve document chunks by dataset names using the retrieval API. Returns raw chunks with similarity scores.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number for pagination. Defaults to 1. | |
| query | Yes | Search query or question | |
| top_k | No | Number of chunks for vector cosine computation. Defaults to 1024. | |
| page_size | No | Number of chunks per page. Defaults to 10. | |
| use_rerank | No | Whether to enable reranking for better result quality. Default: false (uses vector similarity only). | |
| dataset_names | Yes | List of names of the datasets/knowledge bases to search (e.g., ['BASF', 'Legal']) | |
| document_name | No | Optional document name to filter results to specific document | |
| deepening_level | No | Level of DSPy query refinement (0-3). 0=none, 1=basic refinement, 2=gap analysis, 3=full optimization. Default: 0 | |
| similarity_threshold | No | Minimum similarity score for chunks (0.0 to 1.0). Defaults to 0.2. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It mentions return type (raw chunks with similarity scores) but lacks information on side effects, permissions, rate limits, or destructive potential. 'Retrieve' implies read-only but is not explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loading the purpose. It is efficient but could be slightly more structured without adding verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 9 parameters and no output schema, the description is sparse. It omits details on pagination, reranking, deepening_level, and similarity_threshold behavior, leaving the agent to rely solely on the schema for context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds minimal meaning beyond the schema, only briefly noting retrieval by dataset names and return format. No parameter interaction hints are provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (retrieve), resource (document chunks), and distinguishing parameter (by dataset names). It differentiates from siblings like ragflow_retrieval which likely uses different criteria.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives. The description implies usage with dataset names but does not mention exclusions or compare to ragflow_retrieval or other search methods.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
ragflow_get_chunks - First observed
ragflow_list_datasets - First observed
ragflow_list_documents - First observed
ragflow_list_documents_by_name - First observed
ragflow_list_sessions - First observed
ragflow_reset_session - First observed
ragflow_retrieval - First observed
ragflow_retrieval_by_name
TDQS
Scored across 8 tools
Most tools have distinct purposes, but ragflow_list_documents and ragflow_retrieval each have an alternative by-name variant, which could cause confusion if descriptions are not heeded. However, descriptions clarify the difference between ID-based and name-based operations, keeping overlap minimal.
All tools follow a consistent verb_noun pattern with snake_case and the 'ragflow_' prefix. Variations like '_by_name' are systematic and predictable, enhancing readability for agents.
With 8 tools, the set is well-scoped for a knowledge base retrieval server. Each tool serves a clear function, and the count is neither too sparse nor overwhelming for the intended purpose.
The tool surface covers listing datasets, listing documents, retrieving chunks, and managing chat sessions. It lacks create/update/delete operations, but given the likely read-heavy focus of the server, these gaps are acceptable and do not impede the primary retrieval workflow.
Maintenance
Related MCP Connectors
Connect your team's living knowledge base — docs, data, issues, CRM — to Claude and ChatGPT.
Cloud or self-hosted knowledge for AI agents: hybrid search, reranking, GraphRAG, scoped MCP tools.
Ingest, manage, and retrieve documents for RAG-powered AI applications
Search your knowledge bases from any AI assistant using hybrid RAG.
Related MCP Servers
- FlicenseAqualityDmaintenanceIntegrates R2R (Retrieval-Augmented Generation) with Claude Desktop, enabling semantic search across knowledge bases and RAG-based question answering with support for vector, graph, web, and document search.2-
- AlicenseAqualityFmaintenanceConnects Claude Desktop to a RAGBrain knowledge base to enable semantic search, document retrieval, and namespace management. It allows users to browse collections, discover documents by topic, and access full text content through natural language.5MIT
- AlicenseAqualityDmaintenanceProvides semantic search capabilities by connecting Claude Desktop to a Cloudflare Workers backend powered by Vectorize. It enables natural language querying of knowledge bases using vector similarity and edge-based embedding generation.2MIT
- AlicenseNot gradedqualityDmaintenanceEnables semantic retrieval and knowledge base management through the RAGFlow API, including dataset, document, chunk, chat, and graph operations.5MIT