Skip to main content
Glama

Why GraphRAG MCP

To improve information retrieval efficiency within enterprises, there is a need for an Agent capable of extracting user-relevant information from scattered documents. By building enterprise document retrieval as an MCP (Model Context Protocol), other agents within the organization can simply connect to this MCP whenever document retrieval is required. This approach centralizes document search capabilities, making it easier to integrate and scale intelligent agents across the enterprise.

How to Run

  1. Fill in the required environment variables in both .env files located in the project root and in the graphrag/ directory.

  2. Open your terminal and navigate to the project root directory.

  3. Run the following command to start the GraphRAG MCP:

    uv run rag_client.py rag_server.py

After these steps, the GraphRAG MCP will be up and running.

Code Structure

.
├── main.py
├── rag_client.py # MCP client for graphrag
├── rag_server.py # MCP server for graphrag
├── pyproject.toml
├── .env 
├── .gitignore
├── .python-version
├── README.md
└── graphrag/
    ├── .env
    ├── settings.yaml
    ├── cache/
    │   ├── community_reporting/
    │   ├── extract_graph/
    │   ├── summarize_descriptions/
    │   └── text_embedding/
    ├── input/ 
    │   └── test.txt # knowledge base
    ├── logs/
    ├── output/
    │   ├── context.json
    │   ├── stats.json
    │   └── lancedb/
    └── prompts/
        ├── basic_search_system_prompt.txt
        ├── community_report_graph.txt
        ├── community_report_text.txt
        ├── drift_reduce_prompt.txt
        ├── drift_search_system_prompt.txt
        ├── extract_claims.txt
        ├── extract_graph.txt
        └── ...

Folder and File Descriptions

  • rag_client.py: Client for interacting with the RAG server.

  • rag_server.py: Server providing RAG-based tools.

  • pyproject.toml: Python project configuration.

  • .env, .gitignore, .python-version: Environment and versioning files.

graphrag/

  • .env, settings.yaml: Environment and configuration for GraphRAG.

  • cache/: Stores intermediate results for various pipeline stages.

    • community_reporting/: Caches community report data.

    • extract_graph/: Caches graph extraction results.

    • summarize_descriptions/: Caches description summaries.

    • text_embedding/: Caches text embeddings for retrieval.

  • input/: Contains input data files.

  • logs/: Stores log files generated during runs.

  • output/: Stores output data such as context and statistics.

    • context.json: Output context data.

    • stats.json: Output statistics.

    • lancedb/: Database files for vector storage.

  • prompts/: Contains prompt templates for different tasks.

    • basic_search_system_prompt.txt: Prompt for basic search.

    • community_report_graph.txt: Prompt for graph-based community reports.

    • community_report_text.txt: Prompt for text-based community reports.

    • drift_reduce_prompt.txt, drift_search_system_prompt.txt: Prompts for drift analysis.

    • extract_claims.txt, extract_graph.txt: Prompts for claim and graph extraction.

Core Technology: GraphRAG

GraphRAG is a Retrieval-Augmented Generation (RAG) framework that integrates graph-based reasoning with large language models. It enhances traditional RAG by representing entities, relationships, and claims as a knowledge graph, enabling more structured and context-aware retrieval.

Related MCP server: Graforest MCP

How GraphRAG Works

  1. Data Ingestion: Raw text and tabular data are processed to extract entities, relationships, and claims, which are stored in a graph structure.

  2. Prompting: Custom prompts guide the language model to generate responses grounded in the graph data.

  3. Retrieval: When a query is received, relevant nodes and edges from the graph are retrieved using semantic search and graph traversal.

  4. Generation: The language model synthesizes answers using both the retrieved graph context and prompt templates, ensuring responses are accurate and well-grounded.

  5. Caching and Output: Intermediate and final results are cached for efficiency and stored for further analysis.

This approach allows for more explainable, reliable, and context-rich answers compared to standard RAG pipelines.

Available Tools

1 tool
rag_MLC
用于查询特斯拉与安克创新的对比的相关信息
:param query: 用户提出的具体问题
:return: 最终获得的答案
ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, and the description does not disclose return format, potential side effects, or any behavioral expectations beyond 'returns the final answer.' The tool's behavior is largely opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and to the point, with no unnecessary words or repetition. It effectively conveys the core function in a single sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is minimal and does not provide enough context for an agent to understand what kind of information will be returned, how to structure queries, or any edge cases. It lacks examples and clarifications that would make it fully self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole parameter 'query' has no description in the schema. While the name implies it is a search query, no details are given about expected format, length, or examples, leaving the parameter semantics ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to query comparison information between Tesla and Anker Innovations. The verb '查询' (query) is specific, and the object is well-defined. Since there are no sibling tools, differentiation is not needed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool, how to formulate queries, or any limitations. The description mentions only the purpose, leaving usage context entirely implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observedrag_ML

TDQS

C2.9/5.0

Scored across 1 tool

Disambiguation5/5

Only one tool exists, so there is no possibility of overlap or confusion. The tool's purpose is clear from its description.

Naming Consistency4/5

With a single tool, naming consistency is trivially maintained. However, the name 'rag_ML' is somewhat cryptic and does not clearly convey its function, so a small deduction is applied.

Tool Count2/5

The server exposes only one tool, which feels insufficient for a GraphRAG MCP. Typical RAG servers offer at least multiple query modes or additional operations, making this count on the low side.

Completeness1/5

The tool surface is severely incomplete for a RAG system. It only provides a single query function with no capabilities for data ingestion, document management, or configuration, leaving major gaps in expected functionality.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    F
    maintenance
    Provides AI agents with persistent memory and knowledge management through a comprehensive knowledge graph platform. Enables storing, searching, and managing entities, relationships, and observations with advanced features like trending analysis and smart ranking.
    3
    -
  • F
    license
    A
    quality
    C
    maintenance
    Enables AI agents to build, populate, and search knowledge graphs by providing tools for entity extraction, relationship mapping, and graph traversal. It manages the underlying database infrastructure so users can create searchable knowledge bases from text through natural language commands.
    13
    1
    -
  • A
    license
    A
    quality
    D
    maintenance
    Provides persistent long-term memory for AI agents through semantic search and automated knowledge graph extraction. It enables agents to store, recall, and reason over facts, preferences, and relationships across multiple conversations and sessions.
    14
    14 npm
    MIT