Skip to main content
Glama
Akakinad

GraphRAG TypeScript MCP Tools

by Akakinad

GraphRAG TypeScript MCP Tools

一个使用 TypeScript、Neo4j 和 MCP TypeScript SDK 构建的 GraphRAG MCP 服务器的完整实现。本项目演示了如何构建生产级 MCP 服务器,公开基于图的工具、资源以及 LLM 采样和补全等高级功能。

本项目是 Neo4j GraphAcademy — Building GraphRAG TypeScript MCP tools 课程的一部分。


什么是 MCP?

模型上下文协议(MCP)是 Anthropic 制定的开放标准,允许 AI 代理(Claude、Cursor、VS Code Copilot)以标准化方式连接到外部工具和数据源。


Related MCP server: CodeRAG

项目结构

genai-mcp-build-custom-tools-typescript/
├── server/
│ └── index.ts ← Main MCP server: 4 tools + 1 resource + sampling + completions
├── strawberry/
│ └── index.ts ← First MCP server: simple countLetters tool
├── solutions/ ← Course reference solutions
├── .vscode/
│ └── mcp.json ← VS Code MCP configuration
└── README.md

构建内容

步骤 1 — 第一个 MCP 服务器 (strawberry/index.ts)

最简单的 MCP 服务器。只有一个工具,没有数据库,stdio 传输。

server.registerTool("countLetters", {
  description: "Count occurrences of a letter in the text",
  inputSchema: {
    text: z.string().describe("The text to search in"),
    search: z.string().describe("The letter to count"),
  },
}, async ({ text, search }) => ({
  content: [{
    type: "text",
    text: String(text.toLowerCase().split(search.toLowerCase()).length - 1),
  }],
}));

测试结果: countLetters("strawberry", "r") → 3

使用 MCP Inspector 进行测试 — 这是一个基于浏览器的工具,用于探索和测试 MCP 服务器。


步骤 2 — Neo4j 连接 (模块作用域)

与 Python 的 lifespan 上下文管理器不同,TypeScript 使用 模块作用域变量 — 驱动程序在文件顶部创建一次,并由所有工具直接共享。

// Created ONCE when file loads — shared by all tools
const driver: Driver = neo4j.driver(
  process.env["NEO4J_URI"] ?? "neo4j://localhost:7687",
  neo4j.auth.basic(
    process.env["NEO4J_USERNAME"] ?? "neo4j",
    process.env["NEO4J_PASSWORD"] ?? "password"
  )
);
const database = process.env["NEO4J_DATABASE"] ?? "neo4j";

通过 SIGINT 实现优雅关闭:

process.on("SIGINT", async () => {
  await driver.close();
  await server.close();
  process.exit(0);
});

步骤 3 — 工具 1:graphStatistics

统计 Neo4j 中的所有节点和关系。

结果: {"nodes": 28863, "relationships": 332522}


步骤 4 — 工具 2:getMoviesByGenre

按类型搜索电影,并按 IMDB 评分排序。使用 console.error() 进行日志记录 — 在 stdio 服务器中绝不要使用 console.log()(它会破坏 JSON-RPC 通道)。

server.registerTool("getMoviesByGenre", {
  description: "Get movies by genre from the Neo4j database",
  inputSchema: {
    genre: z.string().describe("The genre to search for (e.g., Action, Comedy, Drama)"),
    limit: z.number().default(10).describe("Maximum number of movies to return"),
  },
}, async ({ genre, limit }) => {
  const { records } = await driver.executeQuery(query,
    { genre, limit: neo4j.int(limit) },  // neo4j.int() for 64-bit integer compatibility
    { database }
  );
  ...
});

步骤 5 — 工具 3:browse_movies_by_genre (分页)

使用 Neo4j 的 SKIP 和 LIMIT 进行基于游标的分页:

const skip = parseInt(cursor, 10) || 0;
// Cypher: SKIP $skip LIMIT $limit
const nextCursor = movies.length === pageSize ? String(skip + pageSize) : null;

返回:

{
  "genre": "Action",
  "movies": [...],
  "nextCursor": "2",
  "page": 1,
  "pageSize": 2,
  "hasMore": true,
  "count": 2
}

步骤 6 — 资源:movie://{tmdbId}

使用 ResourceTemplate 通过 TMDB ID 公开完整的电影详情:

server.registerResource(
  "movie",
  new ResourceTemplate("movie://{tmdbId}", { list: undefined }),
  { description: "Get detailed information about a specific movie", mimeType: "application/json" },
  async (uri, { tmdbId }) => {
    // uri.href = "movie://603"
    // returns: contents array with JSON movie data
  }
);

示例: movie://603(《黑客帝国》),movie://13(《阿甘正传》)


步骤 7 — 高级:采样 (explainMovieData)

在执行过程中调用 LLM,将原始 Neo4j 数据转换为自然语言的工具:

const result = await server.server.createMessage({
  messages: [{
    role: "user",
    content: {
      type: "text",
      text: `Describe '${movieData.title}' (${movieData.released})...`,
    },
  }],
  maxTokens: 200,
});

不使用采样: {'title': 'Toy Story', 'released': '1995', 'actors': [...]}

使用采样(VS Code Copilot): "《玩具总动员》— 一部机智、有趣的动画冒险片,讲述伍迪——一个嫉妒的牛仔玩偶——在巴斯光年成为新宠后感到被冷落……"

注意:需要在底层服务器上设置 capability:

server.server["_capabilities"] = { ...server.server["_capabilities"], completions: {} };

步骤 8 — 高级:补全

为类型参数提供实时自动补全建议 — 在用户输入时查询 Neo4j:

import { CompleteRequestSchema } from "@modelcontextprotocol/sdk/types.js";

server.server.setRequestHandler(CompleteRequestSchema, async (request) => {
  if (request.params.argument.name === "genre") {
    const { records } = await driver.executeQuery(
      `MATCH (g:Genre)
       WHERE g.name STARTS WITH $prefix
       RETURN g.name AS name
       ORDER BY name ASC LIMIT 10`,
      { prefix: request.params.argument.value },
      { database }
    );
    return { completion: { values: records.map(r => r.get("name")) } };
  }
  return { completion: { values: [] } };
});

与 Python 版本的主要区别

概念

Python (FastMCP)

TypeScript (McpServer)

工具注册

@mcp.tool() 装饰器

server.registerTool() 方法

共享状态

Lifespan 上下文管理器

模块作用域变量

驱动程序访问

ctx.request_context.lifespan_context.driver

driver(直接)

日志记录

await ctx.info()

console.error()

采样

ctx.session.create_message()

server.server.createMessage()

补全

@server.completion()

server.server.setRequestHandler(CompleteRequestSchema)

文件结构

每个功能独立文件

所有内容都在一个 index.ts 中

数字参数

Python int 类型提示

需要使用 neo4j.int() 包装

提示参数

int, str, float

始终使用 z.string(),手动解析


设置

先决条件

安装

git clone https://github.com/Akakinad/genai-mcp-build-custom-tools-typescript
cd genai-mcp-build-custom-tools-typescript
npm install

配置凭据

cat > server/.env << EOF
NEO4J_URI=bolt://your-sandbox-ip:7687
NEO4J_USERNAME=neo4j
NEO4J_PASSWORD=your-password
NEO4J_DATABASE=neo4j
EOF

验证设置

npx tsx client/test_environment.ts
# Expected: All checks passed!

运行

使用 MCP Inspector 进行测试(浏览器 UI)

cd server
npx @modelcontextprotocol/inspector npx tsx index.ts

打开终端中显示的 URL → 连接 → 工具选项卡 → 列出工具 → 选择工具 → 运行工具。

为 AI 编辑器运行服务器

cd server
npx tsx index.ts

VS Code 配置(.vscode/mcp.json)

{
  "servers": {
    "movies-ts": {
      "type": "stdio",
      "command": "npx",
      "args": ["tsx", "/absolute/path/to/server/index.ts"]
    }
  }
}

在 VS Code Copilot 中测试

使用 movies-ts MCP 工具解释电影《玩具总动员》 使用 movies-ts MCP 工具搜索动作片 使用 movies-ts MCP 工具获取图统计信息


课程

学习路径: Generative AI & GraphRAG

课程: Building GraphRAG TypeScript MCP tools


Building GraphRAG TypeScript MCP Tools

GraphAcademy 课程 Building GraphRAG TypeScript MCP Tools 的配套仓库。

学员将构建一个 MCP(模型上下文协议)服务器,连接到 Neo4j 图数据库,公开工具和资源以供 AI 助手使用。

开始

  1. 将 .env.example 复制为 .env,并使用你的 Neo4j 连接信息更新其中的值。

  2. 安装依赖:

npm install
  1. 启动服务器:

npm start
  1. 使用 MCP Inspector 检查服务器:

npm run inspect

解决方案

solutions/ 目录包含每个课程检查点的完整代码。

Available Tools

4 tools
browse_movies_by_genreC

Browse movies in a genre with pagination support

ParametersJSON Schema
NameRequiredDescriptionDefault
genreYesGenre name (e.g. Action, Comedy, Drama)
cursorNoPagination cursor - position in the result set0
pageSizeNoNumber of movies to return per page

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must disclose behavioral traits. It mentions pagination support but does not state whether this is a read-only operation, whether it requires authentication, or what happens with invalid genres. The mutation safety profile is unclear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that conveys the core purpose efficiently. It contains no unnecessary words, but could benefit from a second sentence on when to use this vs getMoviesByGenre.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description does not explain return values (e.g., movie details format, total result count). For a paginated browsing tool with three parameters and sibling overlap, more context is needed to guide correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters. The description adds minimal value beyond listing genres and pagination, but it does not clarify cursor semantics (e.g., whether it is a page number or token). Baseline 3 is appropriate given full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Browse' and resource 'movies in a genre' with pagination support. While it differentiates from siblings like getMoviesByGenre (similar purpose) and explainMovieData (different purpose), it could be more explicit about how browsing differs from getting movies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool vs getMoviesByGenre, which appears to have overlapping functionality. It does not specify exclusions or alternatives, leaving the agent to infer usage from the description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

explainMovieDataA

Get a natural language explanation of movie data using LLM sampling

ParametersJSON Schema
NameRequiredDescriptionDefault
movieTitleYesThe title of the movie

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description bears full responsibility for behavioral disclosure. It mentions 'using LLM sampling', which implies non-deterministic generative behavior, but omits details on authorization, rate limits, or potential side effects. The minimal context is acceptable for a read-like tool but lacks depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence of 10 words, front-loading the core functionality. Every word contributes meaning, and there is no redundant or irrelevant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given low complexity (1 required parameter, no output schema), the description adequately states the tool's purpose but fails to specify what aspects of movie data are explained (e.g., plot, cast, ratings) or the nature of the 'natural language explanation'. The output format is left entirely to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% coverage with a single parameter 'movieTitle' described as 'The title of the movie'. The description adds no additional parameter semantics beyond the schema, so it meets the baseline for well-documented parameters without extra value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get', the resource 'movie data', and the method 'natural language explanation using LLM sampling'. It distinguishes itself from siblings like getMoviesByGenre and graphStatistics, which serve different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus its siblings. It does not mention when-not-to-use, prerequisites, or alternatives, leaving the agent to infer based solely on the tool name and sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

getMoviesByGenreC

Get movies by genre from the Neo4j database

ParametersJSON Schema
NameRequiredDescriptionDefault
genreYesThe genre to search for (e.g., Action, Comedy, Drama)
limitNoMaximum number of movies to return

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description alone must disclose behavioral traits. It only states 'Get movies by genre' without confirming read-only behavior, authentication needs, or pagination behavior (though the schema reveals a default limit of 10). This minimal disclosure leaves significant behavioral ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence of 8 words, with no superfluous information. It is front-loaded with the core action. However, it could be slightly expanded with additional context (e.g., limit behavior) without becoming verbose, so it does not achieve a perfect score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of an output schema, the description should explain what the tool returns (e.g., list of movie objects, format). It does not. Additionally, it fails to differentiate from the sibling tool 'browse_movies_by_genre', making the overall context incomplete for an agent to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for both parameters, so the baseline is 3. The description does not add any extra meaning beyond what the schema provides (e.g., it does not clarify whether genre matching is exact or fuzzy). It scores neither higher nor lower than the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get movies by genre') and the data source ('Neo4j database'). It is specific and provides a direct understanding of the tool's function. However, it does not differentiate itself from the sibling tool 'browse_movies_by_genre', which appears to have a very similar purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives like 'browse_movies_by_genre'. It lacks any context about prerequisites, limitations, or typical use cases. Without such guidance, an agent may select the wrong tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graphStatisticsB

Count the number of nodes and relationships in the graph

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden. It states a read operation (count), but doesn't disclose whether the count is real-time, cached, or if it requires permissions. For a zero-parameter tool, more context on performance or scope would help.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence, perfectly sized for the tool's simplicity. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters, no output schema, and simple purpose, the description is largely adequate. However, it doesn't mention the format or granularity of the counts (e.g., separate numbers for nodes vs. relationships), leaving minor ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has zero parameters with 100% coverage, so the description doesn't need to add param details. The baseline is 4 due to schema fully covering the (empty) parameter list.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Count') and the resources ('nodes and relationships in the graph'). It differentiates from siblings (which filter by genre or explain data) by being a global count tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for getting graph size, but it doesn't explicitly say when to use this vs. siblings (e.g., 'Use this for an overview; use getMoviesByGenre for filtering'). No when-not or alternatives mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv1.0.0
    • First observedbrowse_movies_by_genre
    • First observedexplainMovieData
    • First observedgetMoviesByGenre
    • First observedgraphStatistics

TDQS

C2.9/5.0

Scored across 4 tools

Disambiguation2/5

The two tools for movies by genre (getMoviesByGenre and browse_movies_by_genre) have overlapping purposes; the only distinction is pagination support, which may cause an agent to misselect. Other tools are distinct.

Naming Consistency2/5

Mixes camelCase (getMoviesByGenre, explainMovieData, graphStatistics) and snake_case (browse_movies_by_genre) with different verb conventions ('get' vs 'browse'), showing inconsistency.

Tool Count3/5

With only 4 tools, the surface is minimal but perhaps appropriate for a read-only movie graph query server. It borders on being too few but is not extreme.

Completeness2/5

Missing basic operations like fetching a specific movie by ID, listing all genres, or querying actors/relationships. The tool set covers only a small subset of expected graph queries, leaving significant gaps.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that transforms codebases into knowledge graphs using Neo4J, enabling AI assistants to understand code structure, relationships, and metrics for more context-aware assistance.
    28
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    An MCP server that enables LLMs to perform semantic and fulltext searches within Neo4j while executing complex, search-augmented Cypher queries for GraphRAG applications. It provides tools for database schema discovery and supports multi-provider embeddings to facilitate advanced graph traversals.
    5
    3
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    MCP server for Neo4j graph database operations, enabling Cypher queries, node/relationship management, and schema discovery.
    1
    BSD 3-Clause