GraphRAG TypeScript MCP Tools
GraphRAG TypeScript MCP Tools
A complete implementation of a GraphRAG MCP server built with TypeScript, Neo4j, and the MCP TypeScript SDK. This project demonstrates how to build production-quality MCP servers that expose graph-backed tools, resources, and advanced features like LLM sampling and completions.
Built as part of the Neo4j GraphAcademy — Building GraphRAG TypeScript MCP tools course.
MCP란 무엇인가?
Model Context Protocol(MCP)는 Anthropic이 만든 개방형 표준으로, AI 에이전트(Claude, Cursor, VS Code Copilot)가 외부 도구와 데이터 소스에 표준화된 방식으로 연결할 수 있게 해줍니다.
Related MCP server: CodeRAG
프로젝트 구조
genai-mcp-build-custom-tools-typescript/ ├── server/ │ └── index.ts ← 메인 MCP 서버: 도구 4개 + 리소스 1개 + 샘플링 + 자동 완성 ├── strawberry/ │ └── index.ts ← 첫 번째 MCP 서버: 간단한 countLetters 도구 ├── solutions/ ← 과정 참조 솔루션 ├── .vscode/ │ └── mcp.json ← VS Code MCP 구성 └── README.md
구축된 내용
1단계 — 첫 번째 MCP 서버 (strawberry/index.ts)
가장 단순한 MCP 서버입니다. 도구 하나, 데이터베이스 없음, stdio 전송.
server.registerTool("countLetters", {
description: "Count occurrences of a letter in the text",
inputSchema: {
text: z.string().describe("The text to search in"),
search: z.string().describe("The letter to count"),
},
}, async ({ text, search }) => ({
content: [{
type: "text",
text: String(text.toLowerCase().split(search.toLowerCase()).length - 1),
}],
}));테스트 결과: countLetters("strawberry", "r") → 3
MCP Inspector를 사용하여 테스트했습니다 — MCP 서버를 탐색하고 테스트하기 위한 브라우저 기반 도구입니다.
2단계 — Neo4j 연결 (모듈 범위)
Python의 lifespan 컨텍스트 관리자와 달리 TypeScript는 모듈 범위 변수를 사용합니다 — 드라이버가 파일 상단에서 한 번 생성되고 모든 도구가 직접 공유합니다.
// Created ONCE when file loads — shared by all tools
const driver: Driver = neo4j.driver(
process.env["NEO4J_URI"] ?? "neo4j://localhost:7687",
neo4j.auth.basic(
process.env["NEO4J_USERNAME"] ?? "neo4j",
process.env["NEO4J_PASSWORD"] ?? "password"
)
);
const database = process.env["NEO4J_DATABASE"] ?? "neo4j";SIGINT를 통한 정상 종료:
process.on("SIGINT", async () => {
await driver.close();
await server.close();
process.exit(0);
});3단계 — 도구 1: graphStatistics
Neo4j의 모든 노드와 관계를 계산합니다.
결과: {"nodes": 28863, "relationships": 332522}
4단계 — 도구 2: getMoviesByGenre
IMDB 평점 순으로 영화를 장르별로 검색합니다. 로깅에는 console.error()를 사용합니다 — stdio 서버에서는 절대 console.log()를 사용하지 마세요(JSON-RPC 채널을 손상시킵니다).
server.registerTool("getMoviesByGenre", {
description: "Get movies by genre from the Neo4j database",
inputSchema: {
genre: z.string().describe("The genre to search for (e.g., Action, Comedy, Drama)"),
limit: z.number().default(10).describe("Maximum number of movies to return"),
},
}, async ({ genre, limit }) => {
const { records } = await driver.executeQuery(query,
{ genre, limit: neo4j.int(limit) }, // neo4j.int() for 64-bit integer compatibility
{ database }
);
...
});5단계 — 도구 3: browse_movies_by_genre (페이지네이션)
Neo4j의 SKIP 및 LIMIT를 사용하는 커서 기반 페이지네이션:
const skip = parseInt(cursor, 10) || 0;
// Cypher: SKIP $skip LIMIT $limit
const nextCursor = movies.length === pageSize ? String(skip + pageSize) : null;반환값:
{
"genre": "Action",
"movies": [...],
"nextCursor": "2",
"page": 1,
"pageSize": 2,
"hasMore": true,
"count": 2
}6단계 — 리소스: movie://{tmdbId}
ResourceTemplate을 사용하여 TMDB ID로 전체 영화 세부 정보를 노출합니다:
server.registerResource(
"movie",
new ResourceTemplate("movie://{tmdbId}", { list: undefined }),
{ description: "Get detailed information about a specific movie", mimeType: "application/json" },
async (uri, { tmdbId }) => {
// uri.href = "movie://603"
// returns: contents array with JSON movie data
}
);예시: movie://603 (The Matrix), movie://13 (Forrest Gump)
7단계 — 고급: 샘플링 (explainMovieData)
실행 중에 LLM을 호출하여 원시 Neo4j 데이터를 자연어로 변환하는 도구:
const result = await server.server.createMessage({
messages: [{
role: "user",
content: {
type: "text",
text: `Describe '${movieData.title}' (${movieData.released})...`,
},
}],
maxTokens: 200,
});샘플링 없음: {'title': 'Toy Story', 'released': '1995', 'actors': [...]}
샘플링 사용(VS Code Copilot): "Toy Story — Buzz Lightyear가 새로운 인기 캐릭터가 되면서 자리가 밀려난 느낌을 받는 질투심 많은 카우보이 인형 Woody에 관한 영리하고 재미있는 애니메이션 모험..."
참고: 저수준 서버에 capability를 설정해야 합니다:
server.server["_capabilities"] = { ...server.server["_capabilities"], completions: {} };
8단계 — 고급: 자동 완성
장르 매개변수에 대한 실시간 자동 완성 제안 — 사용자가 입력하는 동안 Neo4j를 쿼리합니다:
import { CompleteRequestSchema } from "@modelcontextprotocol/sdk/types.js";
server.server.setRequestHandler(CompleteRequestSchema, async (request) => {
if (request.params.argument.name === "genre") {
const { records } = await driver.executeQuery(
`MATCH (g:Genre)
WHERE g.name STARTS WITH $prefix
RETURN g.name AS name
ORDER BY name ASC LIMIT 10`,
{ prefix: request.params.argument.value },
{ database }
);
return { completion: { values: records.map(r => r.get("name")) } };
}
return { completion: { values: [] } };
});Python 버전과의 주요 차이점
개념 | Python (FastMCP) | TypeScript (McpServer) |
도구 등록 |
|
|
공유 상태 | lifespan 컨텍스트 관리자 | 모듈 범위 변수 |
드라이버 접근 |
|
|
로깅 |
|
|
샘플링 |
|
|
자동 완성 |
|
|
파일 구조 | 기능별 별도 파일 | 모든 것이 하나의 |
숫자 매개변수 | Python int 타입 힌트 |
|
프롬프트 매개변수 |
| 항상 |
설정
사전 요구 사항
Node.js 20+
npm
Neo4j Sandbox — sandbox.neo4j.com의 Recommendations 데이터셋
설치
git clone https://github.com/Akakinad/genai-mcp-build-custom-tools-typescript
cd genai-mcp-build-custom-tools-typescript
npm install자격 증명 구성
cat > server/.env << EOF
NEO4J_URI=bolt://your-sandbox-ip:7687
NEO4J_USERNAME=neo4j
NEO4J_PASSWORD=your-password
NEO4J_DATABASE=neo4j
EOF설정 확인
npx tsx client/test_environment.ts
# Expected: All checks passed!실행
MCP Inspector로 테스트 (브라우저 UI)
cd server
npx @modelcontextprotocol/inspector npx tsx index.ts터미널에 표시된 URL을 엽니다 → Connect → Tools 탭 → List Tools → 도구 선택 → Run Tool.
AI 에디터용 서버 실행
cd server
npx tsx index.tsVS Code 구성 (.vscode/mcp.json)
{
"servers": {
"movies-ts": {
"type": "stdio",
"command": "npx",
"args": ["tsx", "/absolute/path/to/server/index.ts"]
}
}
}VS Code Copilot에서 테스트
movies-ts MCP 도구를 사용하여 영화 "Toy Story"를 설명해 보세요 movies-ts MCP 도구를 사용하여 액션 영화를 검색해 보세요 movies-ts MCP 도구를 사용하여 그래프 통계를 가져오세요
과정
학습 경로: Generative AI & GraphRAG
과정: Building GraphRAG TypeScript MCP tools
Building GraphRAG TypeScript MCP Tools
GraphAcademy 과정 Building GraphRAG TypeScript MCP Tools의 동반 리포지토리입니다.
학생들은 Neo4j 그래프 데이터베이스에 연결되는 MCP(Model Context Protocol) 서버를 구축하여 AI 어시스턴트가 사용할 수 있는 도구와 리소스를 노출합니다.
시작하기
.env.example을.env로 복사하고 Neo4j 연결 세부 정보로 값을 업데이트하세요.종속성을 설치합니다:
npm install서버를 시작합니다:
npm startMCP Inspector로 서버를 검사합니다:
npm run inspect솔루션
solutions/ 디렉토리에는 각 레슨 체크포인트에 대한 완성된 코드가 포함되어 있습니다.
Available Tools
4 toolsbrowse_movies_by_genreC
Browse movies in a genre with pagination support
| Name | Required | Description | Default |
|---|---|---|---|
| genre | Yes | Genre name (e.g. Action, Comedy, Drama) | |
| cursor | No | Pagination cursor - position in the result set | 0 |
| pageSize | No | Number of movies to return per page |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must disclose behavioral traits. It mentions pagination support but does not state whether this is a read-only operation, whether it requires authentication, or what happens with invalid genres. The mutation safety profile is unclear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that conveys the core purpose efficiently. It contains no unnecessary words, but could benefit from a second sentence on when to use this vs getMoviesByGenre.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description does not explain return values (e.g., movie details format, total result count). For a paginated browsing tool with three parameters and sibling overlap, more context is needed to guide correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters. The description adds minimal value beyond listing genres and pagination, but it does not clarify cursor semantics (e.g., whether it is a page number or token). Baseline 3 is appropriate given full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Browse' and resource 'movies in a genre' with pagination support. While it differentiates from siblings like getMoviesByGenre (similar purpose) and explainMovieData (different purpose), it could be more explicit about how browsing differs from getting movies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool vs getMoviesByGenre, which appears to have overlapping functionality. It does not specify exclusions or alternatives, leaving the agent to infer usage from the description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explainMovieDataA
Get a natural language explanation of movie data using LLM sampling
| Name | Required | Description | Default |
|---|---|---|---|
| movieTitle | Yes | The title of the movie |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description bears full responsibility for behavioral disclosure. It mentions 'using LLM sampling', which implies non-deterministic generative behavior, but omits details on authorization, rate limits, or potential side effects. The minimal context is acceptable for a read-like tool but lacks depth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence of 10 words, front-loading the core functionality. Every word contributes meaning, and there is no redundant or irrelevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given low complexity (1 required parameter, no output schema), the description adequately states the tool's purpose but fails to specify what aspects of movie data are explained (e.g., plot, cast, ratings) or the nature of the 'natural language explanation'. The output format is left entirely to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% coverage with a single parameter 'movieTitle' described as 'The title of the movie'. The description adds no additional parameter semantics beyond the schema, so it meets the baseline for well-documented parameters without extra value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Get', the resource 'movie data', and the method 'natural language explanation using LLM sampling'. It distinguishes itself from siblings like getMoviesByGenre and graphStatistics, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus its siblings. It does not mention when-not-to-use, prerequisites, or alternatives, leaving the agent to infer based solely on the tool name and sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
getMoviesByGenreC
Get movies by genre from the Neo4j database
| Name | Required | Description | Default |
|---|---|---|---|
| genre | Yes | The genre to search for (e.g., Action, Comedy, Drama) | |
| limit | No | Maximum number of movies to return |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description alone must disclose behavioral traits. It only states 'Get movies by genre' without confirming read-only behavior, authentication needs, or pagination behavior (though the schema reveals a default limit of 10). This minimal disclosure leaves significant behavioral ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence of 8 words, with no superfluous information. It is front-loaded with the core action. However, it could be slightly expanded with additional context (e.g., limit behavior) without becoming verbose, so it does not achieve a perfect score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description should explain what the tool returns (e.g., list of movie objects, format). It does not. Additionally, it fails to differentiate from the sibling tool 'browse_movies_by_genre', making the overall context incomplete for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage for both parameters, so the baseline is 3. The description does not add any extra meaning beyond what the schema provides (e.g., it does not clarify whether genre matching is exact or fuzzy). It scores neither higher nor lower than the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get movies by genre') and the data source ('Neo4j database'). It is specific and provides a direct understanding of the tool's function. However, it does not differentiate itself from the sibling tool 'browse_movies_by_genre', which appears to have a very similar purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to use this tool versus alternatives like 'browse_movies_by_genre'. It lacks any context about prerequisites, limitations, or typical use cases. Without such guidance, an agent may select the wrong tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
graphStatisticsB
Count the number of nodes and relationships in the graph
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full burden. It states a read operation (count), but doesn't disclose whether the count is real-time, cached, or if it requires permissions. For a zero-parameter tool, more context on performance or scope would help.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence, perfectly sized for the tool's simplicity. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters, no output schema, and simple purpose, the description is largely adequate. However, it doesn't mention the format or granularity of the counts (e.g., separate numbers for nodes vs. relationships), leaving minor ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has zero parameters with 100% coverage, so the description doesn't need to add param details. The baseline is 4 due to schema fully covering the (empty) parameter list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Count') and the resources ('nodes and relationships in the graph'). It differentiates from siblings (which filter by genre or explain data) by being a global count tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for getting graph size, but it doesn't explicitly say when to use this vs. siblings (e.g., 'Use this for an overview; use getMoviesByGenre for filtering'). No when-not or alternatives mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.0.0- First observed
browse_movies_by_genre - First observed
explainMovieData - First observed
getMoviesByGenre - First observed
graphStatistics
TDQS
Scored across 4 tools
The two tools for movies by genre (getMoviesByGenre and browse_movies_by_genre) have overlapping purposes; the only distinction is pagination support, which may cause an agent to misselect. Other tools are distinct.
Mixes camelCase (getMoviesByGenre, explainMovieData, graphStatistics) and snake_case (browse_movies_by_genre) with different verb conventions ('get' vs 'browse'), showing inconsistency.
With only 4 tools, the surface is minimal but perhaps appropriate for a read-only movie graph query server. It borders on being too few but is not extreme.
Missing basic operations like fetching a specific movie by ID, listing all genres, or querying actors/relationships. The tool set covers only a small subset of expected graph queries, leaving significant gaps.
Maintenance
Related MCP Connectors
MCP server unifying ERPs, CRMs, APIs and knowledge base for Claude, ChatGPT and Gemini.
MCP server for AI dialogue using various LLM models via AceDataCloud
MCP server connecting AI agents to 100+ apps (Gmail, Slack, Notion, GitHub) via one-click OAuth.
- ZapierOAuthcom.zapier
Hosted MCP server connecting AI assistants to 9,000+ apps and 40,000+ actions via Zapier.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceAn MCP server that enables graph database interactions with Neo4j, allowing users to access and manipulate graph data through natural language commands.-
- AlicenseNot gradedqualityDmaintenanceAn MCP server that transforms codebases into knowledge graphs using Neo4J, enabling AI assistants to understand code structure, relationships, and metrics for more context-aware assistance.28MIT
- AlicenseAqualityCmaintenanceAn MCP server that enables LLMs to perform semantic and fulltext searches within Neo4j while executing complex, search-augmented Cypher queries for GraphRAG applications. It provides tools for database schema discovery and supports multi-provider embeddings to facilitate advanced graph traversals.53MIT
- AlicenseNot gradedqualityBmaintenanceMCP server for Neo4j graph database operations, enabling Cypher queries, node/relationship management, and schema discovery.1BSD 3-Clause