site-context-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@site-context-mcpSearch the cached site corpus for 'LangGraph' and return snippets with source URLs."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
site-context-mcp
Problem
Agents invent project and site claims without citations. When you ask what a public site says about a product page, a topic note, or an essay, you often get fluent prose and no way to verify it against the live corpus.
Related MCP server: DuckDuckGo MCP Server
Solution
A stdio MCP server that:
Fetches public pages starting from
llms.txt, then follows same-host links to core pagesCaches them in-process (~1-hour TTL)
Exposes tools that return structured answers with a top-level
sources[]of public URLs on every response
No secrets, no LinkedIn scraping, no private admin / WOD-APP links — only HTTPS to the wired public site host.
This package is for site/project context. It does not duplicate resume identity or resume-summary tools (see sibling resume-mcp for that).
The default corpus is YongBo Yu’s public site. The same fetch/cache/source-URL shape can be pointed at any public llms.txt-indexed site corpus.
Requirements
Python 3.11+
Network access to the wired public host (default:
yongbo-yu.vercel.app) when running the live server
Install
# from this repo
pip install -e .
# or with uv
uv pip install -e .Dev / smoke tests (mocked HTTP; no live network required):
pip install -e ".[dev]"
pytest -qRun (stdio)
# module entrypoint
python -m site_context_mcp
# console script (after install)
site-context-mcp
# with uv
uv run python -m site_context_mcpThe process speaks MCP over stdin/stdout. Do not pipe unrelated stdout into the same process.
Cursor mcp.json
Add a server entry (Cursor: Settings → MCP, or edit ~/.cursor/mcp.json / project .cursor/mcp.json):
{
"mcpServers": {
"site-context-mcp": {
"command": "python3",
"args": ["-m", "site_context_mcp"],
"cwd": "/absolute/path/to/site-context-mcp"
}
}
}If the package is installed into a venv, point command at that interpreter (or use the site-context-mcp console script).
With uv:
{
"mcpServers": {
"site-context-mcp": {
"command": "uv",
"args": ["run", "python", "-m", "site_context_mcp"],
"cwd": "/absolute/path/to/site-context-mcp"
}
}
}After saving, reload MCP servers in Cursor. The tools below should appear as site-context-mcp tools.
Tools
Tool | Purpose |
| Catalog of cached public pages from |
| One page by path ( |
| Keyword search over the cached corpus (snippets + source URLs) |
| Project evidence page (default: |
Every tool response is JSON text that includes a top-level sources array of public URLs. Partial fetch failures are reported in cache_notes when present.
Example prompts (in Cursor / Claude)
“List the public pages this MCP has cached and cite the source URLs.”
“Get
/projects/kilodockand quote the engineering metrics with sources.”“Search the site corpus for ‘LangGraph’ / ‘Codex’ and return snippets with URLs.”
“Use
get_projectfor KiloDock evidence — do not invent stack claims.”
Data sources (public only)
On startup (and on first tool use if needed), the server fetches:
https://yongbo-yu.vercel.app/llms.txtSame-host pages linked from
llms.txt(home, about, project pages, topic pages, essays)Optionally
resume.jsonwhen linked — used only for light project cross-links, not as a resume clone
Binaries and off-host / private targets are skipped. No API keys.
Author
YongBo Yu (also Yong Yu) — Toronto, Canada · GitHub YongBoYu1
KiloDock evidence: https://yongbo-yu.vercel.app/projects/kilodock
Site index: https://yongbo-yu.vercel.app/llms.txt
License
MIT — see LICENSE.
Available Tools
4 toolsget_pageA
Fetch one cached public page by path (e.g. /about-yongbo-yu) or absolute HTTPS URL.
| Name | Required | Description | Default |
|---|---|---|---|
| path_or_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description itself must carry behavioral information. It discloses that the page is cached, implying possible staleness, and public, implying no auth requirement. 'Fetch' also indicates a read operation; cache-miss behavior is not covered but is secondary for a single-page fetch.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence front-loads the action and resource, and the parenthetical gives a working example without redundancy. There is no filler or restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one required parameter and an output schema present, the description covers input semantics, cache/public behavior, and single-page scope sufficiently. It lacks only explicit routing guidance versus siblings and cache-miss specifics, both minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only the parameter name path_or_url with no description, so the description fully compensates by specifying accepted forms: a relative path like '/about-yongbo-yu' or an absolute HTTPS URL. This is exactly the semantic detail the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb 'Fetch' plus resource 'cached public page' and the qualifier 'one' make the tool's scope explicit. The path and absolute HTTPS URL examples disambiguate it from list_site_pages and search_site, which operate over collections.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'one cached public page by path or absolute HTTPS URL' tells an agent this is the right tool when a specific page identifier is already known. It does not explicitly name sibling alternatives like list_site_pages or search_site, so the decision is left to inference rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_projectB
Return public project evidence for a site project page (default: kilodock).
| Name | Required | Description | Default |
|---|---|---|---|
| name_or_slug | No | kilodock |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description must carry the safety/behavior burden. 'Public' and 'Return' imply an auth-free, read-only operation, which is useful, but nothing is disclosed about errors, edge cases, or any other behavioral traits beyond that minimal implication.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler or redundant clauses. It is appropriately sized for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even with one optional parameter and an output schema, the description is incomplete for confident selection: it leaves 'project evidence' undefined and gives no usage guidance relative to sibling tools. It is enough for a trivial default invocation but not for understanding the tool's real role.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only repeats the default value 'kilodock' without explaining how name_or_slug is matched or what happens when it is omitted. The parameter is simple, but the description does not add meaningful semantics beyond the schema field title.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and names the resource ('public project evidence for a site project page') with a default value. However, 'project evidence' is somewhat opaque and the description does not explicitly differentiate this tool from siblings like get_page or search_site.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given for when to use this tool versus list_site_pages, get_page, or search_site. The phrase 'for a site project page' only weakly implies a context, with no exclusions or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_site_pagesB
List cached public site pages discovered from llms.txt (paths, titles, URLs).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states it lists cached public pages, implying a read-only operation, but does not disclose any potential rate limits, pagination, or staleness of the cache. It doesn't contradict annotations (none), but the description is thin on behavioral nuance.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, efficient, and front-loaded with the key action. It's concise and to the point, though it could add a hint about usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so the description need not explain the return format. The tool is simple (no params) and the description covers core aspects. However, missing usage guidance and behavioral details (cache freshness) make it minimally adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has zero parameters, so there is no need for parameter explanations. The description adds value by specifying the source (llms.txt) and the contents (paths, titles, URLs), which helps the agent understand the output without needing to inspect the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb+resource: lists cached public site pages. Mentions source (llms.txt) and what's included (paths, titles, URLs), distinguishing it from get_page (single page) and search_site (search). Slight lack: does not explicitly state it returns all pages or any filtering.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies it is for discovering cached public pages from llms.txt, which suggests use when you need an inventory of site pages. It doesn't explicitly compare with get_page or search_site, but the distinction is fairly obvious from the sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_siteB
Keyword search over the cached site corpus; returns snippets with source URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears the full burden of behavioral disclosure. It mentions 'cached' implying the data is a snapshot, which is useful, but it does not explain whether results are ordered, what happens on no matches, or any rate limits or permissions. For a search tool, behavior around result limits or pagination is not disclosed, which is a gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence (~12 words) that front-loads the core purpose and output. It does not waste words, though it could benefit from expanding on parameters or limitations without sacrificing brevity. The structure is appropriate for a straightforward search tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has a simple input schema (query and limit) and an output schema (which presumably describes the snippet format), the description covers the primary action and output type. However, it lacks guidance on result ordering, default behavior when 'limit' is omitted, and any special characters or operators supported in 'query.' These omissions make it minimally complete but not comprehensive for an agent unfamiliar with the corpus.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, meaning neither 'query' nor 'limit' is described in the schema beyond their names and types. The description adds no information about the parameters: it does not explain that 'query' is the search string or that 'limit' controls result count. With no parameter documentation in the schema, the description fails to compensate, leaving the agent to guess parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the verb 'search' over a specific resource ('cached site corpus') and indicates the output ('snippets with source URLs'). This distinguishes it from sibling tools like list_site_pages, which lists pages, and get_page, which retrieves a specific page; the description makes the search-and-snippet behavior explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for keyword-based discovery but does not explicitly state when to use this tool versus siblings. It does not mention alternatives or provide exclusion criteria, such as 'use list_site_pages to browse all pages.' The context of search is clear, but guidance on alternatives is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
get_page - First observed
get_project - First observed
list_site_pages - First observed
search_site
TDQS
Scored across 4 tools
Each tool has a clear, non-overlapping purpose: listing pages, fetching a single page, searching the corpus, and retrieving project evidence. No two tools could be confused for the same operation.
All tools follow a consistent verb_noun pattern in snake_case: list_site_pages, get_page, search_site, get_project. The naming is uniform and predictable.
Four tools is a well-scoped set for a site-context server—enough to cover enumeration, retrieval, search, and specialized project data without redundancy or bloat.
The surface covers the core needs for browsing a cached site corpus: list, get, search, and project-specific lookup. There are no obvious missing operations for a read-only context server.
Maintenance
Related MCP Connectors
LLM-ready web search + instant answers + URL-to-clean-text fetch for agents and RAG.
Collaborative, cache-first web search for agents — cited answers from a shared live-web pool.
Shared copies of public web pages for AI agents. Search stored pages or fetch a URL.
Web search and clean-text page fetch for AI agents, with SSRF protection.
Related MCP Servers
- FlicenseNot gradedqualityNot gradedmaintenanceEnables querying and retrieving content from webpages by parsing sitemap.xml files and fetching HTML content from specified URLs. Includes rate limiting for abuse protection.-
- AlicenseBqualityDmaintenanceEnables web search through DuckDuckGo and webpage content fetching with intelligent text extraction. Features built-in rate limiting and LLM-optimized result formatting for seamless integration with language models.2MIT
- AlicenseNot gradedqualityNot gradedmaintenanceProvides controlled access to llms.txt documentation files through MCP tools, allowing AI assistants to fetch and read documentation from user-approved domains with full audit visibility of tool calls and context retrieval.MIT
- AlicenseNot gradedqualityDmaintenanceConverts URLs into clean, LLM-ready markdown, respecting robots.txt and never bypassing anti-bot measures or paywalls.MIT