Evidence MCP
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Evidence MCPRetrieve evidence package for '32.768 kHz load capacitance'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Layout-Aware-RAG
Layout-Aware RAG for Engineering Datasheets
A retrieval-augmented generation pipeline for layout-bearing technical documents. The retrieval unit is a page-native layout region — a set of detected layout elements grouped as one reading unit, such as a table with its caption and footnote — rather than a fixed-length token window. Retrieval fuses a multilingual dense index with a part-number-aware BM25 index by reciprocal rank fusion, and each result carries document identity, page number, and PDF-point bounding boxes, giving region-level provenance back to the source page.
Zikang Zhou · Xiamen University · 2026
Live demos
Two datasheet corpora, searchable in the browser. No account, API key, or local installation is required.
Corpus | Documents | Pages | Chunks | |
TXC — crystals, oscillators, TCXO/VCXO/OCXO | 108 | 226 | 875 | |
TKD — crystals and oscillator modules | 74 | 112 | 547 |
Each demo ships one-click example queries. Representative cases:
Part number, whole or partial —
7M 26.000MHz load capacitance ESR, or only the fragment a user recalls. A standard tokenizer splits these into7 / M / 26 / 000; this one indexes them intact, as subfields, and as character 3-grams.Parameter lookup —
32.768 kHz crystal load capacitance,VCXO 3.3V phase jitter package. Hits resolve to the table region holding the value, rather than to a text window that merely contains the words.Drawing or footprint query —
package dimensions land pattern,pin connection tri-state output, where the answer is a figure rather than prose.Any result, opened. Each hit carries its document, page number, and bounding boxes; the evidence image — the crop of that page region — is bundled for the example queries above. The chunk browser and document tree show how the pages were segmented.
The hosted pages execute BM25 in the browser against an inverted index shipped
with the snapshot: every chunk is searchable and every query is scored at request
time, not replayed from a fixture. Dense retrieval, RRF fusion, and answer
generation are local-pipeline capabilities and are reported as offline in the
results. Evidence images are pre-rendered for the example queries rather than for
all 875 / 547 chunks, which holds each snapshot to a few megabytes instead of
several hundred; running the pipeline locally produces an image for every hit.
The demo interface is in Chinese; the pipeline, the code, and these docs are in English.
Related MCP server: PDFDashboardWithMCP
The problem
Engineering datasheets are layout-bearing documents: the semantics of a value depend on its position within the page structure. A load-capacitance figure is interpretable only in conjunction with the table row containing it, the part-number column heading above it, and the footnote printed beneath the table. Fixed-length token-window chunking discards that structure, and with it the provenance needed to verify a retrieved value against its source.
The retrieval unit is therefore defined at the layout level: a set of detected elements constituting a single reading unit, which retains its page geometry through indexing and into the returned result.
How it works
flowchart LR
PDF["PDF page"] --> R["1. render<br/>200 dpi"]
R --> L["2. layout<br/>YOLO + dedup"]
L --> X["3. extract text<br/>born-digital layer"]
X --> G["4. group<br/>VLM, offline"]
G --> M["5. evidence image"]
M --> I["6. index"]
I --> D["dense"]
I --> B["part-aware BM25"]
D --> F["7. search<br/>RRF fusion"]
B --> F
F --> E["evidence package"]# | Stage | What it does |
1 |
| Rasterize pages at 200 dpi, keeping PDF-point coordinates. |
2 |
| DocLayout-YOLO detects tables, figures, captions, text; duplicate boxes are removed. |
3 |
| Pull the born-digital text layer for each element. |
4 |
| Group elements into chunks; attach a description and a TOC path. Offline. |
5 |
| Stack each chunk's crops into one readable evidence image. |
6 |
| Build the dense and BM25 indexes. |
7 |
| Retrieve from both paths, fuse with RRF. |
Stages 1–6 are offline. Only stage 7 runs at query time, and it needs no API key, no GPU, and no network.
The evidence package
Every result is a record, not an opaque model response:
chunk_id · doc_id · source_pdf · page · bboxes_pdf[]
block_type · section_title · toc_path
native_text · description · crop_images[]Any hit can be replayed against the source PDF independently of the system that produced it. Because the field set is stable under component substitution — layout detector, embedding model, fusion strategy, or answer model — the same contract serves the Python pipeline, the demonstration sites, and the MCP server without being reimplemented for each.
bboxes_pdf is a list, not a union rectangle: each member element keeps its own
box and its own crop.
Four decisions worth explaining
1. PDF points are the stored coordinate; pixels are derived. bbox_pdf is
authoritative, while bbox_px = bbox_pdf × dpi / 72 serves only cropping and
visualization. Re-rendering the corpus at a different DPI therefore invalidates no
stored result.
2. Duplicate boxes are resolved by area, not confidence. YOLO often gives the
box containing more content a lower score. In document 6u, the complete
table_footnote box scores 0.37 against 0.73 for a truncated one — ranking by
confidence truncates the start of the Note line, and does so silently, since the
truncated crop remains well-formed. Ranking by area retains the box preserving more
of the source.
3. Part numbers are tokenized three ways. A standard tokenizer splits
7M-26.000MAAJ-T into 7 / M / 26 / 000 / MAAJ / T, which defeats part-number
search entirely. Here a part-like token is indexed as the intact token (exact
match), as alphanumeric subfields (partial recall), and as character 3-grams
(substring match, so 26.000M alone still retrieves it).
4. The two paths are fused on ranks, not scores. Dense cosine similarity and
BM25 scores are not on a comparable scale, and BM25 in particular varies with query
length and corpus statistics. RRF (1/(k + rank), k = 60) is invariant to both,
so the embedding model can be replaced without re-tuning a fusion weight. Results keep
dense_rank, bm25_rank, and both raw scores, so any hit can be traced to the
path that produced it.
More detail — including the composition of each index, the validation guards on the
grouping stage, and the per-stage artifact formats — is in
docs/METHOD.md.
Getting started
Install (Python 3.11–3.12):
git clone https://github.com/zzkws/Layout-Aware-RAG.git
cd Layout-Aware-RAG
python -m venv .venv
pip install -e ".[embedding,mcp]"Get the data. Download the corpus snapshot from release
v0.1.0 and
restore it under corpora/txc or corpora/tkd. v0.1.0 is the canonical data
reference; later tags are code-only, and retrieval results are not comparable
across corpus revisions.
Search:
RAG_CORPUS=tkd python -m pipeline.search "32.768 kHz load capacitance"$env:RAG_CORPUS = "tkd"
python -m pipeline.search "32.768 kHz load capacitance"Each result prints its fused score, both path ranks, the block type, the page, and the source PDF — enough to open the file and check it.
Run the local demo (standard-library HTTP server, no extra dependencies):
python webapp/server.pyRebuild an index from PDFs. Every stage takes --docs <doc_id> ...; render,
merge_crops, and index_build also take --all:
python -m pipeline.render --docs 6u 7m
python -m pipeline.layout --docs 6u 7m
python -m pipeline.extract_text --docs 6u 7m
python -m pipeline.vlm_blocks --docs 6u 7m
python -m pipeline.merge_crops --docs 6u 7m
python -m pipeline.index_build --docs 6u 7mFor a whole corpus, use the resumable runner — it skips any document whose artifacts already exist and stops gracefully if the Gemini quota runs out:
python -X utf8 -u run_full_corpus.pyOnly stage 4 needs GEMINI_API_KEY.
Use it from an agent (MCP). The server exposes three tools —
build_evidence_package, get_chunk, get_evidence_image — so an agent can ask
for evidence without knowing anything about the indexing:
python evidence_mcp_server.pyConfiguration
All optional; read from the environment only. See .env.example.
Variable | Default | Meaning |
|
| Corpus profile: |
|
| Source PDFs. |
|
| Pipeline output root. |
|
|
|
| — | Stage 4 only. Never read at query time. |
| — | Optional OpenAI-compatible answer generation and query rewriting for the webapp. Unset = retrieval only. |
Tuning constants live in config.py: RENDER_DPI (200), LAYOUT_CONF
(0.25), RRF_K (60), TOP_K (15), and K1/B (1.5 / 0.75) in
pipeline/search.py. BM25 and RRF both score at query time,
so changing any of them needs no index rebuild.
Repository layout
Layout-Aware-RAG/
├── config.py # corpus profiles, coordinate + model constants
├── evidence_service.py # query-time retrieval, fusion, package assembly
├── evidence_mcp_server.py # MCP tools over the evidence contract
├── run_full_corpus.py # resumable whole-corpus runner
├── common/ # part-aware tokenizer, embedding backends, drawing
├── pipeline/ # the seven stages, one file each
├── docs/METHOD.md # design decisions in detail
├── docs/EVALUATION.md # what has and has not been measured
├── webapp/ # local demo: stdlib server + 4 pages
├── sites/{txc,tkd}-demo/ # the two static public deployments
├── tools/ # public export, disclosure scan, release packaging
└── tests/ # core behavior and public-data contract testsStatus and limitations
This is a demo and design study, not a benchmarked system. It has not been
quantitatively evaluated — the design choices above are engineering arguments and
observed cases, not measurements. No accuracy claim, no SOTA claim, and no claim of
superiority over pixel-based document RAG. docs/EVALUATION.md
records what a real evaluation would need.
Known scope limits:
No OCR. Scanned PDFs are out; the pipeline assumes a born-digital text layer.
No cell-level tables. A query about one cell returns the whole table region.
No cross-page table joining. A table split by a page break stays two chunks.
Single-column reading order. Ordering by vertical position is wrong for genuinely multi-column pages.
Grouping is frozen. Descriptions and TOC paths came from one VLM run; regenerating them changes the index.
One document family. Both corpora are crystal/oscillator datasheets from adjacent vendors.
Data rights are separate from code rights. Apache-2.0 covers the code, not the datasheets.
Built on
DocLayout-YOLO for layout detection (Zhao et al., 2024, arXiv:2410.12628) · PyMuPDF for rendering and text extraction · sentence-transformers and fastembed for embeddings · Okapi BM25 (Robertson & Zaragoza, 2009) and Reciprocal Rank Fusion (Cormack et al., 2009) for retrieval and fusion.
Citation
@software{zhou2026layoutawarerag,
author = {Zhou, Zikang},
title = {Layout-Aware-RAG: Layout-Aware Retrieval-Augmented Generation
for Engineering Datasheets},
year = {2026},
url = {https://github.com/zzkws/Layout-Aware-RAG},
license = {Apache-2.0}
}Machine-readable metadata is in CITATION.cff.
License
Apache-2.0 for the code. The datasheets and third-party models carry their own terms — see DATA_NOTICE.md and THIRD_PARTY_NOTICES.md. Security reports: SECURITY.md.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Evidence-readiness MCP server: validate, audit, and score briefs, memos, and evidence packs.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Remote MCP for C2PA intake verifier MCP, structured receipts, audit logs, and reviewer-ready evidenc
Tamper-evident proof creation and verification for AI agents via MCP, A2A, and REST.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceTransforms PDF collections into a searchable knowledge base using TF-IDF indexing and proximity matching. It enables users to search documents, retrieve specific page content, and manage document libraries through natural language via MCP clients.5
- AlicenseAqualityDmaintenanceEnables MCP clients to list indexed PDF document collections and perform semantic search queries on them using locally extracted text and embeddings.2AGPL 3.0
- AlicenseNot gradedqualityAmaintenanceAn MCP-native evidence retrieval platform that ingests source material, builds lexical and vector indexes, performs hybrid retrieval, and returns structured evidence packages for AI assistants.1Apache 2.0
- AlicenseBqualityBmaintenanceProvides a read-only MCP interface to query and retrieve verifiable evidence from a local memory bank, supporting search, dossier, chronology, source, and evidence tools.6BSD Zero Clause
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/zzkws/Layout-Aware-RAG'
If you have feedback or need assistance with the MCP directory API, please join our Discord server