asset-aware-mcp
# asset-aware-mcp
> Citation-ready document infrastructure for AI agents: editable native documents,
> reusable evidence, and portable Foam/LightRAG wikis.
[](https://opensource.org/licenses/Apache-2.0)
π [ηΉι«δΈζ](README.zh-TW.md) Β· [Docs Site](https://u9401066.github.io/asset-aware-mcp/#/overview) Β· [GitHub Wiki](https://github.com/u9401066/asset-aware-mcp/wiki)
## v1.4.1 Consolidated native-document release
This release brings together the document and evidence work below. We keep the
**1.4.x** line and consolidate verified changes instead of releasing each feature.
Check [GitHub Releases](https://github.com/u9401066/asset-aware-mcp/releases) for
publication status and the [changelog](CHANGELOG.md#141---2026-09-22) for details.
| Format | Supported workflows |
|---|---|
| PDF | Pages/regions, page assembly, annotations, AcroForm fields, actual previews, historical evidence |
| DOCX | Creation, scoped body/table edits, table grids/layout, headers/footers, footnotes/endnotes |
| XLSX/XLSM | Cells, worksheet/grid/layout operations, native Tables, A2T application and independent XLSX creation |
| ODS | Creation, exact logical-cell edits/clearing, dependency-checked worksheet rename and optional Calc previews |
| CSV/TSV | Literal fields and row/column CRUD with explicit dialects and byte-preservation checks |
| PPTX | Slides, text, notes, pictures, editable tables and optional whole-slide previews |
| Raster images | Frames/regions, PNG/TIFF derivatives, composition and checked candidate revisions |
| Evidence/Wiki | Immutable selections/derivations, complete receipts, CSL/custom citations and exact source attachments |
Discover the installed runtime's complete contract and enabled flags before acting.
Since v1.4.0, `native-contract-v2` can page schemas: inspect `schema_delivery`, select
`for_op`, or follow the hash-pinned `schema_request`. Follow `contract_request` for
paged policies too; abbreviated previews are not complete operation instructions.
## π― Why Asset-Aware MCP?
Model document analysis may be enough for summaries and questions. This project's
purpose is ongoing work on native documents with reusable, checkable results:
which source revision was used, which component changed, what stayed intact, and
which evidence still supports a claim after later edits or a change of model.
An asset carries identity, revision, native locators, representations, relationships,
available operations and validation state. Wiki notes link these assets; citation
formatting preserves the underlying evidence. Complete operation receipts and
historical sources remain available after component edits or deletion.
**MCP checks source/version integrity, native constraints and operation results.
The Agent performs complete semantic and visual review and coordinates corrections.**
Static previews do not establish interactive viewer parity or universal fidelity.
Arbitrary PDF body editing and ODS worksheet/grid lifecycle remain open work.
See [native usage and limits](docs/wiki/Native-File-Assets.md),
[capability gaps and upstream references](docs/agent-asset-gap-analysis.md),
[actual Agent evaluations](docs/wiki/Release-And-Testing.md) and [contracts](docs/spec.md).
## β¨ Features
Version 1.3.0 adds a native DOCX bridge with revision-bound `read_docx` and
`update_docx`, reusing the existing DFM checks before committing a managed version.
It preserves unrelated package parts and supports explicit tracked text changes;
source writeback and agent layout review remain separate. DOCX block references now
support bounded reads and native verification; wiki snapshots include full block
records, the original DOCX and exact package-part attachments. Existing 1.2.0 wiki
outputs retain their identities and bytes. See the [DOCX bridge guide](docs/wiki/Native-File-Assets.md#130-native-docx-bridge).
- π **Asset-Aware ETL** - PDF β Markdown with a pluggable multi-engine parser (`ETL_ENGINE`):
- **PyMuPDF** (default) - Fast extraction (~50MB), no models required
- **PyMuPDF4LLM** (`[pdf-plus]`) - Drop-in layout-aware upgrade, no GPU
- **Docling** (`[docling]`) - MIT-licensed layout+table+formula+chart engine; bridges through an isolated `.venv-docling` interpreter when the main environment can't install it directly (see [docs/docling-setup.md](docs/docling-setup.md))
- **MinerU** - Adapter retained, but the packaged extra is on security hold while MinerU pins a vulnerable `transformers<5` chain
- **Marker** - Adapter retained for evaluation, but production selection fails closed while upstream `marker-pdf` conflicts with the patched Pillow floor. The legacy `use_marker` parameter now means βprefer the configured structured extractorβ; it does not bypass this hold.
- π§© **Unified Segmentation Export** - Normalized `segmentation.json` merges manifest, blocks, reading order, and persisted markdown line spans for downstream tools and extensions.
- π©Ί **Safe PDF Preflight Router** - `document(op="preflight")` classifies each page as native, sparse, image, scanned, or hybrid; returns 1-based top-left locators, source SHA-256, OCR reasons, and a bounded extraction-engine recommendation from a process-isolated inspector.
- π¦ **Reusable Agent Asset Bundles** - `document(op="export_assets")` writes deterministic `manifest.json`, `assets.jsonl`, copied media, and a portable Foam `index.md`/`notes/**` subtree while preserving stable IDs, hashes, locators, and citation refs.
- π‘οΈ **PDF Safety/Structure/Coverage/Accessibility Audits** - OpenDataloader-inspired artifact-only reports flag suspicious hidden/off-page/prompt-injection text, native structure signals, segmentation coverage gaps, and accessibility/readability readiness via the existing `document` facade. `document(op="prepare_ai")` and `document(op="auto")` expose agent-ready status and next actions without adding public tools.
- π§ **Structural Pointer Retrieval** - Proxy-Pointer-inspired `document(op="pointer_index")`, `document(op="structural_retrieve")`, and `document(op="compare")` preserve section breadcrumbs, line/char/byte locators, source hashes, asset IDs, and evidence-span provenance without adding MCP tools.
- πΌοΈ **Layout Overlay Debugging** - Render page overlays from `original.pdf` to inspect bbox, segment type, and reading order visually.
- π€ **On-Demand OCR Preprocessing** - Optional `ocrmypdf` preprocessing path for scanned PDFs before ETL.
- π§ **Section Navigation** - Dynamic hierarchy section tree through the `section` facade: browse, search, detail, content reading, and block extraction for any depth of headings.
- π **Async Job Pipeline** - Supports asynchronous ingest, configured structured parse, OCR, and conversion jobs with progress tracking.
- π **Mixed-Format Batch Ingestion** - `document(op="auto", file_paths=[...])` auto-detects a batch mixing PDF with DOCX/DOC/ODT/ODS, ingests each file through its correct existing engine in one background job, isolates per-file failures so one bad file cannot abort the rest, and reports per-file progress β no new public tool required.
- πΊοΈ **Document Manifest** - Provides a structured "map" of the document for precise data access by Agents.
- π§ **LightRAG Integration** - Knowledge Graph + Vector Index, supporting cross-document comparison and reasoning.
- π§Ύ **Verified Citation Bundles** - `citation_bundle`, Foam evidence packs, citation health checks, table/figure evidence notes, and claim promotion export citation-ready spans with locator, quote/hash, context, CRAAP scaffold, and verification status.
- π **Docx Editing (DFM)** - Edit .docx files in Markdown via **Docx-Flavored Markdown** format. Supports legacy `.doc`, `.odt`, and `.ods` ingest via LibreOffice auto-conversion. The balanced surface keeps 6 DOCX/DFM public entrypoints for ingest, read, save, validation, conversion, table edit planning, and Docx β A2T bridges.
- π‘οΈ **DFM Integrity Checker** - Automatic validation and auto-repair at every pipeline stage (post-ingest, pre-save, post-save). Catches orphan markers, column mismatches, and format inconsistencies.
- π **A2T (Anything to Table)** - 7 operation-based tools for building professional tables from **any source** (PDF assets, Knowledge Graph, URLs, user input). Features: stable row IDs, row search/filter/paging, citation coverage, artifact-only large-table render, skipped-large-table UX, **Citations** (AssetRef), **Audit Trail**, **Schema Evolution**, **Templates**, **Drafting**, and **Token-efficient resumption**.
- π₯οΈ **VS Code Management Extension** - Graphical interface for monitoring server status, ingested documents, document artifacts, citation spans, and **A2T tables/drafts** with one-click Excel export.
- π **MCP SDK 2 Server** - Uses the official Python SDK `MCPServer` API, runtime-injected context, and v2 clients. MCP SDK v1 is intentionally unsupported.
- π¬ **Research-ready, domain-neutral assets** - Works with scholarly, technical, policy, and operational documents; bounded image bytes let compatible multimodal clients analyze figures instead of relying on server-local paths.
## ποΈ Architecture
```
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β AI Agent (Copilot) β
βββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββ
β MCP Protocol (Tools & Resources)
βββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββ
β MCP Server (Modular Presentation) β
β βββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β tools/: 30 public tools (balanced surface) β β
β β 17 facade tools + 13 high-frequency shortcuts β β
β β compact=17 β legacy/direct compatibility=63 β
β βββββββββββββββββββββββββββββββββββββββββββββββββββ β
β βββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β resources/: 13 resources in 2 modules β β
β βββββββββββββββββββββββββββββββββββββββββββββββββββ β
βββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββ
β
βββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββ
β ETL Pipeline (DDD) β
β ββββββββββββ ββββββββββββ ββββββββββββ β
β β PyMuPDF β β Asset β β LightRAG β β
β β Adapter ββ β Parser ββ β Index β β
β ββββββββββββ ββββββββββββ ββββββββββββ β
βββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββ
β
βββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββ
β Local Storage β
β ./data/ β
β βββ {doc_id}/ # PDF document artifacts β
β βββ docx_{id}/ # Docx IR + DFM + Assets β
β βββ tables/ # A2T Tables (JSON/MD/XLSX) β
β β βββ drafts/ # Table Drafts (Persistence) β
β βββ lightrag_db/ # Knowledge Graph β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
```
## π Project Structure (DDD)
```
asset-aware-mcp/
βββ src/
β βββ domain/ # π΅ Domain: Entities, Value Objects, Interfaces
β βββ application/ # π’ Application: Doc Service, Table Service (A2T), Asset Service
β βββ infrastructure/ # π Infrastructure: PyMuPDF, LightRAG, Excel Renderer
β βββ presentation/ # π΄ Presentation: MCP SDK 2 MCPServer
βββ data/ # Document and Asset Storage
βββ docs/
β βββ spec.md # Technical Specification
βββ tests/ # Unit and Integration Tests
βββ vscode-extension/ # VS Code Management Extension
βββ pyproject.toml # uv Project Config
```
## π Architecture and workflows
The maintained, versioned references are the [documentation site](https://u9401066.github.io/asset-aware-mcp/),
[architecture guide](docs/wiki/Architecture.md), [PDF workflow](docs/wiki/PDF-Document-Workflow.md),
[MCP tool catalog](docs/wiki/MCP-Tools.md), and [release checklist](docs/wiki/Release-And-Testing.md).
They are generated and checked with the implementation so tool counts, engine holds, and release behavior do not drift inside obsolete screenshots.
## π Quick Start
```bash
# Install dependencies (using uv) β default install stays on the fast PyMuPDF backend
uv sync
# Optional high-fidelity PDF->asset engines:
# uv sync --extra pdf-plus # PyMuPDF4LLM: drop-in layout-aware upgrade
# uv sync --extra docling # Docling: MIT layout+table+formula+chart engine
# MinerU and Marker packaged extras are temporarily empty security holds.
# Then set ETL_ENGINE=pymupdf4llm|docling.
# Run MCP Server
uv run python -m src.presentation.server
# Or use the VS Code extension for graphical management
```
Runtime note:
The VS Code extension prefers a managed Python 3.11 runtime when launching the MCP server via version-pinned `uv tool run`, with Python 3.10 fallback for older machines. This avoids native package builds on end-user machines, especially macOS systems without Xcode Command Line Tools, while keeping the project itself compatible with newer Python versions.
Installation scope note:
- The VS Code extension installs once per user. In a trusted workspace, the native VS Code MCP provider may use workspace-scoped `DATA_DIR`, cache, settings, and `.env`; local source is accepted only in Extension Development/Test mode (or a future explicit opt-in).
- Global Codex and Cline entries always launch `asset-aware-mcp==<extension-version>` from extension global storage. They do not inherit workspace-local source, workspace-scoped settings, or repository `.env` values. Restricted Mode skips external config writes and assistant-asset sync entirely.
Engine selection note:
`ETL_ENGINE` picks the extraction backend (default `pymupdf`). The active packaged structured engines (`pymupdf4llm`, `docling`) lazy-load and gracefully fall back to PyMuPDF when their extra is not installed. Marker remains on hold because `marker-pdf` requires `Pillow<11`; MinerU is also on hold because MinerU 3.4.4 pins `transformers<5` while current security fixes require `transformers>=5.5`. Both adapters remain in-tree, but this package will not install a known-vulnerable dependency chain. Use `document(op="preflight", pdf_path="...")` to choose between fast native extraction, OCR, and Docling before ingest.
Agent asset / Foam handoff:
```text
document(op="preflight", pdf_path="/papers/source.pdf")
document(op="auto", file_paths=["/papers/source.pdf"])
document(op="export_assets", doc_id="doc_...", output_dir="agent-assets")
```
The exported directory is deterministic and portable: `manifest.json` is the
bundle contract, `assets.jsonl` is the agent-readable inventory, and
`index.md` plus `notes/**` can be mounted or copied into a Foam workspace.
## π MCP Tools
The default runtime surface is **balanced**: 30 public tools that keep the full document workflow available without overwhelming agents. It is made of 17 operation-based facade tools plus 13 high-frequency shortcuts. Set `ASSET_AWARE_MCP_TOOL_SURFACE=compact` for the 17 facade-only surface, or `ASSET_AWARE_MCP_TOOL_SURFACE=legacy` / `ASSET_AWARE_MCP_ENABLE_LEGACY_TOOLS=true` for the full 63-tool compatibility inventory.
| Area | Balanced public tools |
|------|------------------------|
| Documents, assets, evidence, conversion | `document`, `document_asset`, `evidence`, `convert_document`, `ingest_documents`, `list_documents`, `parse_pdf_structure`, `fetch_document_asset`, `find_evidence_spans`, `verify_citation_ref`, `citation_bundle` |
| DOCX / DFM | `docx`, `docx_table`, `ingest_docx`, `get_docx_content`, `save_docx`, `docx_table_edit_plan` |
| Sections, jobs, KG, ETL profiles | `section`, `job`, `get_job_status`, `list_jobs`, `knowledge`, `etl_profile` |
| A2T tables | `plan_table`, `table_manage`, `table_data`, `table_cite`, `table_history`, `table_draft`, `discover_sources` |
See [MCP Tools](docs/wiki/MCP-Tools.md) and [Tool Consolidation](docs/wiki/MCP-Tool-Consolidation.md) for operation details, shortcut rationale, and legacy direct-tool mapping.
Agent handoff note:
Use `document(op="auto", file_paths=[...])` for new PDFs and `document(op="auto", doc_id="...")` or `document(op="prepare_ai", doc_id="...")` for existing documents. `document(op="prepare_ai", output_format="json")` returns the v2 readiness contract with `status`, `blockers`, `warnings`, `capabilities`, `artifacts`, `missing_audits`, `invalid_audits`, `audit_artifacts`, and `next_actions`. `document(op="audit", doc_id="...")` reuses current audit artifacts only when they are present and valid; pass `refresh=true` to rebuild safety, native-structure, coverage, and accessibility reports. Use `document(op="pointer_index")`, `document(op="structural_retrieve", query="...")`, and `document(op="compare", doc_b_id="...", criteria="...")` when an agent needs section-level structural retrieval or comparison without new public tools. Readiness and job-status artifact discovery are read-only, so status checks do not create document directories.
PDF audit caveat:
The audit reports are inspired by OpenDataloader-style artifact workflows, but they are not a sanitizer, a PDF/UA certification, or an OpenDataloader compatibility layer. They preserve source artifacts and report conservative diagnostics for review.
## π§ Tech Stack
| Category | Technology |
|----------|------------|
| Language | Python 3.10+ |
| Package Manager | **uv** (all pip/setup-python removed) |
| ETL | **PyMuPDF** (default) + secure optional **PyMuPDF4LLM** / **Docling** engines; MinerU and Marker adapters are on dependency security hold |
| RAG | LightRAG (lightrag-hku) |
| MCP | Official Python MCP SDK 2 (`MCPServer`); SDK v1 unsupported |
| Storage | Local filesystem (JSON/Markdown/PNG) |
## π Documentation
Installation guidance:
- Default install: `uv sync` (slim ~227 MB; no LightRAG/KG dependencies).
- LightRAG / Knowledge Graph backend (optional, since v0.6.34): `uv tool install --upgrade --python 3.11 'asset-aware-mcp[lightrag]'` for uvx/published users, or `uv sync --extra lightrag` for local source checkouts. Required before setting `ENABLE_LIGHTRAG=true`.
- VS Code extension: run the command `Asset-Aware MCP: Install LightRAG Backend` from the Command Palette; it auto-detects source vs published mode and emits the matching install command.
- OpenRouter optional preset (since v0.6.35): set `LLM_BACKEND=openrouter`, `OPENROUTER_API_KEY=...`, and optionally `OPENROUTER_MODEL=liquid/lfm-2.5-1.2b-instruct:free` for fast low-cost summaries and draft RAG answers. LightRAG retrieval still uses the configured embedding backend.
- High-fidelity PDF engines: `uv sync --extra pdf-plus` (PyMuPDF4LLM) or `uv sync --extra docling` (Docling), then set `ETL_ENGINE` accordingly. Docling ships a cross-platform isolated installer; see [docs/docling-setup.md](docs/docling-setup.md).
- MinerU and Marker backends: their adapters remain available for upstream testing, but the packaged extras are empty security holds until their dependency caps permit patched `transformers` and Pillow releases.
- VS Code extension: `assetAwareMcp.enableMarkerBackend` is retained as a setting, but the launcher will not install `marker-pdf` while the security hold is active.
- [Technical Spec](docs/spec.md) - Detailed technical specification
- [Architecture](ARCHITECTURE.md) - System architecture
- [Constitution](CONSTITUTION.md) - Project principles
- [Competitive Analysis](docs/competitor-analysis.md) - MCP + DOCX ecosystem landscape
## π License
[Apache License 2.0](LICENSE)
TDQS
Scored across 30 tools
The tool set contains multiple consolidated facades and direct tools that heavily overlap: document, document_asset, fetch_document_asset, section, docx, docx_table, and evidence all cover document/asset territory, while job duplicates get_job_status/list_jobs. Even with detailed descriptions, an agent must read long contracts to pick the right entrypoint. Several tools have unclear boundaries despite some well-written individual descriptions.
Names are consistently lowercase and snake_case, but they follow no single pattern: action-first names like ingest_documents and get_docx_content coexist with noun-only entrypoints like document, job, section, and evidence, plus the mixed table_manage/table_data versus plan_table style. Most names are readable and descriptive, but the generic single-word facades are vague and the overall convention is uneven.
30 tools exceeds the 25-tool threshold and is inflated by redundant consolidated/direct pairs, e.g., document vs document_asset vs docx, and job vs get_job_status/list_jobs. The underlying domain is broad, but many tools could be merged or removed without losing capability, making the surface feel heavy for agents to navigate.
The server covers the document lifecycle well: PDF/DOCX ingestion, document listing and asset retrieval, table CRUD/drafts/citations, evidence spans/bundles, DOCX write-back, and job tracking. Minor gaps remain, such as the lack of a clearly exposed direct PDF create/update/delete tool outside the monolithic document facade and no OCR path, but core workflows have no dead ends.