EvidenceLens MCP
EvidenceLens MCP is a read-only, deterministic evidence-review server exposing a single review_evidence tool that accepts evidence metadata and returns a schema-validated skeleton response.
Submit a review request with a
reviewIdand anobjective, optionally constrained bylimits(max evidence items, max objective length).Provide up to 20 evidence items, each with an ID, role (
assignment_brief,rubric,teacher_instructions,solution,other), and type (text,pdf,image,screenshot,table).Include inline content as plain text,
contentBase64, or areference; optionally specifyformat/mimeType.Reference filesystem evidence using
filesystem://-style entries with an allowedrootIdand normalizedrelativePathwhen explicit filesystem roots are configured.Operate read-only with no network/provider calls; filesystem access is disabled by default and fails closed rather than falling back to broad directories.
Validate inputs against strict schema constraints: root IDs, reference/relative path patterns, evidence count, and length limits; malformed roots fail with a stable error message.
Return deterministic, schema-validated findings with typed provenance, preserving idempotent and read-only behavior.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@EvidenceLens MCPreview this PDF evidence file against contract constraints"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
EvidenceLens MCP
EvidenceLens MCP is a TypeScript Model Context Protocol server for deterministic, read-only evidence review. The single review_evidence tool accepts four distinct required course roles plus optional evidence, normalizes inline and explicitly configured filesystem sources, and returns schema-validated findings with typed provenance. It remains provider-independent and makes no network or provider calls.
Local Development
npm install
npm run dev
npm test
npm run buildnpm run dev starts the MCP server over stdio from src/server.ts. npm test runs the full contract, normalizer, fixture, safety, and MCP protocol test suite. npm run build type-checks and compiles the server.
Related MCP server: Observability Agent MCP
MCP Contract
See docs/mcp-contract.md for the exact four-role request contract, duplicate-ID and stable errors, deterministic finding fields and citation mapping, root configuration grammar, byte limits, filesystem:// provenance, transient analysis boundary, and read-only/no-provider/no-Docker behavior.
Platform note: the default filesystem reader uses descriptor-relative component walking on Linux. On macOS, where this project has no supported Node openat/openat2 binding, default anchored filesystem reads fail closed with sanitized ACCESS_DENIED and no bytes; there is no pathname fallback. Phase 2 inline evidence remains supported, and embedding tests may inject a reviewed safe filesystem adapter for portable filesystem-read coverage.
Optional filesystem roots
Filesystem access is disabled when no roots are configured; the server never falls back to the current directory, home directory, repository, or another broad default. Configure explicit roots with repeated id=absolute-path entries separated by commas or semicolons:
EVIDENCELENS_ALLOWED_ROOTS='course=/absolute/path/course;examples=/absolute/path/examples' npm run devRoot IDs match [A-Za-z][A-Za-z0-9_-]{0,31}. Empty entries, duplicate IDs, canonical path collisions, unreadable roots, non-directories, relative paths, and malformed entries fail with the stable message Invalid filesystem root configuration; paths are never echoed.
Available Tools
1 toolreview_evidenceReview EvidenceCRead-onlyIdempotent
Accept Phase 1 evidence review metadata and return a deterministic skeleton response.
| Name | Required | Description | Default |
|---|---|---|---|
| limits | No | ||
| evidence | No | ||
| reviewId | Yes | ||
| objective | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already document the safety profile (readOnlyHint=true, idempotentHint=true, destructiveHint=false), and the description adds the behavioral fact that the response is deterministic and skeleton-shaped. However, it does not say what the skeleton response contains or how input is handled, leaving a partial transparency gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence with no filler; the core action and outcome are front-loaded. It is appropriately sized but terse to the point of vagueness, so not a full 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a nested input schema and no output schema, the description is too thin: it does not explain the return shape, what 'Phase 1' means, or what evidence metadata is expected. Annotations cover safety but not operational context needed to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to compensate for parameter meaning, but it mentions no parameter names or semantics. The schema itself provides constraints and enums, but the tool description contributes nothing to clarify reviewId, objective, evidence, or limits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a concrete verb-resource pair ('Accept Phase 1 evidence review metadata') and names the outcome ('return a deterministic skeleton response'), so an agent can tell this is a submission/stub endpoint for the review workflow. However, 'skeleton response' is unexplained and there is no sibling differentiation, though no siblings are listed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance about when to call this tool versus any alternative, no workflow conditions, and no preconditions or exclusions. The only hint is the phrase 'Phase 1,' which is not elaborated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.2- First observed
review_evidence
TDQS
Scored across 1 tool
There is only one tool, so there is no ambiguity between tools, but the single tool's purpose is vague and could overlap with future tools. Since there is no set to disambiguate, it scores minimal.
With only one tool, naming consistency is trivial but the naming is not informative; 'review_evidence' is a generic name that doesn't indicate what it does with evidence. No pattern can be assessed.
A single tool for a server called 'EvidenceLens MCP' that claims to 'review evidence' is too thin. The scope suggests a larger toolset for evidence review, but only one trivial tool is provided.
The tool description says 'Accept Phase 1 evidence review metadata and return a deterministic skeleton response.' This covers only a tiny fraction of what an evidence review workflow would need; there's no way to submit actual evidence, get details, update reviews, or handle later phases.
Maintenance
Related MCP Connectors
Read-only game, setup, place, evidence and travel decision tools with explicit provenance.
Agent memory that refuses to guess: evidence-gated recall, exact-source reads, verifiable deletion.
Read-only finance and operations controls for AI agents with evidence and safe next actions.
Browser-local and CLI static evidence for deployed AI model artifacts.
Related MCP Servers
AlicenseNot gradedqualityFmaintenanceProvides a universal read path for agents to extract facts, metadata, and provenance from local files and guarded remote URLs without using generative LLMs.MIT- AlicenseBqualityAmaintenanceA portable, read-only Model Context Protocol server for turning observability data into bounded evidence that AI agents can inspect safely.7Apache 2.0
- AlicenseNot gradedqualityDmaintenanceEnables step-debugging, deterministic replay, and signed audit evidence for AI agents, compliant with EU AI Act.MIT
- AlicenseNot gradedqualityCmaintenanceEnables secure, read-only access to local project files (including text, DOCX, PDF, and XLSX) through MCP, with strict directory whitelisting and no write, edit, or command-execution tools.1MIT