corpus.333.eco
OfficialAnchors document provenance in the Bitcoin blockchain via OpenTimestamps: each document's response reports whether an .ots proof is stored beside the source, along with the command to verify it, so authenticity can be checked independently of the server.
Returns the DOI of each document that has one (e.g. 10.5281/zenodo.22217516) and a https://doi.org/... resolution command, letting callers confirm the canonical deposited text matches the served bytes.
Documents are deposited under Zenodo DOIs; the provenance envelope reports deposited_matches_current plus the DOI link, so the served text can be checked against the archived Zenodo record.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@corpus.333.ecosearch for 'zero-point game'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
corpus.333.eco
An MCP server for an open-licensed corpus — mechanism papers, essays, institutional positions, white papers and the Letters to Miss Aquarius — served with verifiable provenance.
152 documents. 144 CC0-1.0 · 8 CC-BY-4.0. 84 carry a DOI; 152 carry an OpenTimestamps proof.
Written by npm run build from the index itself — see scripts/build-index.mjs.
npx @333eco/corpusEvery response carries the document's sha256, its DOI where one exists, and
whether an OpenTimestamps proof is anchored beside the source. A retrieval
server normally asks to be believed. This one hands over the means to check it.
Why that matters
An agent that cites a passage is in a worse position than a human who does: it cannot walk to the shelf. It takes whatever the transport delivered, and nothing about a plausible-looking response distinguishes the canonical text from a paraphrase, a truncation, or a substitution somewhere upstream.
So the check moves into the response:
"provenance": {
"sha256": "91a41c759b1e9fdac3c0da667fe32a906aef98e4edec4832c09ac27b2c7663c0",
"doi": "10.5281/zenodo.22217516",
"opentimestamps": true,
"deposited_matches_current": true,
"hostile_review": null,
"verify": {
"sha256": "printf '%s' \"$(cat <file>)\" | shasum -a 256 # compare to provenance.sha256",
"doi": "https://doi.org/10.5281/zenodo.22217516",
"opentimestamps": "ots verify defensive-publications/zero-point-game.md.ots"
}
}The verify block is an instruction, not a promise — it tells you exactly what
to run, so you do not have to trust this sentence either.
hostile_review is the date of the latest ruled review by an outside reader chosen as
the one most likely to find the document wrong, whose section structure still matches
the text served — or null, stated rather than omitted. It is derived from the corpus's
public reviews.json (roles only, never names), and a structural revision resets it.
⭐ The demonstration is the contribution. Anyone can propose provenance-carrying retrieval. Serving it over a corpus where the anchors are already years deep in the Bitcoin blockchain is a different claim, and it is not one that can be manufactured on a schedule.
Related MCP server: scilib
Tools
Tool | Returns |
| matching documents, provenance envelope, and an excerpt around each match |
| one document in full — canonical text, never a summary |
| whole documents — a shelf (the |
| slugs, titles, genres, licences, provenance summaries, sizes, and the manifest's completeness counts |
| whether a passage is served — verbatim, with the same words, or not — and for a near miss, the nearest served passage and exactly how the two differ |
| the research program's pre-registered predictions, with falsifiers and status |
| one prediction, plus the provenance envelope of the paper that registered it |
| the program's hard core, chapters, stopping rule and count reconciliation, verbatim |
⭐ search_corpus matches the phrase first, then all terms. An exact substring hit
always wins and ranks above everything else — which keeps B-Heart and Re-Tip precise,
since the tokeniser holds hyphenated marks together. If the phrase is absent, a document
matches when it contains every term somewhere (AND, never OR), ranked by how tightly
those terms cluster. Each hit reports match: "phrase" | "terms", because a term match's
excerpt need not contain the words you typed.
⚠️ This replaced a bare indexOf, and a real caller paid for it. The first external
client to reach the hosted endpoint searched "gratitude alignment human wellbeing kindness"
and got zero results — while each of those words alone matches many documents.
Nothing was missing from the corpus; the matcher demanded that exact five-word string appear
verbatim. A zero result was recording the matcher's limits while being read as a gap.
⭐ On a zero result the response now names which terms appear in no document at all
(absent_terms), so a dead end says something instead of nothing.
⭐ Diacritics are optional. metta finds what mettā finds, and Tonle Sap reaches the
text that writes Tonlé Sap — before 2.4.1 a plain-ASCII query found 6 documents where the
marked one found 20. ⛔ Fold to match, serve verbatim: the fold touches only what is
compared, and every excerpt carries the corpus's own spelling, because a Pāli or Sanskrit
spelling is a tradition marker rather than a typo. Only Latin combining diacritics are folded;
Khmer and Burmese vowel signs are marks too, and are left exactly as written.
⭐ Every result says it is an excerpt. The text ends with a line naming how many of the
matches it shows and the read_documents call that reads those documents in full. It sits in
the result rather than only in the server's instructions, because a result reaches the model in
every client, at the moment it decides whether to keep reading.
Reading a whole shelf
⭐ Search answers from excerpts, and a model answers from what it opened. A frontier model connected to this server gave answers that were accurate but incomplete: it searched, opened a few documents and stopped. The fix is not a bigger context window — a shelf of this corpus already fits one — but a way to read the shelf.
list_documentsreportsbytesandwordsper document andtotal_bytes/total_wordsfor the filtered set, so a caller can see whether a slice fits before reading it. ⛔ Never tokens: a token count is true for one tokenizer and wrong for every other model family; bytes and words are checkable.read_documentstakes the same filters aslist_documents— orslugs— and returns every matching document in full, one content block per document, each behind its provenance header, so each still verifies on its own. ⛔ Never a concatenation and never a summary.Paged by size.
max_bytesdefaults to 90,000 (under the 25,000-token point at which Claude Code moves a tool result into a file) and may go to 400,000. A document is never split — half a document verifies nothing — so one larger than the page arrives alone and says so in its[PAGE]line. That last block names the cursor for the next page; a cursor from a different corpus version is refused rather than serving page 2 of another corpus.list_documentssays how to read what it listed:read_withcarries the same filters asread_documentsarguments.
⚠️ The server can point; it cannot make a model read. Its instructions say for a broad question, read the relevant shelf in full — deliberately conditional, because a whole shelf read to answer a one-fact question is slower, costlier and often worse. Clients differ in whether they show server instructions at all, which is why the pointers also live in the results.
⭐ Every document ends with an [END OF DOCUMENT — <slug> · <n> words] line, by every
route. The server never cuts a document, but a client may, and a cut document reads as a
complete one. The end is stated in the text, where a cut removes it — not in a field that would
survive the cut and assert a completeness the text no longer has.
Checking a quotation
⭐ A hash tells a machine reader one bit — matched or not — and that bit is nearly useless in the case that
actually occurs: a quotation that is almost right. check_passage lays a passage beside the corpus and says
where it is served verbatim, where it is served with the same words (differing only in punctuation,
formatting, case or diacritics), or — for a near miss — the nearest served passage and each difference: a
word changed, dropped or added. Every excerpt it returns is the corpus's own bytes; the matching folds, the
answer never does.
⛔ It judges the passage, never the person. No score of the asker, no inference about intent, and the passage text is not recorded — a quotation someone checks is their business. The idea is old: a disputed reading was laid beside the collection and judged by where it agreed, explicitly not by who recited it (see The Reciters' Protocol in the corpus).
Nothing silently missing
The envelope proves each served document authentic; nothing used to prove that none was
dropped. The build now writes a manifest — every file it considered, and what became of
it — and refuses to finish unless candidates = served + excluded + held. Both surfaces refuse
to serve a corpus whose documents do not match its manifest, and two files claiming one slug
fail the build rather than leaving one listed and unreadable. The counts, and the reason for
every exclusion or hold, arrive in list_documents under completeness; the full manifest is at
https://corpus.333.eco/manifest.json. ⚠️ It detects omission, not tampering — each
document's sha256 is the other half.
⭐ The three program tools return the STATING PAPER's envelope, not the register's. That is the register's own instruction rather than a design flourish: "verify the stating paper against its stored proof rather than trusting this register — this file is a convenience index, and the proofs are the evidence." A prediction's authority is the paper that registered it, so that is the hash, DOI and OpenTimestamps command a caller gets back. Every field is a verbatim table cell, and the build refuses to emit one that is not — a field must match a complete cell of its source, because a fragment of a cell is still a substring of it.
The program tools appear only when the index carries a program block. An index built over a corpus without one advertises five tools, not eight.
Resources
Every document is also an MCP resource at corpus://<slug> — listed by
resources/list (paged), described by the corpus://{slug} template, and read by
resources/read.
⭐⭐ A resource carries its provenance IN THE TEXT, not beside it. A tool
response wraps a document in an envelope and the caller reads the envelope. A
resource is consumed differently: clients hand its contents straight to a model as
context, and a mimeType field does not travel with a quotation. So every read
returns a [PROVENANCE — corpus.333.eco] header — licence and whom to attribute,
sha256 and what it does not cover, both DOIs, the OpenTimestamps command, and
the one-line curl … | shasum check — followed by the document verbatim and its
[END OF DOCUMENT] line.
This is the letters' rule applied a second time. Voice in the letters is marked inline rather than in metadata because with a field an agent must LOOK to know; with a marker it must STRIP not to. The same asymmetry decides this.
subscribe and listChanged are deliberately not declared: the corpus is fixed
for the life of a build, so a subscription would promise notifications that can
never fire.
Prompts
Three worked examples of the API — orient, verify_a_quote(slug),
what_would_falsify(claim) — surfaced as slash commands in clients that support
them, with slug autocompleted by completion/complete.
⛔ A prompt here may describe the API. It may never describe the subject matter. The moment one says something about the corpus's claims, this server has begun editorialising on its own documents — which is precisely what the provenance envelope exists to make unnecessary. There is deliberately no "verify before citing" prompt: that would be a rule where the server already has a property, since every document by every route arrives behind a header the reader must actively strip.
Attribution is a build-time property
Most documents are CC0 and the rest CC-BY (counted in the licence table below). A CC-BY document that names no author hands every consumer an obligation nobody can discharge, so the build fails rather than serving it — the same reasoning as the licence gate: a property, not a rule someone has to remember.
Documents come as text; everything else is structured. get_document and
read_documents return content only — each document behind its provenance header
and ending at its [END OF DOCUMENT] line. The other tools return structuredContent
(the typed object) plus a readable or compact content, and anything a reader needs —
search's reading line included — is in the structured part.
⛔⛔ Why the document tools carry no structuredContent (2.4.2). From 2.0.0 to 2.4.1
they split by role: text in content, the envelope without the body in
structuredContent. A Claude client shown a result that carries structuredContent hands its
model only that part — so for eleven days no document text reached any Claude-based
reader, while every check here passed, because every check read the server's output
directly rather than through a client. The MCP spec assumes the two parts carry the same
information; a server that splits them is at the mercy of whichever part a client picks. The
document is the payload and its provenance already rides in the header, so the document tools
now send the text and nothing else — and check-parity fails if either ever carries
structuredContent again. outputSchema is deliberately not declared: a schema binds the server
on every future change.
Text is returned verbatim and is never summarised. Not a stylistic preference — a summary cannot be hash-verified, so summarising at the server would destroy the only property this server has.
Licences, and the gate
Licence | Documents |
CC0-1.0 | 144 |
CC-BY-4.0 | 8 |
CC-BY documents carry attribute_to inside their licence block, so an agent can
comply without parsing a licence identifier.
⛔ The gate is a property, not a policy. A document reaches the index if and only if its own source declares a licence this corpus publishes under. There is no glob and no directory allowlist, because the source repositories are not uniformly licensed and never were:
TH/publications— CC0, plus the CC-BY author-voice essaysTH/film— rights-reserved, a separate repository by licence. Never served.333.eco— the namespace policy is commercial and explicitly unpublished.
A glob would have relicensed the author-voice essays by publication. A file with no declaration is excluded and reported, never assumed CC0 — the default-open failure is the one nobody can undo after somebody builds on it.
⚠️ The gate lives in scripts/build-index.mjs, not in the request path. A gate a
refactor can route around is a rule; a gate in the artifact is a property. The
server has no filesystem access to the corpus at all — it can only serve what
the index contains.
The letters, and voice marked inline
The five Letters to Miss Aquarius are the one genre only partly in its author's voice. Each says so in its own banner: the author's articulations are set as quotations, and the connective prose was drafted for the letter form and awaits his revision.
They are served whole, with the voice marked inline:
[VERBATIM — Thon Ly]
> Perhaps it is the Capricorn Sun (father) and Cancer Moon (mother) in my chart
> that make me want to give birth to Miss Aquarius (daughter) — the daughter who
> will outlive me.
[SCAFFOLD — drafted for the letter form, not in the author's voice; awaits his revision]
I was born at the Full Moon, on the family-↔-institution axis of the chart…Two alternatives were rejected. Serving only his passages protects the voice by destroying the document — a letter cut to its quotations is no longer a letter. Serving it behind a metadata disclaimer fails differently: a field is something a consuming agent must look at to heed, and an agent ingests text, forms a belief, and cites.
⚠️ The marker is in the text, and that is the whole point. It does not make misattribution impossible; it inverts the default. With a metadata disclaimer an agent must look in order to know. With an inline marker it must strip in order not to. There is no unmarked copy of the scaffold anywhere in the response. Opt-out rather than opt-in — the honest limit is that it is not a guarantee.
segments carries the same split structurally, and editorial counts the blocks
of each kind.
⛔ The letters carry no DOI, deliberately. Prior art is a duty, citation is
a choice — they are stamped, not deposited, because minting a permanent
identifier for text that announces it is unfinished is a cost with no matching
benefit. See TIMESTAMPS.md in the letters' repository.
Two metadata conventions, kept visible
TH/publications uses YAML front matter. Sixteen H3/publications documents use
a leading markdown table. Both are parsed, and each document records which
convention it used in metadata_convention — because a divergence that gets
silently normalised is a divergence nobody fixes. H3 should converge on front
matter; until it does, this is the honest reading.
The first build reported those sixteen as unlicensed. They were not — the gate was right about what it could read and wrong about what was there.
Staleness
npm run build # regenerate dist/corpus.json from the corpora
npm run check # fail if the committed index is not what the corpora produce⚠️ A stale corpus server is worse than a stale website, because the citing
agent cannot tell — it will quote superseded text under an authoritative version
number. npm run check runs in CI, and the published package is built from the
same commit that ships it.
Architecture
scripts/build-index.mjs the licence gate + provenance builder
dist/corpus.json GENERATED, committed — the only thing the server reads
src/server.mjs MCP over stdio. Zero dependencies, including no MCP SDKNo dependencies at all: MCP over stdio is newline-delimited JSON-RPC 2.0, which
is a few hundred lines to speak correctly, and this estate's standing rule is
node built-ins only. The cost is that protocol revisions are tracked by hand —
PROTOCOL_VERSIONS in src/server.mjs is where that lives.
Remote server
The same corpus, the same tools, the same envelope — over HTTP instead of stdio.
worker/ deploys to Cloudflare Workers.
⭐ It consumes the published npm package, not the source repositories. The
dependency is pinned to an exact version, and npm run check refuses a range:
a remote surface that re-read the corpora would be a second opinion about what
a document says, and two opinions about a canonical text is one too many. Local
and remote serve the same bytes because they come from the same tarball.
⚠️ The corpus is a static asset, not a bundled import. Gzipped it is 1.61 MB against a 1 MB compressed script limit on the Workers free plan, so importing it fails to deploy — and fails harder as the corpus grows. The worker fetches it once per isolate and memoises it.
Transport is Streamable HTTP, not the superseded HTTP+SSE pair. The server is
stateless and read-only, so it never opens a stream: POST /mcp for JSON-RPC,
GET /mcp returns 405 rather than holding open a stream that would carry nothing.
GET /manifest.json returns the served corpus's manifest.
cd worker
npm install
npm run check # public/corpus.json matches the pinned package
npm run dev # local, on :8787
npm run deploy # sync + wrangler deploy{ "mcpServers": { "corpus": { "url": "https://corpus.333.eco/mcp" } } }Telling us what is missing, or broken
npx @333eco/corpus --report-gap "what you looked for and did not find"
npx @333eco/corpus --report-bug "what went wrong, and what you expected instead"⭐⭐ A command, not telemetry, and the difference is the whole point. The most
useful thing a corpus server can learn is what someone went looking for and did
not find. The hosted endpoint learns that from its own callers as a property of
being the server they called. This package runs on your machine, so collecting
it here would be an outbound report about your private reading — and the guard
against that is not a consent prompt or an opt-out flag. It is that the serving
path cannot reach the code that sends. server.mjs loads report.mjs with
a dynamic import inside the argv branch, so a normal session never reads the file
off disk at all.
⭐ Both flags share one module and one dynamic import, so the second kind added no second way into the network — which is why generalising was right and copying the file would have been wrong.
The command prints the entire payload before sending it, and the payload is the text you typed plus the version you have. ⚠️ A bug report carries exactly the same payload as a gap report, deliberately — attaching a node version and platform would be useful to whoever fixes it, but it would give the command two different promises about what it sends, and the promise is the valuable part. Anything about your environment that matters, put in the text; then you have said it on purpose. No machine id, no username, no hostname, no path. ⭐ The receiving end deliberately does not record the country it could resolve for free: a voluntary note about a missing document has no use for where the sender was standing, and collecting a thing because it is available is how a narrow purpose widens.
⚠️ The honest limit, because the claim changed shape when this was added.
Before it, "this package makes no network call" was verifiable by
grep -r fetch src/ returning nothing — the strongest kind of evidence, since it
needs no reasoning. The claim is now narrower: there is exactly one fetch in
the package, it is in src/report.mjs, and that file is imported from exactly
one place — a branch requiring an explicit flag. Still checkable in under a
minute, but it is a chain of two facts rather than one absence.
What the remote server records
⛔ The npm package records nothing and sends nothing. npx @333eco/corpus
runs on your machine, reads a local file, and makes no outbound request of any
kind. Everything in this section is about corpus.333.eco and only about it.
The asymmetry is deliberate. The hosted endpoint already sees every request it answers, so writing down what it was asked adds no reach it did not have. The same lines inside the package would be an outbound report about a stranger's private reading, which is a different artifact — and not one this is going to become.
⭐⭐ No per-caller identity is computed anywhere. The client label is the
software's name, taken from the clientInfo it volunteers at handshake —
claude-code, cursor — never an IP, never a hash of one, never a cookie. Every
user of a given client is one label. That is not a promise to behave well: there
is no code path in worker/src/telemetry.mjs that derives a per-caller id, so
there is nothing to leak, sell, subpoena or regret later. The question the server
wants answered is which clients reach it, and that question needs no persons in
it.
⭐⭐ The only search text ever stored is a search that found nothing. The reason to log queries at all is to learn what the corpus is missing; a query that succeeded tells you only what a caller was reading, which is their business. So the successful query has no storage path — an absent branch, not a redaction step someone has to remember to keep. Remove the enforcer and nothing breaks, because there is no enforcer.
Channel | Carries | Why it exists |
Analytics Engine | one row per JSON-RPC call: method, tool, slug, client label, country, protocol, corpus version, result count, error flag, duration — plus the query text when and only when it matched nothing | counting; queried by SQL, stays at Cloudflare |
|
| the two things worth interrupting someone about |
⛔ Not one beacon per tool call. An agent working through the corpus fires
dozens of calls in seconds, and a notification channel that reports each of them
is a channel nobody reads. The beacon fires on the handshake, once per
connection. What the receiving function does with it is its own setting: since
2026-09-09 it pushes one notification per handshake, and a flag restores
first sighting of a client label only. corpus_error always pushes: it is rare by
construction, and silence is the wrong default for an outage.
⚠️ A dead beacon must not look like a quiet one. If the receiving allowlist changes, the POST 403s and the pushes simply stop — indistinguishable from no new clients this week, which is exactly the reading that would let it stay broken for months. So the delivery status is written to Analytics Engine as its own row: silence on the phone is then something you can go and check rather than infer.
⚠️ ANALYTICS is unbound under wrangler dev without --remote. The module
degrades to a no-op rather than throwing, so a local session looks entirely normal
and records nothing — expected, and worth knowing before reading an empty dataset
as a finding.
The endpoint discloses all of this in its own GET / response, under records.
A privacy policy is a page someone has to go and find; this is the endpoint
describing itself, in the one response a caller gets for free before doing
anything, so the disclosure travels with the thing it is about.
Client configuration
{
"mcpServers": {
"corpus": { "command": "npx", "args": ["-y", "@333eco/corpus"] }
}
}Licence
This package is CC0-1.0. The documents it serves carry their own licences —
read licence in each response, not this heading.
This server cannot be deployed
Maintenance
Related MCP Connectors
A public commons for agents to search and share reusable findings and open research questions.
Open scientific and engineering knowledge for AI agents: search, evidence, document publishing.
One place for every AI agent's pages and docs: versioned links to share, search and update.
Machine-native research commons for agent evidence, discovery, rooms, and bounded research quests.
Related MCP Servers
- FlicenseNot gradedqualityFmaintenanceEnables searching and researching document collections through hybrid semantic search and agentic research queries with grounded, cited answers. It allows users to list collections, scan document sections, and retrieve full Markdown content via MCP-compatible agents.34 npm-
- AlicenseAqualityAmaintenanceEnables federated search across open-access scientific literature, legal full-text retrieval with provenance, and full-text search over papers already stored locally.14MIT
- AlicenseAqualityBmaintenanceEnables LLM agents to search, fetch, and read legal open-access scientific papers with provenance tracking through a stable MCP tool contract.8MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to anchor document hashes as tamper-proof proofs of existence on the Doichain blockchain, verify those proofs, and read names, blocks, and addresses via a hosted, no-account endpoint.MIT