Corpuscribe
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Corpuscribecompare phrasing of 'may underestimate' and 'could underestimate' in my literature"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Let your literature inform your language
How do researchers express uncertainty? Is a formulation common across several papers, or repeated in just one? What wording accompanies a particular kind of estimate?
Corpuscribe brings those questions into your AI writing workflow. Its local MCP server searches a prepared literature database, compares formulations, and returns sentences with their article and page references. Your writing assistant uses those examples to suggest revisions you can inspect.
Find the wording | Keep the context | Check the evidence |
Compare formulations across articles. | Filter examples using existing Jev classifications. | Return article identifiers, pages and source identifiers. |
Early release · v0.1. Three read-only tools are available today. PDF ingestion, paragraph revision and automatic classification are outside this server's current scope.
Related MCP server: ares-mcp
From literature to a revision
flowchart LR
A[Prepared literature corpus] --> B[(Local SQLite database)]
J[Optional Jev classifications] --> B
B --> C[Corpuscribe MCP]
C --> D[Your writing assistant]
D --> E[Revision with sourced examples]
classDef paper fill:#F4F0E8,stroke:#182C32,color:#182C32
classDef accent fill:#182C32,stroke:#182C32,color:#ffffff
class A,B,J,D,E paper
class C accentThe server queries saved data. It makes no model calls and requires no API key. An optional upstream classifier can attach rhetorical labels before you search.
Quick start
Use Node.js 22 or newer for the documented setup. The repository's CI tests Node.js 22.
git clone https://github.com/tofunori/corpuscribe.git
cd corpuscribe
npm ci
npm run buildPoint the server at an existing compatible database, or try the included fictional two-sentence corpus with Python 3:
python3 - <<'PY'
import sqlite3
from pathlib import Path
path = Path('data/demo.sqlite')
path.parent.mkdir(exist_ok=True)
if path.exists():
raise SystemExit('Demo database already exists; use it without recreating it.')
with sqlite3.connect(path) as db:
db.executescript(Path('examples/demo.sql').read_text())
PY
node dist/server.js --database "$PWD/data/demo.sqlite"The server uses stdio: it waits for an MCP client rather than opening a webpage or an interactive terminal prompt.
Connect your assistant
Add this entry to a client that accepts an mcpServers configuration, replacing the two paths with absolute paths on your machine:
{
"mcpServers": {
"corpuscribe": {
"command": "node",
"args": [
"/absolute/path/to/corpuscribe/dist/server.js",
"--database",
"/absolute/path/to/linguistic.sqlite"
]
}
}
}You can alternatively supply the database path through CORPUSCRIBE_DB. The server does not automatically load .env files. Client configuration formats vary; the JSON above is a configuration example, not an automatic installer.
Three tools, one corpus
Tool | What it returns | Example input |
| Eligible sentence, article and source counts; stored Jev coverage. |
|
| Up to 50 matching sentences with provenance and optional label filters. |
|
| Matching sentence counts, distinct article/source counts and up to three examples per term. |
|
Find cautious language
Call corpus_search with:
{
"query": "underestimate",
"limit": 5,
"rhetoricalFilters": [
{ "label": "certainty", "choice": "possible", "minProbability": 0.9 }
]
}The fictional demo corpus returns this sentence inside the tool's result:
{
"id": "demo-1",
"article": "demo:paper-a",
"title": "Fictional snow study",
"pagePdf": 4,
"sourceId": "demo:source-a",
"text": "These assumptions may underestimate melt.",
"wordCount": 6
}In a writing session, try: “Find examples of ‘may underestimate’, show their sources, and use them to suggest a cautious revision of my sentence.” The assistant proposes the revision; Corpuscribe supplies the examples.
Read the counts correctly
Search is case-insensitive substring matching, not semantic retrieval or whole-word matching. PDF line breaks can affect phrase matches.
occurrencescurrently counts matching sentences, not individual repetitions of a term.Article counts help distinguish broad usage from repeated wording in one source. Source duplication depends on the upstream database.
A common expression is evidence of usage, not proof of scientific correctness or a universal style rule.
Jev labels are optional, probabilistic and dependent on upstream coverage. This version does not select a classifier schema: use an index containing one intended schema to avoid mixing label versions.
See the database contract for required columns and coverage definitions.
Your corpus stays under your control
The SQLite connection is read-only. The repository contains code and fictional examples; personal databases, PDFs and environment files are ignored by Git.
Retrieved sentences are returned to your MCP client. If that client uses a hosted model, the excerpts it includes in its prompts may be sent to that provider. Choose a client appropriate for your corpus.
Development
npm test
npm run check
npm run buildTests use a synthetic SQLite fixture. See Contributing for reporting problems and proposing changes.
Next directions
Linguistic feature queries for verbs, adverbs and connectors.
Richer classification provenance and schema selection.
Corpus-aware paragraph diagnostics.
Easier corpus preparation and onboarding.
These are planned directions, not tools available in v0.1.
This server cannot be deployed
Maintenance
Related MCP Connectors
Academic paper search, scientific literature, citation analysis, arXiv & semantic related-work.
Search 340M+ academic papers — citation graphs, semantic similarity, and AI literature reviews.
Federated search of books and papers, BibTeX/RIS citations, open-access retrieval and reading.
- TBRP CodexOAuthapp.tbrp
Search, read and cite a researcher's own library and the canonical corpora on tbrp.app.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables semantic search and conversational querying across a personal research library of PDFs, DOCX, and other documents using a vector database. It provides tools for document summarization, finding related papers, and high-accuracy retrieval for AI clients like Claude Desktop.-
- AlicenseAqualityBmaintenanceProvides local academic literature search and writing support by querying OpenAlex, arXiv, and Crossref, downloading open-access PDFs, extracting IMRaD sections, and appending BibTeX references.42MIT
- AlicenseAqualityBmaintenanceEnables AI writing assistants to retrieve citable evidence from local PDFs and verify draft citations against their sources locally, providing verifiable support for claims.21MIT
- AlicenseNot gradedqualityCmaintenanceEnables coding agents to verify verbatim quotes against a local corpus of scientific PDFs, returning exact matches with page locators or typed refusals without using an LLM.AGPL 3.0