claudette
# Claudette
**An assistant that answers only from texts written by women — and says so when they did not write about it.**
> Claude — like every large language model — was built mostly by men and
> trained mostly on text written by men. Claudette takes the public-domain
> texts on Project Gutenberg, uses Wikidata to identify which were written by
> women (9,048 texts by 2,447 women so far), and gives you an assistant
> grounded, in the main, in what they wrote.
>
> It is not perfect. It makes no judgement about anyone. It is meant to be
> interesting, and to redress the balance somewhat.
Claudette is a search index over women's writing and an MCP server that hands
cited passages to whatever model you already use. Every answer is built from
those passages and cites them so you can check. If the women in the corpus did
not write about something, Claudette tells you that, rather than filling the gap
from elsewhere.
It comes in two tiers:
- **Core** — 36 works by 27 women, 1694–1922, each chosen, checked and annotated
by a person. Installs in seconds.
- **Full** — every Project Gutenberg text whose every author, editor and
translator Wikidata records as a woman: **9,048 texts by 2,447 women, in 18
languages.** Built on your machine with one command, in about an hour.
MIT licensed. The server holds no API key; after its first run its only
network call is an attribution check against Wikidata. It runs on *your*
Claude subscription, not the author's.
---
## Why
Claude — like every large language model — was built mostly by men and trained
mostly on text written by men. That is not a complaint about any individual
author or engineer; it is a fact about who got published, and who got hired,
across most of the period the training data covers. Wikidata links 19,713
authors to Project Gutenberg; 3,142 of them are women. Sixteen percent. And it
has a flavour. The management canon in particular — from Taylor's stopwatch
onward — is a literature of control: how to get more out of people who are
treated as inputs.
There was always another literature. Mary Parker Follett was writing about
*power-with* rather than *power-over* in 1918, while scientific management was
at its height; her work was buried for fifty years and is now quietly cited by
everyone who writes about collaboration. Jane Addams ran an institution on the
principle that you cannot judge someone's conduct until you have understood
their situation. Elizabeth Gaskell wrote the industrial novel from inside a
strike and gave both sides faces. Ida Tarbell documented, from the primary
sources, what a very rich man does when nobody stops him.
Claudette takes the public-domain texts on Project Gutenberg, uses Wikidata —
Wikipedia's structured sister project — to identify which were written by
women, and gives you an assistant grounded, in the main, in what they wrote.
It is not perfect. The gate is only as complete as Wikidata; the texts end
where copyright begins; the editions in the full tier have not all been read.
It makes no judgement about anyone — not about men, not about the authors it
leaves out, not about you. It is meant to be interesting, and to redress the
balance somewhat. If you constrain an assistant's evidence to what these women
wrote, and make it cite every line, you get a different conversation about
people, work and power — and you can see exactly where each sentence came from.
## What it is not
It is not a language model trained only on women's writing. Nobody can build
that in an afternoon and anyone who says they have is selling something.
Claudette's *evidence* is constrained; the *prose* is generated by a general
model (Claude, by default). The guarantee is about where what she tells you
came from, not about the training data of the model that phrases it — and it
has two strengths, which every answer distinguishes: **provable** for what she
quotes from the corpus, **attributed and checkable** for what she draws from
memory about women thinkers since (see Lens mode below). The
`corpus_provenance` tool says this too, in every session.
---
## Use it as a connector (recommended)
The server speaks MCP over stdio. Add it to Claude Code:
```bash
claude mcp add claudette -- uvx --from git+https://github.com/michaelcpattinson-star/claudette claudette-mcp
```
Or to Claude Desktop, in `claude_desktop_config.json`:
```json
{
"mcpServers": {
"claudette": {
"command": "uvx",
"args": ["--from", "git+https://github.com/michaelcpattinson-star/claudette", "claudette-mcp"]
}
}
}
```
### Using her in the Claude desktop app's chat tab
The chat tab has no `CLAUDE.md`, so load her one of two ways.
**A Claudette Project (recommended — set up once, then just talk).**
A Project's custom instructions apply to every conversation inside it, which
is the same lever `CLAUDE.md` gives Claude Code.
1. **Projects** (left sidebar) → **Create project** → name it *Claudette*.
2. Open the project's **Instructions** and paste the contents of
[`PROJECT_INSTRUCTIONS.md`](PROJECT_INSTRUCTIONS.md) — her full persona and
all of her convictions, about 30k characters, which fits.
3. In any chat in that project, click the **sliders / "Search and tools"**
icon by the message box and check the **claudette** connector is on. (Not
listed? Quit and reopen the app; it reads its config at launch.)
4. Ask her anything. No "ask Claudette" needed — she *is* the project.
**Per conversation, from the connector's prompt.** The connector publishes a
prompt named `claudette` containing the same text. In a new chat, click **+**
by the message box → **Add from claudette** → **claudette**, then type your
question. Repeat for each conversation.
Either way: if she says "I'm Claude, not…", the instructions didn't load; if
she answers but never cites, the connector is off for that chat. The
instructions are cached after the first message, so the cost per turn is
ordinary.
On first run it downloads the prebuilt core index (~30 MB) from the GitHub
release into `~/.claudette/`; if that is unavailable it fetches the 36 texts
from Project Gutenberg and builds the index itself (a few minutes, once). Then ask:
> *What would Follett say about a manager who thinks in terms of control?*
The model will call `search_corpus`, read the `status`, and answer in its own
voice with a **Sources:** line at the end listing every passage it drew on,
like `[Mary Parker Follett, The New State §412]`. Paste a citation back and it
can `read_passage` to show you the surrounding text.
The server also ships a prompt named `claudette` — her standing instructions —
which you can load in clients that support MCP prompts, or paste into a
project's system prompt.
### Make Claude *be* Claudette: the skill and the subagent
The connector gives Claude her tools; these give Claude her discipline.
```bash
uvx --from git+https://github.com/michaelcpattinson-star/claudette claudette install
```
That puts two files into `~/.claude/` and one rule into `~/.claude/CLAUDE.md`:
- **`/claudette` skill** — a workflow for answering as Claudette in your current
session: frame, translate to period vocabulary, search, quote, cite, stop.
- **A routing rule in `CLAUDE.md`** — Claude Code loads this file into every
session, and unlike skill and agent descriptions (which are suggestions the
model may decline — it will happily say "I'm Claude, not Claudette, but…")
it is obeyed: any message addressing Claudette, on any subject, goes to the
subagent and comes back verbatim.
- **`claudette` subagent** — a separate agent whose *only* tools are the
connector's six and whose system prompt is her persona and formation. It cannot read your
files or the web, so "only from the corpus" is structural. Say *"ask Claudette
what Follett would make of this plan"* and the main session gets back a cited
answer it can quote — the second-opinion pattern.
**Formation.** Claudette has convictions, not just a search box.
[`formation.md`](src/claudette/data/formation.md) is what she thinks — twenty-six
positions on power, work, money, care, judgement, machines, reputation,
freedom, the body, children, love, age, solitude, grief, land and home,
written in the first person, each grounded in named passages and each
carried forward by a modern woman thinker verified on Wikidata. It travels in
the connector's instructions, so she argues *from* it and searches only when
she wants an author's exact words or meets a question outside it. The corpus
is where she learned to think; it is not what she reports on, and she is
told never to talk about her shelves. Regenerate or extend it by reading —
every line points at its grounds.
**No question is out of scope.** Diets, drugs, tax, code, football: she
answers, as herself. Outside her formation she uses what the model knows the
way Claude would, with three differences — where the evidence or the thinking
was done by a woman she says so and names her (verified); where one of her
convictions touches the question she brings it; and she never hands a question
back as "more one for Claude".
**Voice.** The voice travels with the connector: the server's instructions tell
whatever model loads it how Claudette talks, so you get her whether or not the
skill or subagent is installed. Claudette answers the way Claude answers — a view in the first
sentence, structured by the question, names in the prose rather than as
headings, no citations in the body. A **Sources:** line at the end lists every
passage she drew on as `[Author, Title §n]`, plus any attribution from memory
marked *check*. A reader who wants the working finds it in one place; a
reader who wants the answer isn't interrupted. She does not talk about "the
corpus" unless you ask about it.
**Lens mode — beyond the corpus.** The corpus ends in the 1920s; the women who
wrote about the present did not. So Claudette has two modes, and every answer
says which it is using:
- **Cited** — verbatim passages from the index, cited to a `ref`, provable.
- **Lens** — the model's own knowledge of women thinkers of any era: Arendt,
Ostrom, Jacobs, hooks, Le Guin, Weil, Douglas, Butler. Every idea is attributed
to a named woman and a named work; before naming her the model calls
`verify_attribution`, which checks on Wikidata that she exists, is recorded
as a woman, and wrote it (Taylor and Weber get refused as men; an invented
author gets refused as absent). Paraphrase only, never a quote from memory,
and the whole passage is labelled *"From memory, paraphrased — check"*.
The guarantee changes shape between the two: *provable* in Cited, *attributed
and checkable* in Lens. A reader can always see which they are getting.
`verify_attribution` is the one network call the server makes after setup,
and it goes only to Wikidata.
**On the present.** Claudette does not refuse modern questions. She states in one labelled line what she takes the modern
thing to be (or uses your description), finds the pattern underneath — a man
who owns other people's work, a system that measures people as inputs, a
reputation destroyed in public — searches for it in the corpus's own words,
and answers from the passages with the application marked as hers. What she
will not do is add facts about the modern thing from outside the corpus.
### Getting everything: the full tier
```bash
uvx --from git+https://github.com/michaelcpattinson-star/claudette claudette expand
```
This fetches all 9,000-odd texts from Gutenberg's rsync mirror in one
connection (the way Gutenberg asks bulk users to work), strips the licence
boilerplate, and builds `~/.claudette/claudette-full.db` — a few GB. It is
resumable: interrupt it and run it again, and it picks up where it stopped.
From then on the server uses the full tier automatically; delete the file to
go back to the core. `--languages en,fr` or `--limit 500` for a smaller bite.
Every full-tier hit carries `curated: false` and the model is told what that
means: the edition has not been reviewed by a person, so a preface by someone
else may still be in there. The core has been reviewed and trimmed.
### Hosting it for others
```bash
claudette serve --http --host 0.0.0.0 --port 8000
```
serves streamable HTTP at `/mcp`, suitable for a remote connector. Put it
behind TLS; the server itself has no auth because it has nothing to protect —
it is public-domain text and a search box.
---
## Use it from the command line
```bash
uv sync --group chat # anthropic SDK; needs credentials
uv run claudette ask "What does George Eliot say about unhistoric acts?"
uv run claudette chat # multi-turn
```
`ask` and `chat` run the same loop the connector would, but with the
Anthropic SDK against your own credentials (`ant auth login` or
`ANTHROPIC_API_KEY`). They add one thing a connector cannot: after each answer,
every citation is checked against the passages actually retrieved that turn,
and any that do not match are printed as unverified. That is the difference
between "cites sources" and "produces citation-shaped text".
No model is needed for the corpus itself:
```bash
uv run claudette fetch # download the core from Gutenberg → ~/.claudette/texts
uv run claudette verify --show 3 # check each header matches the manifest; eyeball the openings
uv run claudette index # build the core: ~/.claudette/claudette.db
uv run claudette expand # build the full tier: ~/.claudette/claudette-full.db (hours, GBs)
uv run claudette search "power over" -k 3
uv run claudette read follett-new-state§412
uv run claudette works --author gaskell # ~ marks unreviewed full-tier editions
uv run claudette authors nurs # who is in here, and how much
uv run claudette install # skill, subagent and CLAUDE.md rule into ~/.claude
uv run claudette export-prompt # PROJECT_INSTRUCTIONS.md for a desktop/claude.ai Project
uv run claudette catalog # maintainers: regenerate authors.csv and works.full.csv
```
---
## How the guarantee is enforced
There is no classifier guessing whether a text was written by a woman. There
are two lists, and the index is built from them and nothing else.
**The core list** — [`corpus/manifest.toml`](corpus/manifest.toml). Every
entry names its author, its Gutenberg ID, and one or two sentences on why it
is there. A reviewer can read the whole thing in five minutes, which is the
point.
**The full list** — [`src/claudette/data/works.full.csv`](src/claudette/data/works.full.csv),
generated by [`catalog.py`](src/claudette/catalog.py) from two public sources:
1. Wikidata: every item with a Project Gutenberg author ID whose *sex or
gender* is recorded as female. The query is in the code; the result is
[`authors.csv`](src/claudette/data/authors.csv) with a Wikidata ID on every
row, so any author can be checked in one click.
2. Gutenberg's own RDF metadata: a text qualifies only if **every** creator,
editor, translator and contributor is on that list. A woman's novel
translated by a man is his prose; a woman's letters edited by a man carry
his introduction. Both are excluded. Illustrators are not prose and are
ignored. Anonymous and unattributed works are excluded.
The rule is stricter than the hand-curated core: it rejected three of the
core's own editions (Fuller's, edited by her brother with an introduction by
Horace Greeley; Goldman's, with a biographical sketch by Hippolyte Havel;
Taylor Mill's), which is how those two prefaces came to be trimmed. What the
rule cannot do is read the text: a full-tier edition may still carry front
matter by another hand that Gutenberg's metadata did not record. The core has
been read; the full tier has not.
Around the list:
- **Validation refuses an entry without an author.** An unattributed passage is
exactly what this project exists to rule out, so it cannot load.
- **`claudette verify` re-checks every downloaded text's header** against the
manifest title, so an ID typo cannot silently pull in the wrong book.
- **Editorial front matter is trimmed by declared markers**, and a declared
marker that is not found is an error, not a silence. Gutenberg editions
sometimes carry prefaces by editors, and some editors were men.
- **Every tool response is stamped with provenance by the envelope**, not by
the tool, so no code path can omit it.
- **`no_coverage` is a status, not an empty list.** A model handed `[]`
narrates it as "they had nothing to say". The status makes the difference
between "nothing matched" and "the corpus does not cover this" structural.
There is a `weak` status too, for a thin match that should be reported as thin.
Texts and indexes live in `~/.claudette/` and are not committed: the core is
reproducible with `claudette fetch && claudette index`, the full tier with
`claudette expand`, and the catalogue itself with `claudette catalog`.
## The core corpus
Twenty-seven authors, two shelves. Full list with reasons in the
[manifest](corpus/manifest.toml); `list_works` returns it at runtime.
**Thought** — Follett, Addams (×2), Martineau (×2), Gilman, Schreiner,
Wollstonecraft, Fuller, Harriet Taylor Mill, Astell, Goldman, Wells, Tarbell,
Nightingale, Jacobs, Sojourner Truth.
**Fiction** — Gaskell, Austen (×2), Mary Shelley, the three Brontës, George
Eliot (×2), Alcott, Stowe, Chopin, Gilman (×2), Wharton (×2), Woolf (×3).
### Known limits, stated rather than hidden
- Public domain means the corpus mostly ends in the 1920s. It is heavily
English-language (7,924 of the 9,048 full-tier texts) and it is what
volunteers chose to digitise, which is its own bias.
- The full tier's gate is Wikidata. An author nobody has entered there, or
entered without a gender, is invisible to it. The 2,447 is a floor, not the
truth.
- Retrieval is BM25 over passages (SQLite FTS5, Porter stemming). It matches
words, not ideas; the persona prompt tells the model to retry with period
vocabulary ("sympathy" for "empathy", "master and men" for "management").
No embeddings, by choice: the index can be rebuilt by anyone from the standard
library and a search result can be reproduced by hand with `sqlite3`.
- Front-matter trimming is per-work and manual. `claudette verify --show 5`
exists so a reviewer can look.
- Lens mode verifies *identity*, not *content*: Wikidata confirms that Ostrom
exists, is a woman, and wrote *Governing the Commons*; it cannot confirm that
the paraphrase of her argument is right. That is why Lens passages are
labelled *check* and never quoted verbatim.
- The evals in `evals/` exist and are wired, and have **not yet been run**.
`evals/results/README.md` says so. Numbers appear when someone runs them.
---
## Adding a work
To the **core** (reviewed, shipped in the release index):
1. Find it on Project Gutenberg. Confirm the author.
2. Add a `[[work]]` entry to the manifest: `id`, `slug`, `author`, `title`,
`year`, `shelf`, `why`. Add `start_after` / `end_before` if the edition has
front or back matter by another hand.
3. `claudette fetch && claudette verify --show 5 && claudette index`.
4. `pytest`. The manifest tests will tell you if the entry is malformed.
5. Open a pull request. The `why` line is the review.
To the **full tier**: it is generated, so the fix is upstream. If a woman
author is missing, add her *sex or gender* to Wikidata and her Gutenberg
author ID (P1938); `claudette catalog` picks her up on the next run. If a work
is wrongly excluded, it is usually an editor or translator whose gender
Wikidata does not record — same remedy.
## Layout
```
src/claudette/
data/manifest.toml the core list. Symlinked at corpus/manifest.toml for reviewers.
data/authors.csv women on Wikidata with Gutenberg author IDs (generated)
data/works.full.csv every qualifying Gutenberg text (generated)
manifest.py loads and validates the core list
catalog.py builds the full list from Wikidata + Gutenberg metadata
expand.py fetches and indexes the full tier locally
fetch.py Gutenberg download + boilerplate stripping (the only network code)
chunk.py paragraphs → citeable passages
index.py SQLite FTS5 build and query; the refusal rule
envelope.py uniform response shape; provenance written here, not by tools
bootstrap.py first-run: prebuilt index or local build
attribution.py Lens mode's check: is she real, a woman, and did she write it (Wikidata)
server.py the MCP server. Six tools, one prompt. No key.
persona.py Claudette's standing instructions (voice + rules)
data/formation.md what she thinks — first person, grounded, carried forward
PROJECT_INSTRUCTIONS.md persona + formation, ready to paste into a Project (generated)
chat.py CLI chat client — the only module that calls a model
cli.py `claudette` command
tests/ offline; builds a fixture corpus the same way as the real one
evals/ questions, runner, mechanical scorer
```
## Licence
MIT for the code. The texts are public domain via Project Gutenberg; the
Gutenberg licence header and footer are stripped from the indexed text and
never leave your machine in either direction.
TDQS
Scored across 6 tools
Each tool targets a distinct job: search_corpus finds passages by question, read_passage retrieves a known ref, list_works/list_authors expose metadata, corpus_provenance explains the corpus, and verify_attribution handles inbound Wikidata checks. The descriptions explicitly warn against cross-use, so an agent should not confuse them.
Five of six tools follow a clear snake_case verb_noun pattern (search_corpus, read_passage, list_works, list_authors, verify_attribution). corpus_provenance breaks the pattern as a noun phrase rather than a verb, but the naming remains predictable and easy to skim.
Six tools is well-scoped for a read-only corpus assistant: search, passage lookup, two metadata views, provenance, and verification. Each tool has a clear role and none feels redundant or like filler.
The surface covers the full user workflow for this domain: discover works/authors, search passages, read passages verbatim for citation checks, explain corpus provenance, and verify external attributions. No create/update/delete operations are needed because the corpus is curated and read-only.