Skip to main content
Glama
README.md
# Lector

**Agentic RAG over the EU AI Act.** Ask in English or Italian; every sentence of the
answer is cited to the exact article, paragraph or point of Regulation (EU) 2024/1689,
and every citation is checked against the text the agent actually retrieved.

```
$ lector ask "Can an employer use AI to read the emotions of its staff?"
```

The answer comes back with citations like `[Article 5(1)(f)]`, each linked to EUR-Lex,
plus the tool calls the agent made, tokens, cost and latency.

- **Agent** — a LangGraph state machine around Claude's tool use: it searches, reads
  whole articles when a passage is cut, follows cross-references (Article 6 → Annex III),
  and has to back every citation with retrieved text before the answer leaves the graph.
- **Retrieval** — structure-aware chunks (one per paragraph or point, never across
  them) with multilingual embeddings in pgvector. BM25, hybrid fusion (RRF) and
  cross-encoder reranking are built in and measured against dense retrieval; the
  default is whatever the evaluation says works best.
- **Knowledge graph** — 398 cross-references between articles and annexes, extracted
  from the text and exposed as a tool.
- **MCP server** — the same tools for Claude Code, Claude Desktop or any MCP client.
- **Measured** — 50 hand-written questions with the articles that answer them; hit@k and
  MRR for every retrieval mode.

## How it works

```mermaid
flowchart LR
    Q([question]) --> G[guard<br/>injection, length, language]
    G -->|blocked| E([end])
    G --> A[agent<br/>Claude + tools]
    A -->|tool calls| T[tools<br/>search · get_article · get_references]
    T --> A
    A -->|answer| V[verify<br/>every citation retrieved?]
    V -->|no: one revision| A
    V -->|yes| E
```

| Step | What happens |
|---|---|
| **guard** | Rejects empty, oversized and prompt-injection inputs before any model call; detects the language so tools search the English or Italian text. |
| **agent** | Claude (`claude-opus-5` by default) decides whether to search, read an article in full or follow its references, for at most 6 tool rounds. |
| **tools** | Run against Postgres; every unit shown to the model (`art5.1.f`, `anxIII.4`, `rct27`) is recorded. |
| **verify** | Extracts the `[citations]` from the answer. If one was never retrieved, or there are none, the agent gets one revision round; if it still fails, the answer is returned flagged `grounded: false` with the offending ids. |

Conversations are kept per `thread_id` with a LangGraph checkpointer, so follow-up
questions work. Each response reports input/output/cache tokens and cost in USD.

### Ingestion

The official XHTML from the EU Publications Office gives every article, paragraph,
recital and annex a stable element id. `corpus.py` walks that structure and emits:

- **306 documents** per language: 113 articles, 180 recitals, 13 annexes;
- **~1,050 chunks** per language: one per paragraph, split into its points (a), (b)…
  when a paragraph is too long to embed, each point keeping the paragraph's lead-in so
  it still reads on its own;
- **citation labels** in both legal styles: `Article 5(1)(f)` and
  `Articolo 5, paragrafo 1, lettera f)`;
- **cross-references**, with references to other acts (`Article 16 of Regulation (EU)
  2016/679`, `Article 114 TFEU`) filtered out.

## Evaluation

`eval/questions.jsonl`: 50 questions (25 English, 25 Italian) written the way a
compliance officer or a developer would ask them — *"Is an AI tool that screens job
applications considered high-risk?"*, *"Il fornitore deve monitorare il sistema anche
dopo averlo immesso sul mercato?"* — each with the articles, annexes or recitals that
answer it. A question counts as a hit at *k* if one of them is in the top *k* results.

| Retrieval | hit@1 | hit@3 | hit@5 | hit@10 | MRR@10 | median ms |
|---|---:|---:|---:|---:|---:|---:|
| keyword (BM25) | 0.48 | 0.78 | 0.84 | 0.92 | 0.641 | 3 |
| **vector** (default) | **0.64** | **0.92** | **0.94** | **0.98** | 0.780 | 107 |
| hybrid (RRF) | 0.52 | 0.86 | 0.92 | 0.96 | 0.698 | 118 |
| hybrid + rerank | 0.66 | 0.90 | 0.92 | 0.98 | 0.790 | 4579 |

`intfloat/multilingual-e5-large` embeddings, `jina-reranker-v2-base-multilingual`
reranker, CPU only (Ryzen 9 7940HX). Reproduce with `lector eval retrieval`.

- **Dense retrieval is the default.** It puts a correct article in the top 3 for 92% of
  the questions, and the agent reads 8 passages per search.
- **BM25 is computed on the lexemes Postgres stems for English and Italian**, with
  inverse document frequency — Postgres' own `ts_rank` has none, so "AI" and "system"
  would weigh as much as "subliminal". It is strong for exact terms but weaker on
  paraphrased questions, and in an equal-weight fusion it pulls the dense ranking down.
- **Reranking** adds +0.01 MRR for 40× the latency on CPU, and the only multilingual
  cross-encoder available is CC-BY-NC, so it is off unless `LECTOR_RERANK_MODEL` is set.

## Quickstart

Requirements: Docker, Python 3.11+, and an [Anthropic API key](https://console.anthropic.com)
for the agent (retrieval, the API's `/search` and the MCP server work without one).

```bash
git clone https://github.com/namespaceMarcello/lector && cd lector
docker compose up -d db                  # Postgres 17 + pgvector on localhost:5433
python -m venv .venv && . .venv/bin/activate
pip install -e ".[dev]"
lector ingest                            # loads the pre-computed vectors from data/
export ANTHROPIC_API_KEY=sk-ant-...
lector ask "Quali pratiche di IA sono vietate?"
```

The first run downloads the embedding model (`intfloat/multilingual-e5-large`,
2.2 GB) into the fastembed cache; it is needed to embed queries. Passage vectors for the
whole Act ship in `data/`, so ingestion takes seconds instead of ~20 minutes per
language on CPU (`lector ingest --recompute` rebuilds them).

Or everything in containers:

```bash
ANTHROPIC_API_KEY=sk-ant-... docker compose up -d
docker compose run --rm api lector ingest
curl -s localhost:8000/ask -H 'content-type: application/json' \
     -d '{"question": "Do chatbots have to disclose that they are AI?"}'
```

## Usage

```bash
lector search "real-time remote biometric identification" -k 5   # retrieval only
lector search "sistemi ad alto rischio" --lang it --mode keyword
lector ask "Which obligations do deployers of high-risk systems have?" --json
lector eval retrieval                                            # the table above
lector eval answers --limit 10 --out results/answers.json        # needs an API key
lector serve                                                     # http://127.0.0.1:8000/docs
lector fetch                                                     # re-download and re-parse the Act
```

### HTTP API

| Method | Path | |
|---|---|---|
| `POST` | `/ask` | `{"question", "thread_id"}` → answer, citations with labels and EUR-Lex links, `grounded`, usage, latency, trace |
| `GET` | `/search?q=&lang=&k=&mode=&rerank=` | passages only, no LLM |
| `GET` | `/documents/{id}` | an article, annex or recital (`art6`, `Annex III`, `rct27`) with its cross-references |
| `GET` | `/health` | chunk counts per language |

### MCP

```bash
claude mcp add lector -- lector mcp
```

Tools: `search_regulation(query, lang, k)`, `get_article(doc_id, lang)`,
`get_references(doc_id, lang)`. The client's model does the reasoning; Lector supplies
the passages and the citation ids.

## Configuration

| Variable | Default |
|---|---|
| `ANTHROPIC_API_KEY` | — |
| `LECTOR_MODEL` | `claude-opus-5` (server-side refusal fallback enabled) |
| `LECTOR_DATABASE_URL` | `postgresql://lector:lector@localhost:5433/lector` |
| `LECTOR_EMBED_MODEL` | `intfloat/multilingual-e5-large` |
| `LECTOR_SEARCH_MODE` | `vector` (`hybrid`, `keyword`) |
| `LECTOR_RERANK_MODEL` | empty (off); e.g. `jinaai/jina-reranker-v2-base-multilingual` |
| `LECTOR_MODEL_CACHE` | fastembed default |

## Project layout

```
src/lector/
  corpus.py      fetch + structure-aware parsing + cross-references
  store.py       Postgres: pgvector HNSW, BM25 on stemmed lexemes, reference graph
  retrieve.py    dense / BM25 / hybrid (RRF) retrieval, optional reranking
  tools.py       search_regulation · get_article · get_references
  agent.py       LangGraph: guard → agent ⇄ tools → verify
  llm.py         Anthropic SDK, cost accounting, refusal fallback
  api.py         FastAPI
  mcp_server.py  MCP server (stdio)
  evaluate.py    retrieval and answer metrics
eval/questions.jsonl
data/            parsed corpus and cached passage vectors
tests/           parser, graph, agent (scripted model), API, MCP, Postgres
```

## Tests

```bash
pytest -q                                   # no network, no API key, no model downloads
LECTOR_TEST_DATABASE_URL=postgresql://lector:lector@localhost:5433/lector_test pytest -q
```

The agent tests replace the Anthropic client with a scripted one, so the graph —
tool loop, citation checks, revision round, tool budget, per-thread memory, refusal
fallback handling — is tested deterministically.

## Data

The text of Regulation (EU) 2024/1689 in `data/` is © European Union
([EUR-Lex](https://eur-lex.europa.eu/eli/reg/2024/1689/oj)), reused under Commission
Decision 2011/833/EU. Lector is not legal advice.

Built by Marcello Costagliola with Claude Code. MIT license.