quant-rag
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@quant-ragWhat do papers in the library say about bid-ask spread components? Include citations."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
quant-rag
Retrieval over a closed library of quantitative-finance research, where every retrieval choice is settled by a paired measurement and every citation can be checked without a language model.
418 documents (papers, working papers, books; 1971–2026) · 26,120 passages · 10 MCP tools · 85 ms median search, warm · 546 tests that run without the corpus
An LLM agent asks questions about market microstructure, volatility models, backtest statistics or portfolio construction. The server answers with passages that arrive whole (never cut inside a formula, a table row or a word), each with a citable source, a character-level anchor into the document's canonical text, and a trace naming the exact server, corpus and configuration that produced it.
The engine itself is deliberately plain: dense retrieval on Qwen3 embeddings in an embedded Qdrant. The work is in the measurement behind it: two benchmark families built against the known-item bias, an LLM judge whose accuracy and noise are measured on every run, placebo arms, paired bootstrap confidence intervals, and a written protocol before each major experiment. Several popular techniques were measured on this corpus, and most of them were rejected.
Results at a glance
Paired differences in nDCG@10 (same questions in both arms), 95 % bootstrap CI;
* = the interval excludes zero. Full protocols and tables: docs/evaluation.md.
Question | What was measured | Decision |
Does hybrid BM25 + RRF + cross-encoder beat dense? | Best configuration on a hand-written known-item bench (0.935 vs 0.853), but −0.140 [−0.240, −0.043]* on the first leak-controlled natural-language bench | dense by default; hybrid only on request, for remembered identifiers |
Can a rule-based router pick the right path? | +0.021 [+0.002, +0.050]* on the 65 calibration questions, then −0.019 [−0.041, +0.002] on the 130-question bench that followed (−0.027 on its 90 unseen questions); a perfect router would add +0.055 |
|
RRF fusion of dense and BM25 (14 variants) | −0.088 [−0.124, −0.054]* | rejected |
| −0.092 [−0.148, −0.033]* | reranking off in dense mode |
| +0.070 [+0.018, +0.125]*, R@1 0.25 → 0.37, at 15 s median per query | not shipped; reranking only the top 10 keeps half the gain at 3.3 s |
Re-embedding every passage with a clean, consolidated title | +0.031 [+0.003, +0.059]*; questions whose gold passage falls outside the top 50: 39 → 26 | adopted |
Filtering by publication year on dated questions | +0.100 [+0.029, +0.184]* | adopted; a period written in the question ("published before 2010") is parsed into the filter: 15/15 detected, 0 false positives on 160 other questions |
Fusing entity-graph hits into the ranking | −0.105 [−0.135, −0.074]* | rejected; the graph stays a separate exploration tool |
Served passages cut inside a formula, table row or word | 178 of 760 → 0, for +14.1 % context | output contract |
Citation check, no language model involved | 50/50 real quotes found, 0/50 one-word alterations accepted |
|
Questions whose answer is not in the corpus | 20/20 abstentions, 0 fabrications | answer prompt |
Related MCP server: finwatch-mcp
Try it in 30 seconds — no corpus, no model download
git clone https://github.com/Elias-outofsample/quant-rag && cd quant-rag
python3 -m venv .venv && .venv/bin/pip install -e . # Python 3.12+
.venv/bin/python rag/demo.pyThe demo runs the real serving functions on one synthetic document (only the storage layer is swapped for an in-memory text):
1. The output contract: a passage arrives whole, or cleanly interrupted
cap 325 (inside the display formula)
raw cut … \kappa (\theta - s_t)\,dt + \sigma\, dW_t, breaks: an open $$ block
contract cut …$ is where the spread returns after a shock. breaks: nothing · 678 characters announced as remaining
cap 752 (inside a table row)
raw cut …A | 0.35 | 0.00 | 0.12 | 1.98 |⏎| B | 0.08 | breaks: half a table row
contract cut …-|---|---|⏎| A | 0.35 | 0.00 | 0.12 | 1.98 | breaks: nothing · 209 characters announced as remaining
2. Quote verification, without a language model
copied from the passage found · offsets [376, 467] · exacte
retyped: spacing, hyphen, capitals found · offsets [376, 467] · exacte après normalisation
one word changed (more -> less) NOT found · closest passage: similarity 0.933, 2 word edits
stitched from two places NOT found · closest passage: similarity 0.833, 4 word edits
3. The query side: publication periods and exact tokens
'volatility targeting, according to papers published between 2015 and 2020'
-> year_min=2015 year_max=2020, embedded query: 'volatility targeting'How it works
flowchart TB
subgraph ingest["1 · Ingestion, offline"]
direction LR
PDF["PDF"] --> P["MinerU parse"] --> K["passages with<br/>section path"] --> M["metadata with provenance;<br/>tables to Markdown"] --> E["Qwen3-Embedding-0.6B"]
end
subgraph store["2 · Corpus state, named by its signature"]
direction LR
Q[("Qdrant, embedded")]
B[("BM25 index")]
G[("GLiNER2 entity graph")]
Q ~~~ B ~~~ G
end
subgraph serve["3 · Serving"]
direction LR
U["question"] --> D["period clause<br/>to year filter"] --> R{"mode"}
R -->|"auto"| DN["dense top-k"]
R -->|"hybrid, on request"| H["BM25 + RRF<br/>+ cross-encoder"]
DN --> S["drop headings and near-duplicates;<br/>max 2 passages per document"]
H --> S
S --> O["output contract:<br/>whole passage, anchor, trace"]
end
ingest --> store --> serve --> MCP["MCP server, 10 tools"] --> A["LLM agent"]Every step is described, with its measured cost, in docs/architecture.md. The promises made on what is served are in docs/output-contract.md.
What a served passage looks like
Real output of search_documents (the server's labels are French: citer = cite as,
ancre = anchor, interne = internal keys, qualité = known defects of the passage):
routage: dense — dense par défaut (calibration v3 : aucune règle ne bat le dense) · rerank: non · dense_top1=0.719 · 12406 ms
trace: request_id=8a7bdb351ec64fbb · serveur 1.4.0 · contrat 1.0.0 · corpus e1bdf36e2e · config aa5531c56ba9 · fenêtre 10000 c.
[1] cosine=0.719
citer : Ding et al. (2025) — Deep Learning Option Pricing with Market Implied Volatility Surfaces, p. 3
section : III. RESULTS
interne : chunk_id=chunk-169b5be73892b3f3 · document_id=doc-96437e7c996bb218 · mixed (clés de travail, ne pas citer)
ancre : doc-96437e7c996bb218 · caractères 10861–13884 (exacte) · sha256:c41346f98382
qualité : has_html_tagsThe 12.4 s of this first call include loading the embedding model; warm searches take
75–89 ms. The ancre line is what verify_citation(document_id, quote) checks a quote
against: character offsets in the canonical text of the document and the SHA-256 of that
text, so an anchor rebuilt on a different text is detected.
The ten MCP tools
Measured on the running server over JSON-RPC.
Tool | What it returns | Latency |
| ranked passages, with source, anchor and trace; filters by year and author | 84 ms |
| one passage with its neighbours, never truncated mid-structure | 363 ms |
| bibliography, filtered by author, period or text | 1.8 ms |
| how a topic evolves, grouped by publication year | 454 ms |
| found / not found, offsets, and the closest match when not found | 83 ms |
| the same, in batch | 156 ms |
| collection, counts, backend, device, contract, provenance | 2.7 s |
| passages mentioning the entities of a query | 288 ms |
| typed neighbours of an entity ( | 17 ms |
| passages and documents where two entities meet | 181 ms |
How the evaluation is built
The first benchmark was the usual one: 25 questions written by hand while looking at the passage that answers them. It measured an IDF-weighted lexical leak of the question into its answer of 0.533 at the median and 0.916 at worst: retrieval on those questions is string matching. It crowned the hybrid pipeline, and that verdict reversed on questions worded the way a practitioner asks them. The benchmarks that followed are built against that bias:
Question factory. An LLM drafts a question from a passage as a practitioner would ask it; a second call rewrites it without seeing the passage; a leak score above 0.42 (the first quartile of the hand-written set) sends it back with the offending terms named; a verifier rejects questions that are unanswerable, not self-contained, or answered equally well by many other passages. The rejection rate is itself a finding: 12 % of table questions survive, against about a third of text questions.
Negative questions name a source that is provably absent: every spelling of it is checked absent from the corpus text. The retriever is only ever used to refute a negative question, never to certify one.
Two score families, never merged. Retrieval (nDCG@10 with graded gains, recall, MRR) needs no judge and stays valid when the judge model changes. Generation (groundedness, coverage of reference facts extracted before any answer existed) is judged; abstention is an exact string match.
The judge is measured on every run: sentinel answers with a known correct grade (verbatim gold, off-topic, fluent but unsupported, wrongful refusal) give its accuracy (0.889 on the v3 baseline), and a double-graded sample gives its noise floor. Generator and judge are different models.
Placebo arms. Regenerating the same configuration left only 55 % of answers identical and produced a "significant" +0.118 [+0.029, +0.235] change in fabrications on one family, with no variable changed. Since then both arms of a comparison are drawn in the same run, and a difference whose interval overlaps the placebo's is not reported as a result.
Instruments are proven before use. The in-memory dense matrix used to run experiments in parallel reproduces Qdrant's top 50, in the same order, on 155/155 questions; replaying the old bench in the new code reproduces the published numbers with a maximum deviation of 0.00.
Written protocol first. The larger experiments start from a dated pre-registration (question, metric, decision rule) and end with a report and a verdict: keep, reject, or keep as experimental.
Engineering
Explicit corpus state. The document set, the overlays (deduplication, Markdown tables) and the consolidated titles hash to a corpus signature. The corpus is frozen between batches; unfreezing requires a written reason, kept in the history. The BM25 index file is named by that signature, and a check proves that production and benchmark serve the same index.
Ingestion without an operator. A batch driver runs eight steps per PDF (parse, survey, diagnosis, import, registry and index updates, entity graph, retrieval probes, commit), holds an exclusive lock, survives a
kill -9, and rolls an import back symmetrically. On a real 6-page paper: import in 25.9 s, zero LLM calls.One home per decision. The Qdrant backend (embedded or server), the served window, the model prices used for cost accounting: each lives in one module, and tests fail if a literal copy reappears elsewhere.
Installation check.
verifier_installation.pynames every artefact the served path depends on, checks coherence rather than presence (index and collection counts against the registry, unresolved LFS pointers), says how to rebuild each one, and exits non-zero if one is missing.Tests. 674 tests; 546 run in CI without the corpus; those that read the private corpus are skipped with the missing path (rag/conftest.py).
Repository layout
rag/
quant_rag.py retrieval: period parsing, routing, dense / hybrid, filters, CLI
mcp_server.py MCP server, ten tools
contrat.py output contract: safe cut, quality flags, anchor and trace lines
citation.py ancrage.py quote verification; passage anchoring in the canonical text
graph_search.py entity-graph queries: mentions, neighbours, meeting points
reranking.py cross-encoder and Qwen3 rerankers
qdrant_backend.py embedded or server Qdrant, one home for the choice
build_index.py rebuild the collection from the export and the overlays
corpus_overlay.py deduplication, tables and titles overlays, applied wherever the corpus is read
gel_corpus.py corpus freeze, with its history
verifier_installation.py what the served path needs, and how to rebuild it
demo.py corpus-free tour of the served surface
ingestion/ parse, inspect, import, roll back; batch driver; registry
metadata/ titles/ tables/ graph/ corpus curation passes
benchmark/ about 100 experiment scripts and their tests (see its README)
agents/rag-scout.md minimal system prompt for the agent consuming the MCP tools
src/ parsing and lexical primitives (see NOTICE.md)
tests/ tests of the two original parsing modules in src/, and of the demo
docs/ architecture, evaluation, output contractRunning it
uv venv --python 3.12 .venv
VIRTUAL_ENV=.venv uv pip install -e ".[dev]"
.venv/bin/python -m pytest -n auto -rs # no corpus needed; corpus tests report what is missingServing queries needs the model extras (.[models]: torch, sentence-transformers) and a
corpus state on disk: parsed documents, the vector export and the overlays under data/
and rag/. None of it is distributed (see below). With a corpus in place:
.venv/bin/python rag/build_index.py # embedded Qdrant collection
.venv/bin/python rag/quant_rag.py "square-root law of market impact"
.venv/bin/python rag/quant_rag.py "cointegration test" --year-max 2000
.venv/bin/python rag/quant_rag.py --timeline "rough volatility"
claude mcp add quant-rag -- "$PWD/.venv/bin/python" "$PWD/rag/mcp_server.py"Scope
Not distributed: the corpus (third-party copyrighted documents) and everything derived from it: vectors, BM25 index, entity graph, consolidated metadata, benchmark questions and per-run results. Test fixtures that came from the corpus were replaced by structure-preserving scrambles; the three short quotes kept for the citation tests are bibliographic-length.
Research log kept private: about forty dated pre-registrations and reports, in French, quote the corpus. Code comments that point to
docs/…orRAPPORT-…files refer to it.Language: identifiers, comments and test names are mostly in French, the working language of the project; the documentation is in English.
LLMs: none on the served path. Language models only draft, launder and verify benchmark questions and judge answers (Mistral small and medium; Gemini once, as a separate instrument).
Hardware: developed and measured on an Apple M4 with 16 GB (MPS); CPU works, slower.
Stack: Python 3.12 · Qdrant · Qwen3-Embedding-0.6B · BM25 · bge / Qwen3 rerankers · GLiNER2 · MinerU · Model Context Protocol · NumPy.
License
All rights reserved. The code is published for reading and review; see LICENSE
and, for the provenance of src/, NOTICE.md.
Available Tools
10 toolsconnect_entitiesA
Passages et documents qui mentionnent LES DEUX entités (« SVI » et « rough volatility »).
Pour trouver où deux idées se rencontrent dans le corpus. Liste les documents par nombre de passages communs, puis les meilleurs passages avec chunk_id. Aucun passage commun est une information en soi : les deux entités ne sont jamais citées ensemble dans un même passage.
Args: entity_a: première entité (nom ou fragment, variantes réunies). entity_b: seconde entité. top_k: nombre de passages (défaut 10, au plus 2 par document).
| Name | Required | Description | Default |
|---|---|---|---|
| top_k | No | ||
| entity_a | Yes | ||
| entity_b | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and delivers: it discloses the ranking behavior (documents by number of common passages, then passages), the deduplication constraint (top_k but at most 2 per document), and the empty-result semantic ('Aucun passage commun est une information en soi'). It could add read-only/safety notes, but for an obvious search tool the disclosed behavior is substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: purpose in the first line, ranking/empty-result semantics in the second, and a clean Args block. The inline example entities ('SVI' and 'rough volatility') are slightly odd since an agent could momentarily read them as fixed values, but the Args block dispels that, and no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with an output schema present (so return format is covered elsewhere), the description is nearly complete: it covers purpose, use case, ranking order, the top_k-per-document cap, empty-result semantics, and all parameter meanings. Missing only error/not-found behavior, which is a minor gap for a search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description fully compensates by documenting all three parameters with meaning beyond the bare type: entity_a as 'nom ou fragment, variantes réunies' (name or fragment with variants grouped), entity_b by analogy, and top_k with its 'default 10, at most 2 per document' effective behavior. This resolves ambiguity the schema alone cannot.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line states a specific verb+resource: 'Passages et documents qui mentionnent LES DEUX entités' – a two-entity co-occurrence search, clearly distinct from siblings like search_documents (single-entity) and get_passage (single passage retrieval). It also names the search target precisely: documents ranked by common passage count, then the best passages with chunk_id.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit use case: 'Pour trouver où deux idées se rencontrent dans le corpus' (to find where two ideas meet in the corpus). This provides clear context for when to invoke the tool, though it does not name alternative tools or state when not to use it – no exclusions are given despite siblings like search_graph existing for similar relational queries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
corpus_statusA
Décrit l'index : documents, passages, couverture des métadonnées (années, auteurs), backend et device, et l'état du graphe d'entités (nombre d'entités, relations, état du corpus sur lequel il a été bâti).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and does a credible job: it explicitly frames the operation as a read/status report ('Décrit l'index') and enumerates exactly which dimensions it covers, including entity counts and corpus state. It omits minor traits like cost or permissions, but these are low-risk for a status tool with an output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the purpose ('Décrit l'index') before enumerating the report scope. It is efficient and wastes nothing, though the comma-separated run-on structure is slightly harder to parse than a bulleted or semi-colon separated list would be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and the presence of an output schema (which removes the need to describe return values), the description adequately covers the semantics of what this status tool reports. The main gap is the missing usage guidance relative to sibling tools, but the report scope itself is fully specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and an empty input schema, so the baseline of 4 applies. There are no parameter semantics for the description to clarify.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Décrit' — describes) and an enumerated resource list: documents, passages, metadata coverage (years, authors), backend, device, and entity graph state. This concrete scope clearly separates it from the search/verify siblings, though it never names or explicitly contrasts them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied — an agent can infer this is the go-to tool for overall index/corpus status — but there is no explicit guidance on when to choose it over list_documents or search_graph, no when-not conditions, and no mention of alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
expand_entityA
Entités reliées à entity_name dans le graphe, par relation, avec les passages qui l'attestent.
Répond à « quelles méthodes dépendent de la rough volatility ? », « à quoi SVI est-il comparé ? ».
Chaque voisin porte la relation et son sens (→ : l'entité est la tête, ← : la queue), le nombre
de passages qui soutiennent le lien et jusqu'à trois chunk_id à lire avec get_passage. Les
relations viennent de GLiNER2 (seuil 0,5) et sont clairsemées : un lien attesté par un seul
passage est une piste, pas un fait ; lire le passage avant de l'affirmer. Sans relation, la
réponse liste aussi les CO-MENTIONS : les entités citées dans les mêmes passages, pondérées
par leur rareté — la vue la plus utile pour explorer un sujet (auteurs, modèles, mesures
qui vont avec).
Args: entity_name: nom de l'entité (« rough volatility », « Heston model », « Gatheral »). L'entité exacte et ses variantes de nom (« rough volatility models ») sont réunies. relation: une seule relation (measures, predicts, causes, depends_on, correlates_with, applies_to, uses_method, compares_with, is_a, part_of) ; vide = toutes. entity_type: type de l'entité de départ, si le nom est ambigu (person, market_concept…). limit: nombre de voisins (défaut 20), classés par nombre de passages qui les attestent.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| relation | No | ||
| entity_name | Yes | ||
| entity_type | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that relations come from GLiNER2 at a 0.5 threshold, that they are sparse, and that a single passage link is 'une piste, pas un fait' — advising to read the passage before asserting. It also explains the direction indicators (→/←) and the co-mention weighting by rarity. It does not explicitly state that the operation is read-only, but that is strongly implied by its nature as a query tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured with a main paragraph covering purpose and behavior, followed by a clear 'Args' section. It is somewhat verbose but every sentence adds value: examples, relation semantics, co-mention explanation, and parameter details. The core purpose is front-loaded in the first sentence, making it easy for an agent to grasp quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters, an output schema exists (per context signals), and no annotations, the description covers all necessary aspects: input semantics, output structure (neighbors with relation, direction, passage count, chunk_ids), and behavior nuances (sparse relations, co-mentions). It does not detail error cases or edge conditions, but for a read-oriented graph exploration tool, this is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must fully explain each parameter. It does: entity_name includes examples and notes that variants are grouped; relation lists the allowed values and explains that empty returns all; entity_type is described as a disambiguation aid; limit is explained as neighbor count with default 20 and sorting by supporting passages. This completely compensates for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: returns entities related to a given entity in the graph, along with the relation and supporting passages. It provides concrete example questions ('quelles méthodes dépendent de la rough volatility ?') and explicitly describes the output structure. This is a specific verb-resource-object statement that distinguishes it from sibling tools like search_graph or connect_entities.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit context on when to use the tool via example questions and explains the difference between the default behavior (all relations) and the co-mention view when relation is omitted. It does not name alternative tools explicitly, but the examples and the note that co-mentions are 'la vue la plus utile pour explorer un sujet' clearly signal usage scenarios. It lacks an explicit 'when not to use' but is strong on context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_passageA
Retourne le passage complet d'un chunk_id, avec ses voisins immédiats, pour citation.
Args:
chunk_id: la clé interne rendue par search_documents ou search_graph.
max_characters: fenêtre servie. Vide = 15 000, la fenêtre du contrat de sortie. Le
défaut historique était 6 000 et tronquait 30,2 % des fenêtres du corpus
(médiane 4 520 c., max 13 982) : un outil dont le nom promet le passage entier en
rendait moins d'un tiers du temps.
neighbours: nombre de voisins de chaque côté (défaut 1, 0 pour le seul chunk demandé).
Les voisins ne sont pas filtrés comme les résultats de recherche — un voisin
en-tête est un intitulé de section, donc du contexte — mais ils sont déclarés dans
la ligne « voisins : ».
| Name | Required | Description | Default |
|---|---|---|---|
| chunk_id | Yes | ||
| neighbours | No | ||
| max_characters | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Aucune annotation n'étant fournie, la description porte seule la charge de la transparence. Elle révèle des comportements importants : la valeur null de max_characters donne 15 000 caractères, l'ancien défaut tronquait 30,2 % des fenêtres, et les voisins ne sont pas filtrés comme les résultats de recherche. Elle précise aussi que les voisins en-tête sont des intitulés de section et sont déclarés dans la ligne « voisins : ».
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
La description est dense et chaque phrase apporte une information utile. La phrase d'ouverture est immédiate, puis la liste des arguments structure clairement les détails. Même la note historique sur la troncature justifie sa place en expliquant pourquoi le défaut actuel est plus sûr.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Un schéma de sortie existe, donc la description n'a pas besoin de documenter la valeur de retour. Elle couvre les trois paramètres, leurs défauts, les cas limites (neighbours=0) et le comportement des voisins. Pour un outil de lecture avec cette complexité, rien d'essentiel ne manque.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Le schéma n'offre aucune description de paramètre (couverture 0 %), mais la description compense entièrement : chunk_id est expliqué comme une clé interne issue des outils de recherche, max_characters est défini avec sa valeur par défaut et son effet, et neighbours est décrit avec sa valeur par défaut et son comportement de filtrage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
La description commence par un verbe précis (« Retourne ») et une ressource claire (« le passage complet d'un chunk_id »), avec l'objectif « pour citation ». Elle se distingue des outils de recherche comme search_documents et search_graph en se concentrant sur l'extraction d'un passage à partir d'une clé interne.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
La description indique que chunk_id est « la clé interne rendue par search_documents ou search_graph », ce qui donne un contexte d'utilisation clair après une recherche. Elle mentionne l'usage pour citation, mais ne nomme pas explicitement d'alternatives ni de cas où ne pas utiliser cet outil.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_documentsA
Liste bibliographique des documents du corpus, du plus récent au plus ancien, sans recherche vectorielle.
Args: author: sous-chaîne du nom d'un auteur (insensible aux accents et à la casse). year_min: année de publication minimale (inclus). year_max: année de publication maximale (inclus). text: sous-chaîne du titre. limit: nombre maximal de lignes (défaut 50).
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | ||
| limit | No | ||
| author | No | ||
| year_max | No | ||
| year_min | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden and does substantial work: it states sort order, non-vector search behavior, substring matching for author and title, case/accent insensitivity for the author filter, and inclusive year bounds. It does not state side effects or authentication, but the operation is presented as a read-only listing and output shape is covered by the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence plus a compact args list. Every line adds a distinct piece of information; no filler or restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter, all-optional listing tool with an output schema, the description covers every invocation-relevant aspect: fields, filtering semantics, ordering, and limit behavior. Pagination is not needed beyond the limit parameter, and return values are structurally specified elsewhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so all semantic weight falls on the description. It fully compensates by explaining every parameter: author substring with accent/case handling, inclusive year_min/year_max, title substring for text, and limit default. This is exactly the compensation needed at zero schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line names a precise action and object: a bibliographic listing of corpus documents, sorted newest-to-oldest. The qualifier 'sans recherche vectorielle' also separates it from the sibling search_documents, so an agent can select it by exclusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys that this is a filtering/list operation, not a semantic one via 'sans recherche vectorielle,' which implies search_documents is the alternative. However, it never explicitly says 'use this when' or names the sibling, so the guidance remains implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_documentsA
Cherche des passages dans le corpus de recherche quantitative (corpus dont la taille n'a pas pu être lue).
Couvre la volatilité et les options, la microstructure et l'exécution, les facteurs et
l'alpha, la construction de portefeuille et le ML appliqué à la finance (1971–2026).
Retourne des extraits classés, précédés de la décision de routage. Chaque extrait porte
une ligne « citer : » — auteurs, année, titre, page — et c'est la seule à recopier
dans une réponse : l'année fait partie de l'information, le corpus mélangeant des travaux
de 1987 et des preprints de 2026. La ligne « interne : » porte chunk_id et
document_id, qui servent à get_passage et ne sont pas des citations — ils ne
survivent pas à un re-découpage du corpus et ne disent rien à un lecteur.
Les pages affichées sont celles d'un lecteur, et une plage quand le passage en couvre plusieurs : le système ne dispose pas d'un pointeur de phrase, et annoncer une page exacte pour un passage à cheval serait une précision qu'il n'a pas.
Protocole d'usage, mesuré sur 155 questions (benchmark/eval_protocol.py) :
Une période explicite (« sources published in 2022 or earlier », « papers before 2010 », « depuis 2020 ») est détectée dans la question et transformée en year_min/year_max, la clause retirée du texte : rien à faire, sauf quand la période est implicite (« récent », « d'avant la crise ») — alors la traduire soi-même en year_min/year_max (+0,10 nDCG@10).
Question ouverte, conceptuelle : une seule requête en langue naturelle, en anglais, qui décrit le problème ; la reformuler pour l'index n'a pas aidé (chantier C).
N'utiliser mode="hybrid" que pour une requête faite d'identifiants mémorisés (auteur, acronyme, numéro, année : « Fukasawa SVI B2 B3 ») ; sur toute autre question il perd (−0,11 nDCG@10 sur 130 questions, −0,19 sur les tableaux). Depuis le 5 septembre 2026 on sait pourquoi : ce chemin empile deux composants mesurés négatifs — la fusion (−0,088 pooled, 14 variantes essayées, zéro en GO) et le reranker bge-base sur pool dense (−0,076 pooled, mesuré sur ce corpus). Il reste utile pour le cas des identifiants exacts ; ailleurs, le choisir est une erreur documentée.
Question multi-documents ou exploratoire : compléter par search_graph / expand_entity / connect_entities sur les entités nommées de la question, puis lire les passages des deux sources ; ne pas fusionner aveuglément (mesuré : la fusion automatique perd).
Args: query: la question, en anglais de préférence (le corpus est anglophone). limit: nombre de passages à retourner (défaut 5). document_id: restreindre à un seul document. mode: "auto" (défaut) = recherche sémantique seule (dense) — depuis le banc de 150 questions, aucune règle automatique ne fait mieux. "hybrid" ajoute BM25+fusion+rerank : à demander explicitement quand la question est faite d'identifiants exacts dont l'utilisateur se souvient (Fukasawa, SVI, ITRAXX 2007, arXiv 1206.0682), et seulement dans ce cas — sur une question en langue naturelle, et sur les tableaux, l'hybride fait moins bien que le dense. rerank: laisser vide pour suivre le mode (activé en hybrid, désactivé en dense). True/False force. Mesuré : bge-reranker-base aide sur un pool hybride et nuit sur un pool dense (−0,076 nDCG@10 pooled, −0,147 sur les questions simples). Ne pas forcer rerank=True en mode dense. year_min: ne garder que les documents publiés à partir de cette année (inclus). year_max: ne garder que les documents publiés jusqu'à cette année (inclus). Une borne d'année exclut les documents dont l'année est inconnue (un nombre non lu de documents). Laissées vides, une période explicite dans la question est détectée et appliquée (la ligne « période: » de la réponse le dit) ; une borne donnée prime toujours. author: ne garder que les documents d'un auteur (sous-chaîne du nom, insensible aux accents et à la casse : "lopez de prado", "Gatheral", "Jacquier"). reclassement: "selectif" reclasse les dix premiers candidats par Qwen3-Reranker-0.6B en protégeant le rang 1 dense. Laisser vide (défaut) dans l'immense majorité des appels.
**Ce que ça coûte** : ~3,2 s par requête contre 59 ms sans — un facteur **54** —
et 2,4 Go de mémoire résidente tant que le modèle est chargé, sur une machine de
16 Go qui tourne déjà avec 2,4 Go de swap. Ce n'est pas un réglage anodin.
**Ce que ça rapporte** (155 questions, corpus 5530cba145) : l'or entre dans les
cinq passages servis pour 19 questions et en sort pour 5. Net +14, et les cinq
pertes sont un sous-ensemble strict des sept du reclassement plein.
**Ce que ces 19 et ces 5 valent en RÉPONSE — mesuré le 8 septembre 2026** sur les
24 questions concernées, jugées dans les deux bras : sur les 5 « pertes », **une
seule** perd vraiment sa réponse ; les trois autres mesurables avaient une
couverture nulle **dans les deux bras** — elles n'avaient rien à perdre. Sur les
19 « gains », **12** gagnent vraiment une réponse. Le rapport réel est **12 pour
1**, non 19 pour 5, et l'échange net vaut **+11 réponses**.
**Ce que ça change pour toi** : le risque de dégrader une réponse en l'activant est
**cinq fois plus petit** que ce que le paragraphe précédent laisse croire. Ce qui
reste vrai, et qui est désormais la seule raison de ne pas l'activer partout, c'est
le coût — 3,2 s et 2,4 Go sur une machine de 16 Go. Active-le sans hésiter quand la
qualité de la réponse compte plus que trois secondes ; ne l'active pas en rafale
sur une exploration où tu enchaînes dix recherches.
**Quand le mettre** : quand un humain demande explicitement une recherche plus
soignée, ou après avoir constaté qu'une première recherche a rendu des passages
hors sujet sur une question dont on a de bonnes raisons de croire que le corpus
porte la réponse.
**Quand ne pas le mettre — et c'est le point important** : ne l'active pas « au
cas où », ni systématiquement, ni selon une règle que tu te donnerais toi-même
(longueur de la question, présence de chiffres, famille supposée). Une règle
d'activation inventée par l'appelant est une **politique de routage non mesurée**,
et ce dépôt a mesuré six fois que ses politiques de routage perdent. Le gain
ci-dessus n'est valable que sur un usage à la demande ; il ne dit rien de ce que
vaudrait une heuristique automatique, qui n'a jamais été évaluée.
Un échec de reclassement ne casse jamais la recherche : l'ordre dense est rendu.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | auto | |
| limit | No | ||
| query | Yes | ||
| author | No | ||
| rerank | No | ||
| year_max | No | ||
| year_min | No | ||
| document_id | No | ||
| reclassement | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it discloses the routing-decision prefix, the citation line semantics, the page/range display limitation, the measured effects of hybrid and reranker paths, and the fallback behavior on reclassification failure. This goes well beyond what the schema alone could convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and uses bold section labels for the longer protocol notes. It is verbose, especially in the repeated cost/benefit discussion of reclassement, but the detail is warranted for a tool with subtle usage rules; only minor tightening would improve it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter search tool with no schema descriptions and no annotations, this description is complete: it covers invocation semantics, selection among siblings, expected output shape, caveats about citation vs. internal IDs, date-bound behavior, and measured costs. An agent has enough to call the tool correctly without inspecting external docs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the Args section compensates fully: every parameter (query, limit, document_id, mode, rerank, year_min/year_max, author, reclassement) is explained with defaults, semantics, and usage caveats. It even warns about forcing rerank=True in dense mode and the cost of reclassement.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a specific verb and resource — searching passages in the quantitative research corpus — and the topic/date coverage narrows it further. It also implicitly distinguishes itself from sibling tools like search_graph and get_passage by describing passage retrieval with citation output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: natural-language open questions use a single dense query; hybrid mode is reserved for memorized identifiers; rerank is recommended only when a human explicitly asks or after a failed first search; and graph tools are named as complements for multi-document or exploratory questions. It also includes measured losses for misuse, so an agent can make an informed choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_graphA
Cherche dans le graphe d'entités : les passages qui MENTIONNENT une entité dont le nom contient query.
Complémentaire de search_documents (sémantique) : ici la correspondance est nominale et exacte — « Gatheral » retourne les passages où Gatheral est cité, « SVI » ceux où SVI apparaît, même quand la question sémantique ne les ferait pas remonter. Les entités ont été extraites par GLiNER2 sur chaque passage (personnes, organisations, instruments, concepts de marché, mesures, méthodes, jeux de données). Même format de sortie que search_documents : source citable, pages, chunk_id (pour get_passage), plus les entités reconnues dans le passage. Croiser avec search_documents : le graphe ne connaît que les noms, pas le sens.
Args: query: nom d'entité ou fragment (« Gatheral », « rough volatility », « SVI », « S&P 500 »). Casse et accents ignorés ; « Lopez de Prado » trouve « López de Prado ». entity_type: restreindre à un type : person, organization, financial_instrument, market_concept, measure, method, dataset. relation: ne garder que les passages où l'entité est tête ou queue d'une relation de ce type : measures, predicts, causes, depends_on, correlates_with, applies_to, uses_method, compares_with, is_a, part_of. top_k: nombre de passages (défaut 10, au plus 2 par document).
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| top_k | No | ||
| relation | No | ||
| entity_type | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses key behaviors: exact nominal matching ('correspondance est nominale et exacte'), case and accent insensitivity ('Casse et accents ignorés'), the entity extraction method (GLiNER2), the output format similarity to search_documents, and the constraint 'au plus 2 par document' for top_k. This goes well beyond the structured schema and sets correct expectations for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than the minimum, but every sentence earns its place: it covers purpose, differentiation, output format, and parameter semantics. The structure is logical, starting with the core function, then the sibling contrast, then parameter details. It is slightly verbose but not bloated, earning a 4 rather than a 5 for prose efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists (as indicated by 'Has output schema: true'), the description need not detail return values. It still mentions the output format ('Même format de sortie que search_documents... plus les entités reconnues'). It covers all parameters, defaults, and usage context. Nothing essential is missing for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does thoroughly. Each parameter is explained: query with examples and matching behavior, entity_type with the full list of allowed values, relation with all possible relation types, and top_k with default and per-document limit. This gives the agent everything needed to construct correct arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Cherche dans le graphe d'entités : les passages qui MENTIONNENT une entité dont le nom contient query.' It specifies the action (search), the resource (entity graph), and the exact operation (find passages mentioning an entity name containing the query). It also explicitly contrasts with search_documents, making it easy for an agent to choose the right tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage guidance: 'Complémentaire de search_documents (sémantique) : ici la correspondance est nominale et exacte' and later 'Croiser avec search_documents : le graphe ne connaît que les noms, pas le sens.' This clearly tells the agent when to use this tool (exact name matching) versus when to use the semantic sibling, and even suggests cross-referencing both tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
timelineA
Chronologie d'un sujet : le meilleur passage de chaque document, groupé par année de publication.
Utile pour voir comment une idée a évolué (rough volatility 2014 → 2026) ou pour distinguer les travaux fondateurs des preprints récents.
Args: topic: le sujet, formulé comme une question ou un intitulé. year_min: borne basse (inclus). year_max: borne haute (inclus). limit: nombre de documents (défaut 20).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| topic | Yes | ||
| year_max | No | ||
| year_min | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does so well: it reveals the core behavior (one best passage per document, year grouping, inclusive year bounds, limit default) and the intended analytical use. It does not explain the meaning of null year bounds or the criterion for 'best passage', but these are minor given the read-only, query-like nature of the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well structured: a one-sentence definition, a short usage rationale, then a clean Args list. It is front-loaded with the most important behavior and contains no redundant sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description does not need to document return values. Combined with the Args list and usage rationale, an agent has enough to call the tool correctly. The remaining gaps are the behavior of null year bounds and the exact meaning of 'best passage', but these are not blocking for basic invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the Args section is the only source of parameter meaning. It adds helpful semantics for every parameter: topic should be phrased as a question/title, year_min/year_max are inclusive bounds, limit defaults to 20. The phrase 'nombre de documents' is slightly ambiguous (total vs per year), preventing a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states exactly what the tool does: it builds a chronology of a topic by taking the best passage from each document and grouping those passages by publication year. This is a specific, informative description that immediately distinguishes timeline from siblings like search_documents or list_documents. It is not a tautology of the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second paragraph explicitly identifies when timeline is useful: observing how an idea evolves or separating foundational works from recent preprints. It gives a concrete example (rough volatility 2014→2026). It does not explicitly name alternatives or state when not to use them, so it falls short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_citationA
Vérifie qu'une citation se trouve vraiment dans le document qu'elle nomme.
À utiliser avant d'affirmer qu'un document dit quelque chose, et surtout quand la citation vient d'ailleurs que d'un passage servi à l'instant : d'une note prise plus tôt, d'un résumé, d'une réponse antérieure. Une citation recopiée de travers, recollée à partir de deux endroits ou reformulée reste plausible à la lecture — c'est précisément ce que cet outil détecte, et rien d'autre ne le détecte.
La vérification est déterministe : aucune génération, aucun modèle de langue. La
citation est normalisée (ligatures, <sup>/<sub>, tirets, guillemets, blancs, casse) puis
cherchée dans le texte canonique du document, celui dont le sha256 est servi par la ligne
« ancre : ».
Ce que rend l'outil :
trouve— vrai seulement pour une correspondance exacte après normalisation. Une citation dont un seul mot diffère rendfalse. Mesuré : 50 extraits réels sur 50 retrouvés, 0 accepté sur 50 extraits altérés d'un mot.offsets— la position[début, fin]dans le texte canonique, opposable.plus_proche— quand c'est faux : le passage le plus ressemblant, sa ressemblance et le nombre de mots qui diffèrent. C'est ce qui dit où la citation a dérivé.source—texte canonique, outexte servipour les 5 914 tableaux dont le rendu Markdown n'est pas une sous-chaîne du document.
Args: document_id: la clé interne du document, servie par la ligne « interne : ». quote: le texte cité, tel qu'on s'apprête à l'affirmer. Une phrase ou deux suffisent ; au-delà d'un paragraphe, une reformulation invisible fera échouer la vérification pour une raison sans rapport avec l'honnêteté de la citation. chunk_id: facultatif, restreint la recherche du repli « texte servi » à ce passage.
| Name | Required | Description | Default |
|---|---|---|---|
| quote | Yes | ||
| chunk_id | No | ||
| document_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility and meets it thoroughly. It discloses that verification is deterministic (no generation), details the normalization process, specifies exact-match behavior after normalization, provides measured accuracy (50/50 and 0/50), and explains the fallback for tables ('texte servi'). It also warns about quote length limits. This is exceptionally transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every paragraph earns its place: the main purpose, the deterministic behavior, the return fields, and parameter explanations. It is front-loaded with the core action and structured with bullet points for return values. It could be trimmed slightly (e.g., the measurement anecdote), but overall it's appropriately sized for the complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the description is complete: it explains the return fields (trouve, offsets, plus_proche, source), the source lines referenced, the fallback behavior, and parameter nuances. It also addresses edge cases like long quotes and the 5,914 tables. An agent has everything needed to invoke this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does. Each parameter is explained: document_id as 'la clé interne... servie par la ligne « interne : »', quote with guidance on length and phrasing, and chunk_id as optional with a clear purpose ('restreint la recherche du repli « texte servi »'). This adds meaning far beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Vérifie') and resource ('qu'une citation se trouve dans le document qu'elle nomme'), making the tool's function unmistakable. It also contrasts with the sibling tool 'verify_citations' (plural) by emphasizing it checks exact presence, not just similarity. This is a clear, distinguishing statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells when to use the tool: 'À utiliser avant d'affirmer qu'un document dit quelque chose' and especially when the citation comes from somewhere other than the current passage. It also claims 'rien d'autre ne le détecte', effectively routing the agent to this tool. However, it does not explicitly mention alternatives or when NOT to use it, only these strong directives. This is clear context without a full exclusion statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_citationsA
Vérifie plusieurs citations d'un coup — même contrat que verify_citation.
Args:
citations: liste d'objets {"document_id": …, "quote": …, "chunk_id": … (facultatif)}.
Vérifier ensemble les citations d'une même réponse coûte à peine plus qu'une seule :
le texte canonique d'un document n'est lu qu'une fois.
| Name | Required | Description | Default |
|---|---|---|---|
| citations | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the disclosure burden. It reveals a useful efficiency behavior (canonical text read only once) and the batching semantics, but does not state whether the operation is read-only, what it returns, or how failures are handled; 'same contract' borrows those details from a sibling rather than stating them directly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, starts with the core purpose, and contains no filler. Every sentence earns its place: purpose, sibling contract, parameter structure, and a non-obvious performance insight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter batch tool with an output schema, the description covers purpose, when to use it, and parameter format. The main gap is explicit routing guidance between single and batch verification, but the 'same contract as verify_citation' reference keeps the definition sufficiently usable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only says `citations` is an array of open objects (0% schema coverage), while the description supplies the object shape: `document_id`, `quote`, and optional `chunk_id`. That is enough structure to build a valid call, though the meaning of the fields is left largely to naming conventions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific action and object ('Vérifie plusieurs citations d'un coup') and clearly positions it as the batched counterpart to `verify_citation`. It is not a tautology and tells an agent exactly what resource is operated on and at what granularity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete usage context: batch citations from the same response, with a performance rationale ('coûte à peine plus qu'une seule'). It does not explicitly say 'use verify_citation for a single citation', but the 'même contrat' reference and the plural/singular contrast imply the alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v1.4.0- First observed
connect_entities - First observed
corpus_status - First observed
expand_entity - First observed
get_passage - First observed
list_documents - First observed
search_documents - First observed
search_graph - First observed
timeline - First observed
verify_citation - First observed
verify_citations
TDQS
Scored across 10 tools
Most tools have clearly distinct purposes: search_documents (semantic search), search_graph (entity mention search), expand_entity (entity neighbors), connect_entities (co-mentions), get_passage (retrieve chunk), verify_citation/verify_citations (verification), list_documents (bibliographic listing), timeline (chronological overview), corpus_status (index status). The only potential confusion is between search_documents and search_graph, but their descriptions explicitly contrast semantic vs. nominal matching. verify_citation and verify_citations are clearly the same operation with batch vs. single, which is acceptable.
Tool names mostly follow a verb_noun pattern: search_documents, search_graph, get_passage, list_documents, expand_entity, connect_entities, verify_citation, verify_citations, corpus_status, timeline. The pattern is consistent (imperative verb + object), though corpus_status and timeline are noun-only rather than verb_noun, which is a minor deviation. verify_citation vs verify_citations is a clear plural variant.
10 tools is well within the ideal 3-15 range for a RAG/quant-research server. Each tool serves a distinct function: corpus status, search (semantic, graph, timeline), retrieval (get_passage), citation verification (single and batch), document listing, and entity graph exploration (expand, connect). No tool feels redundant or superfluous.
The tool surface covers the full RAG workflow: search (semantic, entity-based, timeline), passage retrieval, citation verification, document listing, corpus status, and entity graph exploration. The domain is a read-only research corpus, so CRUD operations are not expected. The only minor gap might be a tool to get full document text, but get_passage with neighbors and list_documents cover most needs, and the corpus is described as passage-based.
Maintenance
Related MCP Connectors
Finance research agent for thesis, trade note, and market-state review.
Real-time financial news & regulatory intelligence: cited search, fetch, summaries, collections.
Institutional financial data with SEC filing citations, for every AI agent. OAuth 2.1.
ResearchOracle - 11 financial research tools: 10-K parsing, equity, macro, citation graph.
Related MCP Servers
- FlicenseBqualityNot gradedmaintenanceEnables LLMs to retrieve, analyze, and visualize stock prices and financial report data for quantitative trading research and investment analysis. Provides real-time and historical stock data, financial statement analysis, key metric calculations, and trading signal visualization.13-
- AlicenseNot gradedqualityDmaintenanceEnables natural language portfolio monitoring, risk analysis, and compliance checking with tools for real-time portfolio status, risk metrics, anomaly detection, SEC filings search, market KPIs, and compliance validation.MIT
- FlicenseAqualityBmaintenanceProvides an institutional research backend for AI assistants, with 15 tools for company, financial, funding, competitor, industry, and news intelligence, plus Markdown/PDF report generation, featuring deterministic source routing, extraction, validation, and citation generation.15-
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to perform institutional-grade research by retrieving and validating company, financial, industry, news, litigation, and funding data, then generating cited reports in Markdown, HTML, or PDF.-