Skip to main content
Glama

KFabric

CI

KFabric est une plateforme Python-first de fabrication de corpus documentaires. Le projet vise un problème très concret : aider à construire un corpus traçable, pondéré et réutilisable à partir de sources hétérogènes, avant même de brancher un assistant RAG conversationnel.

Au lieu de passer directement du web au chat, KFabric se concentre d'abord sur la qualité du matériau documentaire :

  • découverte de documents candidats

  • collecte et normalisation

  • scoring et décision documentaire

  • récupération de fragments utiles dans des documents rejetés

  • consolidation et synthèse

  • préparation d'artefacts indexables pour des usages RAG futurs

Pourquoi KFabric

Dans beaucoup de pipelines RAG, la vraie faiblesse n'est pas le modèle mais le corpus. KFabric part de l'idée inverse :

  • un bon corpus vaut mieux qu'une mauvaise conversation bien emballée

  • les documents faibles contiennent parfois des signaux utiles à sauver

  • la traçabilité et la prudence documentaire doivent exister dès le MVP

  • un serveur MCP et une API REST doivent exposer exactement le même coeur métier

Related MCP server: knowledge-index

Ce que fait le MVP

Le MVP actuel couvre déjà un flux bout en bout :

  1. créer une requête documentaire

  2. découvrir des documents candidats

  3. collecter et parser un document

  4. attribuer un score global et des sous-scores

  5. accepter, rejeter, ou rejeter avec récupération partielle

  6. consolider les fragments sauvés

  7. générer une synthèse documentaire prudente

  8. construire un corpus final

  9. préparer un artefact d'indexation

Points forts

  • API REST FastAPI pour piloter le pipeline corpus

  • serveur MCP natif en Python

  • workers Celery pour les traitements longs

  • UI légère en Jinja2, HTMX et Alpine.js

  • modèles SQLAlchemy 2 et migration Alembic initiale

  • mode sécurisé activé par défaut

  • approche corpus-first avant chat RAG complet

Architecture

Le projet est structuré comme un monolithe modulaire Python :

Stack technique

  • Python 3.12

  • FastAPI

  • Pydantic v2

  • SQLAlchemy 2 + Alembic

  • Celery + Redis + RabbitMQ

  • PostgreSQL prêt pour la production

  • MCP Python SDK

  • Jinja2 + HTMX + Alpine.js

Démarrage rapide

Installation minimale :

python3.12 -m venv .venv
source .venv/bin/activate
pip install setuptools wheel
pip install -e ".[dev]" --no-build-isolation
cp .env.example .env
uvicorn kfabric.api.app:app --reload

Si tu veux aussi les dépendances plus lourdes liées aux connecteurs et à la préparation RAG étendue :

pip install -e ".[dev,extended]" --no-build-isolation

L'application démarre ensuite sur :

  • UI : http://127.0.0.1:8000/

  • API : http://127.0.0.1:8000/docs

Commandes utiles via Makefile :

make install-extended
make test
make run-api
make stack-up

Async broker-only

Les traitements asynchrones de KFabric fonctionnent maintenant en mode broker-only :

  • les routes et boutons async doivent être dispatchés via Celery

  • RabbitMQ et Redis doivent être disponibles

  • le worker KFabric doit être lancé

  • il n'existe plus de fallback local en thread si le broker ne répond pas

Variables recommandées dans .env.example :

KFABRIC_PREFER_CELERY_TASKS=true
KFABRIC_CELERY_ALWAYS_EAGER=false

Dans docker-compose.yml, les services KFabric utilisent toujours les hôtes internes postgres, redis, rabbitmq et qdrant, même si ton fichier .env contient des URLs localhost pour un lancement hors Docker.

En environnement de test, KFABRIC_CELERY_ALWAYS_EAGER=true reste utile pour exécuter les tâches immédiatement sans broker externe.

Exploitation V1

KFabric dispose maintenant d'un mode d'exploitation local plus stable :

  • docker-compose.yml avec migrations, healthchecks et volume de stockage

  • Makefile pour les commandes courantes

  • readiness détaillée sur /api/v1/readiness

  • mode async broker-only avec worker Celery dédié

Le runbook dédié est disponible dans docs/v1-runbook.md.

Sécurité et accès

KFabric peut fonctionner sans authentification en local, mais dès qu'une clé API est configurée via KFABRIC_API_KEY, l'accès est protégé :

  • l'API REST accepte X-API-Key ou Authorization: Bearer ...

  • l'interface web demande une session via /auth

  • les réponses exposent un trace_id et des headers de sécurité

Exemple :

export KFABRIC_API_KEY="change-me"
curl -H "Authorization: Bearer change-me" http://127.0.0.1:8000/api/v1/version

Démo produit

Deux scénarios de démonstration reproductibles sont fournis dans docs/demo-scenarios.md.

Génération rapide :

export KFABRIC_DATABASE_URL="sqlite:////tmp/kfabric-demo.db"
export KFABRIC_STORAGE_PATH="/tmp/kfabric-demo-storage"
./.venv/bin/python scripts/generate_demo_scenarios.py \
  --base-url "http://127.0.0.1:8010" \
  --output /tmp/kfabric-demo-manifest.json

Les captures de démonstration peuvent ensuite être générées depuis l’UI locale, et les corpus sont exportables en HTML et en Markdown.

Aperçu visuel

Page d'accueil avec les requêtes récentes :

Accueil KFabric

Workflow corpus-first sur le scénario "savon Europe" :

Tableau de bord KFabric - savon Europe

Export HTML prêt pour une démo ou une revue documentaire :

Export corpus KFabric - savon Europe

La galerie complète des scénarios validés est disponible dans docs/demo-scenarios.md.

Vérification locale

Les tests principaux peuvent être lancés avec :

pytest

Le dépôt exécute aussi une CI GitHub Actions sur push et pull_request via ci.yml.

Le MVP a été vérifié localement avec :

  • compilation Python

  • tests API

  • tests REST/MCP

  • tests de scoring et de récupération de fragments

API et MCP

KFabric expose deux surfaces complémentaires :

  • une API REST métier pour piloter tout le pipeline

  • une API REST concordante MCP

  • un serveur MCP natif en Python pour les tools, resources et prompts

Exemples de capacités exposées :

  • create_query

  • discover_documents

  • list_candidates

  • collect_candidate

  • analyze_document

  • accept_document

  • reject_document

  • consolidate_fragments

  • generate_fragment_synthesis

  • build_corpus

  • prepare_index

  • get_corpus_status

Statut du projet

KFabric est aujourd'hui un MVP technique fonctionnel.

Ce qui existe :

  • coeur métier corpus-first

  • contrat REST principal

  • socle MCP

  • UI workflow

  • migration initiale

  • tests de base

Ce qui viendra ensuite :

  • vrais connecteurs documentaires

  • meilleure collecte multi-format

  • embeddings réels et intégration vectorielle étendue

  • scoring plus fin par domaine

  • multi-tenant

  • interface de production plus avancée

Développement avec assistance IA

Ce projet a été conçu et développé avec assistance IA, puis structuré, contrôlé, vérifié et arbitré manuellement.

L'IA a servi d'accélérateur pour :

  • le prototypage

  • l'implémentation initiale

  • la documentation

  • certaines itérations techniques

Les choix d'architecture, la validation du flux MVP et la cohérence produit ont été assumés et consolidés manuellement.

Licence

Licence non encore définie.

Available Tools

13 tools
accept_documentC

Force a manual accept decision for a parsed document.

ParametersJSON Schema
NameRequiredDescriptionDefault
kwargsYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description should disclose behavioral traits, but 'Force a manual accept decision' does not indicate if the action is destructive, reversible, or has side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (one sentence) but omits critical information, making it under-specified rather than efficiently complete.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations, output schema, and parameter documentation, the description fails to provide a complete understanding of the tool's behavior and requirements.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter 'kwargs' is completely undocumented in both the input schema (0% coverage) and the description, providing no semantic meaning to the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Force a manual accept decision') and the resource ('a parsed document'), making the purpose specific and distinguishable from sibling tools like 'reject_document'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, such as whether it should be used after analysis or when to accept vs. reject a document.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_documentC

Parse, score and decide on a collected document.

ParametersJSON Schema
NameRequiredDescriptionDefault
kwargsYes

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided. The description mentions parsing, scoring, and deciding but doesn't disclose side effects, whether it mutates state, or what 'decide' entails (e.g., does it trigger acceptance/rejection?).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise, but it lacks necessary detail. Conciseness is good, but under-specification harms usability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the high number of sibling tools and no output schema, the description fails to explain how this tool fits in the workflow, what it returns, or what 'decide' means in relation to accept/reject.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter 'kwargs' has no schema details, type, or description. Schema description coverage is 0%. The description does not clarify what kwargs expects, making it impossible to invoke correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it parses, scores, and decides on a document, but 'decide' is vague. Among siblings like accept_document and reject_document, it's unclear how this differs—it may be the analysis step, but that's not explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. Siblings include accept/reject documents, but no context on workflow order or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

build_corpusC

Build the final corpus for a query.

ParametersJSON Schema
NameRequiredDescriptionDefault
kwargsYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must convey behavioral traits, but it only states a high-level purpose. It does not disclose side effects, return values, permissions, or dependencies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise, but it omits essential details, making it under-specified rather than efficiently informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and a vague parameter, the description is completely inadequate. It does not explain how to use the tool or what 'kwargs' should contain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one parameter 'kwargs' with no type or description, and the description adds no parameter information. Schema description coverage is 0%, so the description fails to compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Build the final corpus for a query' states a verb and resource, but 'build' is vague and 'corpus' is not clearly defined. It does not distinguish among siblings like collect_candidate or consolidate_fragments.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. No explicit context, prerequisites, or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

collect_candidateC

Collect a candidate document and persist the raw payload.

ParametersJSON Schema
NameRequiredDescriptionDefault
kwargsYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry full burden. It states 'persist' implying a write operation, but does not disclose side effects, authorization requirements, idempotency, or behavior on duplicate candidates. This is insufficient for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise at one sentence, but it lacks substantive information. While brevity is positive, it comes at the cost of clarity and completeness, making it merely adequate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 13 sibling tools, no output schema, and a parameter with 0% coverage, the description is severely incomplete. It does not explain the tool's role within the corpus management workflow, nor does it specify inputs, outputs, or prerequisites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has a single parameter 'kwargs' with no type or description, and schema description coverage is 0%. The description adds no meaning about what kwargs should contain, failing to compensate for the lack of schema detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies a verb ('collect') and resource ('candidate document') and mentions persistence of raw payload, which distinguishes it from sibling tools like accept_document or reject_document. However, the term 'collect' is ambiguous and could imply fetching or gathering, requiring clarification.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as accept_document, discover_documents, or list_candidates. The description lacks context for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

consolidate_fragmentsC

Cluster salvaged fragments for a query.

ParametersJSON Schema
NameRequiredDescriptionDefault
kwargsYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description should fully disclose behavior. It only says 'cluster', which implies grouping but does not explain what happens to the fragments, whether it mutates state, or what the output format is. The mutation/destructive nature is unclear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at six words, which is efficient but sacrifices necessary detail. It is front-loaded with the verb, but does not earn its place as it omits critical information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one opaque parameter, no output schema, no annotations), the description is insufficiently complete. It does not clarify how to use the kwargs, what clustering means, or what the result provides, leaving significant gaps for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole parameter 'kwargs' is an opaque object with no description in the schema or the tool description. With 0% schema description coverage, the description fails to add any meaning to this parameter, leaving the agent with no guidance on what keys or values to provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'cluster' and identifies the resource as 'salvaged fragments for a query', making the general purpose clear. It is distinguishable from related tools like list_salvaged_fragments and generate_fragment_synthesis, but lacks detail on the clustering approach.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives (e.g., when to cluster vs list or synthesize). There are no usage contexts, prerequisites, or exclusions stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_queryC

Create a new KFabric documentary query.

ParametersJSON Schema
NameRequiredDescriptionDefault
kwargsYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description must disclose behavioral traits. It only says 'Create,' implying a state change, but gives no details about side effects, return values, authentication needs, or whether the query is saved or executed. The presence of a generic 'kwargs' parameter further obscures behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise, but it lacks structured content like parameter details or examples. It is front-loaded with the action but omits necessary information, reducing its effectiveness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of output schema, annotations, and parameter descriptions, the description is extremely incomplete. It does not explain the query's purpose, return value, or relation to sibling tools, making it insufficient for reliable use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has a required 'kwargs' parameter with no type or description, and the description provides zero clarification. With 0% schema description coverage, the description should compensate but does not, making invocation impossible without prior knowledge.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Create a new KFabric documentary query,' clearly identifying the action (create) and resource (query) within the KFabric domain. This distinguishes it from sibling tools like accept_document or analyze_document. However, the term 'documentary query' is not explained, leaving some ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. Sibling tools include many actions (e.g., analyze_document, build_corpus), but there is no context, prerequisites, or exclusions, leaving the agent without direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

discover_documentsC

Launch discovery for an existing KFabric query.

ParametersJSON Schema
NameRequiredDescriptionDefault
kwargsYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully convey behavioral traits. It does not indicate whether the tool is read-only, destructive, or what side effects occur (e.g., creating resources, modifying state). The phrase 'launch discovery' implies an asynchronous operation, but no details are given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short (one sentence), but it is under-specified rather than concise. It fails to convey necessary information about the tool's purpose, parameters, and behavior, making it insufficient for reliable selection and invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations, output schema, and parameter descriptions, the description is woefully incomplete. It does not address the complexity of the tool (single generic parameter, no return value info), leaving critical gaps for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter 'kwargs' is generic and lacks description in both the schema and the tool description. With 0% schema description coverage, the description adds no meaning to the parameter, leaving the agent completely uninformed about required or optional input structure.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the action ('launch discovery') and the resource ('existing KFabric query'), but 'discovery' is vague and not elaborated, leaving ambiguity about what the tool accomplishes. It distinguishes from siblings like 'create_query' by specifying an existing query, but lacks clarity on the meaning of 'discovery'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'analyze_document' or 'create_query'. The description does not mention prerequisites, contexts, or exclusions, leaving the agent to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_fragment_synthesisC

Create a synthesis from salvaged fragments.

ParametersJSON Schema
NameRequiredDescriptionDefault
kwargsYes

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, and the description only says 'create,' implying mutation, but lacks details on side effects, permissions, or other behavioral traits. The description carries the full burden but provides minimal information.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short (one sentence), which might be concise but is under-specified. It lacks any structural elements or additional context, making it insufficient for a tool with a vague parameter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of an output schema, no annotation coverage, and a single undocumented parameter, the description fails to provide enough context for an AI to use the tool effectively. It lacks details on input, output, and behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'kwargs' is undefined with no type or description in the schema (0% coverage). The description does not explain what 'kwargs' should contain, leaving the AI with no guidance on how to construct valid input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'Create a synthesis from salvaged fragments.' It indicates the tool's function but does not differentiate from sibling tools like 'consolidate_fragments' or 'build_corpus'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. There is no mention of conditions, prerequisites, or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_corpus_statusC

Read the latest corpus status for a query or corpus id.

ParametersJSON Schema
NameRequiredDescriptionDefault
kwargsYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description carries full burden. It implies a read-only operation but does not disclose side effects, permissions, or return behavior. The description is minimal and lacks transparency beyond the basic operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, making it concise. However, given the vague schema, it is too brief and sacrifices necessary detail. It could be longer to cover parameter usage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and an opaque parameter, the description fails to provide complete context. The agent lacks information about the status object structure, when to use the tool, and how to specify the query or corpus id.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one parameter 'kwargs' with 0% description coverage. The description adds that it is for a 'query or corpus id', but does not specify how to structure the kwargs, leaving the agent to guess the format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Read' and the resource 'corpus status', specifying it is for a query or corpus id. However, it does not differentiate from sibling tools like 'list_candidates' or 'build_corpus', which might also be read-related.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description lacks context about prerequisites or situations where getting the status is appropriate, leaving the agent without direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_candidatesC

List candidate documents for a query.

ParametersJSON Schema
NameRequiredDescriptionDefault
kwargsYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so description carries full burden. It implies read-only but does not disclose any behavioral traits like pagination, sorting, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded, but it is underspecified. It lacks necessary detail, so brevity is not a virtue here.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and a vague parameter, the description is incomplete. The agent needs information on input format and expected output for a list operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the only parameter 'kwargs' is undocumented. The description does not explain what keys or values are expected, leaving the agent entirely in the dark.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists candidate documents for a query, which is a specific verb and resource. However, it does not differentiate from sibling tools like 'discover_documents' or 'collect_candidate'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance provided on when to use this tool instead of alternatives, nor any context about prerequisites or scenarios where it should not be used.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_salvaged_fragmentsC

List salvaged fragments optionally filtered by query.

ParametersJSON Schema
NameRequiredDescriptionDefault
kwargsYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description must disclose behavioral traits. It only states a list operation with optional filtering, but omits critical details such as pagination, ordering, permission requirements, or what happens when no fragments match.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise (8 words) and front-loaded with the main action. However, it is too sparse to be fully effective; it sacrifices completeness for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the vague parameter, lack of output schema, and no annotations, the description is severely incomplete. An agent cannot determine how to invoke the tool correctly or what to expect from it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one parameter 'kwargs' with no type or description, and schema description coverage is 0%. The description mentions 'optionally filtered by query' but does not explain how to represent the query within kwargs, leaving the agent without usable parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'List' and resource 'salvaged fragments', indicating the action clearly. It also mentions optional filtering by query. It distinguishes from sibling tools like accept_document or analyze_document, which have different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, nor any prerequisites or conditions for filtering. It lacks context about its place in the workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prepare_indexC

Prepare an indexable artifact for a corpus.

ParametersJSON Schema
NameRequiredDescriptionDefault
kwargsYes

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It only says 'prepare', which implies a non-destructive operation, but no details on side effects, permissions, or output behavior are given. This is insufficient for an agent to understand the tool's impact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short (five words), but this brevity comes at the cost of necessary detail. It is not appropriately sized for the information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the toolset (13 sibling tools), the lack of output schema, and the completely undocumented parameter, the description is woefully incomplete. It does not provide enough information for an agent to use this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one parameter 'kwargs' with no type constraints or descriptions, and schema description coverage is 0%. The description adds no information about what this parameter expects, so the agent cannot infer how to use it correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the verb 'prepare' and the resource 'indexable artifact for a corpus', making the basic action clear. However, it does not differentiate from sibling tools like build_corpus or generate_fragment_synthesis, and the term 'indexable artifact' is vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no prerequisites, and no context about the expected workflow. The agent is left without any decision-making support.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reject_documentC

Force a manual reject decision for a parsed document.

ParametersJSON Schema
NameRequiredDescriptionDefault
kwargsYes

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It reveals that the tool forces a manual reject decision, implying mutation, but lacks details on side effects, reversibility, permissions, or what happens to the document post-reject.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short (one sentence) but under-specified for the complexity of a tool with a single opaque parameter. It sacrifices necessary detail for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations, output schema, and parameter documentation, the description is woefully incomplete. It fails to cover return values, parameter requirements, behavioral context, or usage scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole parameter 'kwargs' has no schema description, type, or constraints (0% coverage). The description does not explain what kwargs should contain, leaving the agent without meaningful guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('reject') and the resource ('parsed document'), making the purpose understandable. However, it does not differentiate from the sibling 'accept_document' beyond being the opposite, preventing a top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus its siblings (e.g., accept_document) or when not to use it. No context about prerequisites or alternatives is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updates
    • Addedbuild_corpus
    • Addedcollect_candidate
    • Addedconsolidate_fragments
    • Addedcreate_query
    • Addedprepare_index
  2. 8 tool updatesv0.1.0
    • First observedaccept_document
    • First observedanalyze_document
    • First observeddiscover_documents
    • First observedgenerate_fragment_synthesis
    • First observedget_corpus_status
    • First observedlist_candidates
    • First observedlist_salvaged_fragments
    • First observedreject_document

TDQS

B3/5.0

Scored across 13 tools

Disambiguation5/5

Each tool targets a distinct step in the document processing pipeline, from query creation to corpus building. There is minimal overlap between tools like analyze_document and accept_document, which serve different stages.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (e.g., create_query, discover_documents). The style is uniform with no mixing of conventions.

Tool Count5/5

The 13 tools are well-scoped for a document collection and analysis pipeline. Each tool earns its place, covering discovery, collection, analysis, decision, and corpus assembly.

Completeness4/5

The core workflow is fully covered, but missing tools for listing queries or deleting candidates/corpora are minor gaps. Overall, the surface is nearly complete for the stated purpose.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    A private, self-hosted RAG service over MCP that enables document ingestion, hybrid retrieval (BM25 + dense vectors fused with RRF), and notebook management through 14 tools, keeping documents on your own hardware.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    A modular RAG framework exposing knowledge retrieval tools via MCP, enabling AI assistants to perform hybrid search, reranking, and multimodal document queries with full observability and evaluation.
    MIT