rechtssysteem-mcp
rechtssysteem-mcp
Een MCP-server die AI-agents toegang geeft tot een eerlijke voorspelling van Nederlandse rechtszaken: het model heeft de uitkomst nooit gezien.
Waarom "eerlijk"
De meeste uitkomstvoorspellers scoren hoog omdat de uitspraak zelf het antwoord al bevat. Meet je dat, dan blijkt 92% van de teksten de uitkomst letterlijk te verraden. Wij knippen die zinnen er eerst uit (knipregel R2) en meten wat er overblijft: 0,1% restlekkage.
Wat het model daarna nog kan, is dus echt geleerd:
Zaken | 609.715 |
Accuracy | 78,2% |
Macro-F1 | 77,1% |
Meerderheidsbaseline | 43,7% |
Restlekkage na knip | 0,1% (was 92%) |
Validatie | 5-fold CV, out-of-fold |
Per klasse F1: afgewezen 0,827 · gedeeltelijk 0,726 · toegewezen 0,761.
De baseline staat er met opzet bij. 78,2% zegt niets zonder die 43,7% ernaast.
Related MCP server: Vaquill-AI/vaquill-mcp
Tools
Tool | Wat het doet |
| afgewezen / gedeeltelijk / toegewezen, met kansen per klasse |
| de benchmarkcijfers hierboven |
| meet of een tekst de uitkomst al verraadt — bruikbaar om andermans dataset of AI-claim te toetsen |
Installeren
Er is niets te installeren. De client gebruikt uitsluitend de Python-standaard- bibliotheek — geen pip, geen wheels, geen supply chain. Python 3.10 of hoger.
curl -O https://raw.githubusercontent.com/rechtssysteem-ai/rechtssysteem-mcp/main/rechtssysteem_mcp.pySleutel
lekkage_check en rechtspraak_cijfers werken zonder sleutel. De meetlat
hoort controleerbaar te zijn: wie wil nagaan of een dataset of een AI-claim de
uitkomst al verraadt, moet dat kunnen zonder eerst iets te vragen.
voorspel_uitkomst heeft wel een sleutel nodig. Gebruik deze:
242b9a68ddeaf44a53de32590fb6379e13e608397e6a9ed1c82af238ab871ca7Gedeelde proeftier-sleutel, max 20 verzoeken per minuut, kan wijzigen; betaalde sleutels volgen. Hij wordt door iedereen gedeeld, dus de limiet geldt voor het geheel. Bij misbruik draaien we hem en komt de nieuwe hier te staan.
Claude Code
claude mcp add rechtssysteem --scope user \
--env RECHTSSYSTEEM_API_KEY=242b9a68ddeaf44a53de32590fb6379e13e608397e6a9ed1c82af238ab871ca7 \
-- python3 /pad/naar/rechtssysteem_mcp.pyClaude Desktop — claude_desktop_config.json
{
"mcpServers": {
"rechtssysteem": {
"command": "python3",
"args": ["/pad/naar/rechtssysteem_mcp.py"],
"env": { "RECHTSSYSTEEM_API_KEY": "242b9a68ddeaf44a53de32590fb6379e13e608397e6a9ed1c82af238ab871ca7" }
}
}
}Instellingen
Variabele | Standaard | |
| — | alleen voor |
|
| |
|
| seconden |
Privacy
De zaaktekst gaat naar de server van Rechtssysteem.ai om geanalyseerd te worden. Stuur geen tekst die u niet mag delen. Maximaal 20.000 tekens per verzoek.
Wat dit niet is
Een risico-indicatie op grond van vergelijkbare rechtspraak. Geen juridisch advies. Bij een zekerheid onder 55% zegt de tool "weet niet", en dat is dan ook het enige juiste antwoord.
Juridische status & aansprakelijkheid
Geen juridisch advies. Deze MCP-server is uitsluitend bedoeld voor statistische analyse, benchmarking en onderzoek van openbare rechterlijke uitspraken. De output vormt uitdrukkelijk geen juridisch advies, geen bindende proceskansenbeoordeling en geen vervanging van een advocaat.
Beperking. Rechterlijke beslissingen hangen af van individuele feiten, procesvoering, bewijswaardering en de discretionaire bevoegdheid van de rechter. Statistiek over het verleden garandeert niets over de toekomst.
Onzekerheidsdrempel. Bij een zekerheid onder 55% geeft het systeem "onbepaald — onvoldoende statistische significantie". De drempel is een gekozen instelling en kan wijzigen.
Aansprakelijkheid. Rechtssysteem.ai is niet aansprakelijk voor schade die voortvloeit uit gebruik van of vertrouwen op deze software, behoudens opzet of bewuste roekeloosheid. Tegenover zakelijke afnemers is de aansprakelijkheid in elk geval beperkt tot het bedrag dat in de twaalf maanden vóór de schadeveroorzakende gebeurtenis voor de dienst is betaald. Tegenover consumenten geldt deze beperking slechts voor zover zij niet onredelijk bezwarend is; dwingend consumentenrecht blijft onverkort gelden.
AI-transparantie. Alle output wordt volledig gegenereerd door een machine-learning model (LightGBM); er komt geen menselijke beoordeling aan te pas. Rechtssysteem.ai vermeldt dit uit eigen beweging. Het systeem genereert geen synthetische inhoud in de zin van artikel 50, lid 2, van de AI-verordening.
Niet bestemd voor rechterlijke instanties. Deze software is niet bedoeld voor gebruik door of namens een rechterlijke instantie bij het onderzoeken en uitleggen van feiten en recht, noch voor autonome geschilbeslechting zonder menselijke tussenkomst.
Merk, intellectueel eigendom & licentie
Copyright 2026 Rechtssysteem.ai (The Coppola Connection).
Deze client is gelicentieerd onder de Apache License, Version 2.0. Zie LICENSE en NOTICE.
Het model, de trainingsdata en de lekkage-knipregel (R2) zijn niet onder deze licentie vrijgegeven en blijven eigendom van Rechtssysteem.ai.
"Rechtssysteem.ai" en "rechtssysteem-mcp" zijn handelsnamen. Conform Section 6 van de Apache-2.0 licentie verleent deze licentie geen recht op het gebruik van handelsnamen, merken of productnamen van de licentiegever, behoudens redelijk en gebruikelijk redactioneel gebruik ter aanduiding van de herkomst.
Available Tools
3 toolslekkage_checkA
Meet of de overwegingen van een tekst de afloop al prijsgeven — woorden als 'toewijsbaar', 'is ongegrond', 'wordt vernietigd', 'bewezen verklaard' — en of daar na de lekkage-knip R2 nog iets van overblijft. Gemeten wordt het deel vóór de beslissing: het dictum gaat eruit, samen met de 300 tekens aanloop ervóór, en bij een tekst van 200 tekens of meer zonder beslissingszin geldt het slot als dictum. Onder de 200 tekens wordt er niets weggeknipt en telt de hele tekst mee. 'Schoon' betekent dus: geen uitkomst-taal in het deel dat is overgebleven — geef een volledige uitspraak mee, want bij een kort fragment waarin de beslissing vroeg valt, blijft er niets te meten over. Gebruik dit om een dataset, een benchmark of andermans AI-claim te toetsen: 92% van de Nederlandse uitspraken verraadt de afloop woordelijk, waardoor een model dat daarop traint beter lijkt dan het is. Dit is een meting, geen voorspelling; gebruik voorspel_uitkomst als je een oordeel over de afloop wilt. Geeft terug: de lengte in tekens, welke categorieën uitkomst-taal ruw zijn gevonden (beroep_uitspraak, bevestiging_vernietiging, civiel_vordering, straf_uitspraak) en welke daarvan na de knip resteren. Geen sleutel nodig. Privacy: de tekst wordt voor de analyse naar de server van Rechtssysteem.ai gestuurd; stuur geen tekst die u niet mag delen.
| Name | Required | Description | Default |
|---|---|---|---|
| tekst | Yes | Nederlandse tekst om te meten, meestal een uitspraak of een trainingsvoorbeeld. Platte tekst, max 20.000 tekens; langere tekst wordt hier geweigerd voordat er iets wordt verstuurd. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it explains the cut rule (including the 300-character lead, the <200-character exception, and the dictum detection), what 'clean' means, the return values (length, categories, and what remains), that no key is needed, and the privacy implication that text is sent to a server. It even notes that longer text is rejected before transmission, a behavioral edge case.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but each sentence earns its place, covering purpose, algorithm, usage, returns, authentication, and privacy in a logical order. It is not overly verbose given the complexity, though it could be slightly more concise by merging some clauses. The front-loading of purpose and the cut rule is effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a measurement tool with no output schema and no annotations, the description is remarkably complete. It covers the input, the algorithm, the return format (including category names), the usage context, authentication (no key), and a critical privacy caveat. An agent can invoke it correctly and interpret the result without external knowledge.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the parameter fully (tekst, maxLength, plain text), so the baseline is 3. The description adds value by advising to supply a full statement ('geef een volledige uitspraak mee') and explaining why short fragments may be unmeasurable, which helps the agent provide an appropriate input. It also reinforces the max-length rejection behavior, which is not in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: measuring whether a text's reasoning reveals its outcome, using a defined 'lekkage-knip' (cut) that removes the dictum and preceding characters. It explicitly distinguishes itself from the sibling tool voorspel_uitkomst by noting it is a measurement, not a prediction, which fully disambiguates the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when to use this tool: to test a dataset, benchmark, or an AI claim, and explicitly says to use voorspel_uitkomst instead when an outcome judgment is desired. It also advises to provide a full text because a short fragment may leave nothing to measure, giving concrete usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rechtspraak_cijfersA
Geeft de benchmark-cijfers achter voorspel_uitkomst: 609.715 zaken, accuracy 78,2%, macro-F1 77,1% (5-fold CV), F1 per klasse, labelverdeling, restlekkage 0,1% na de knip (92% zonder knip) en de meerderheidsbaseline van 43,7%. Gebruik dit als iemand vraagt hoe goed het model is, of om een cijfer te controleren voordat je het citeert — niet om een zaak te beoordelen (dat is voorspel_uitkomst) of een tekst te meten (lekkage_check). De baseline hoort altijd naast de accuracy: zonder die 43,7% zegt 78,2% niets. Geen parameters, geen sleutel nodig, verstuurt geen tekst.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool has no parameters, requires no key, and sends no text, which clarifies its side-effect-free nature. It also enumerates the exact data it returns (metrics, leakage, baseline), giving a clear picture of behavior. It does not explicitly state the return format, but the listed contents suffice for a zero-parameter retrieval tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place. It front-loads the primary function, then lists the specific metrics, then gives usage guidance and a critical caution about the baseline, and finally states operational facts. No fluff or redundancy; the structure guides the reader logically from what to why to how.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no parameters and no output schema, the description is complete: it fully describes what the tool returns (all benchmark figures), when to use it, when not to, and important caveats (baseline alongside accuracy). An agent can invoke it correctly and interpret the output without ambiguity. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters and the schema is empty, so the description confirms this with 'Geen parameters'. It adds that no key is needed and it sends no text, which is useful operational context beyond the schema. Since there are no parameters to describe, this baseline of 4 is appropriate and the description adds value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool's purpose: providing benchmark figures behind voorspel_uitkomst. It lists specific metrics (accuracy, macro-F1, F1 per class, label distribution, leakage, baseline) and clearly distinguishes it from siblings by stating it is not for assessing cases (voorspel_uitkomst) or measuring text (lekkage_check). This makes the purpose unambiguous and differentiates it from all siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage guidance is explicit: 'Use this if someone asks how good the model is, or to check a figure before citing it — not to assess a case (that is voorspel_uitkomst) or to measure text (lekkage_check).' It also provides a critical rule: the baseline must always accompany accuracy, preventing misuse. This is clear when/when-not guidance with named alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voorspel_uitkomstA
Voorspelt de afloop van een Nederlandse rechtszaak uit de zaaktekst: afgewezen, gedeeltelijk of toegewezen. Gebruik dit voor een volledige zaak- of procestekst (dagvaarding, pleitnota, uitspraak), niet voor een samenvatting van een paar zinnen: na het wegknippen moeten er minstens 200 tekens overblijven, anders komt er een fout terug. Kies deze tool als je een oordeel over de afloop wilt; gebruik lekkage_check als je alleen wilt meten of een tekst de uitkomst al prijsgeeft, en rechtspraak_cijfers als het om de benchmark-cijfers zelf gaat. Werkwijze: het dictum en uitkomst-aankondigende zinnen gaan er eerst uit (lekkage-knip R2), het model oordeelt over wat overblijft. Gemeten over 609.715 zaken (5-fold CV): accuracy 78,2%, macro-F1 77,1%, tegen een meerderheidsbaseline van 43,7% — noem die baseline altijd naast de accuracy. Geeft terug: label, zekerheid, kansen per klasse en het gebruikte rechtsgebied. Zekerheid onder 55% betekent 'weet niet'; presenteer het dan ook zo. AI-gegenereerde risico-indicatie op grond van vergelijkbare rechtspraak, geen juridisch advies, niet bestemd voor gebruik door of namens een rechterlijke instantie. Vereist RECHTSSYSTEEM_API_KEY in de omgeving van deze MCP-server (max 20 verzoeken per minuut per sleutel); zonder sleutel volgt een fout in plaats van een oordeel. Privacy: de tekst wordt voor de analyse naar de server van Rechtssysteem.ai gestuurd; stuur geen tekst die u niet mag delen.
| Name | Required | Description | Default |
|---|---|---|---|
| tekst | Yes | Volledige Nederlandse zaak- of procestekst. Platte tekst, max 20.000 tekens; langere tekst wordt hier geweigerd voordat er iets wordt verstuurd. Dictum en uitkomst-zinnen mogen erin blijven staan: die worden er aan de serverkant uitgeknipt. | |
| rechtsgebied | No | Optioneel. Zet de rechtsgebied-feature van het model. Laat weg om het rechtsgebied door de server uit de tekst te laten afleiden; dat valt terug op 'onbekend' als de tekst te weinig houvast geeft. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden, and it does so richly: it explains the server-side leakage clipping workflow, model performance, returned fields, the 55% confidence threshold, API-key requirements, rate limits, external data transmission, and the non-advice/non-judicial-use disclaimer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but information-dense, with each block serving a purpose: purpose, sibling routing, method/performance, output shape, threshold, auth, and privacy. It is front-loaded with the core purpose. There is slight redundancy between the initial 'Gebruik dit' sentence and the later 'Kies deze tool' routing, so it is not maximally tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description covers everything an agent needs to invoke the tool correctly and interpret its result: input constraints, failure cases, required environment variable, rate limiting, output fields, confidence threshold, privacy implications, and alternative sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already describes both parameters at 100% coverage, so the baseline is 3. The description still adds value by specifying the 200-character post-clipping minimum, clarifying that dictum/outcome sentences can remain because they are removed server-side, and giving the exact class labels returned.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a concrete action ('Voorspelt de afloop'), a specific object ('Nederlandse rechtszaak uit de zaaktekst'), and the exact output classes ('afgewezen, gedeeltelijk of toegewezen'). It also explicitly separates this tool from its siblings, so an agent can distinguish it without opening their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to choose this tool ('als je een oordeel over de afloop wilt'), when to use lekkage_check instead, and when to use rechtspraak_cijfers. It also gives concrete input requirements: use a full case text, not a short summary, and keep at least 200 characters after trimming.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.1- Changed
lekkage_check3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / tekst / descriptionAdded value: +"Nederlandse tekst om te meten, meestal een uitspraak of een trainingsvoorbeeld. Platte tekst, max 20.000 tekens; langere tekst wordt hier geweigerd voordat er iets wordt verstuurd." - added
Input schema / properties / tekst / maxLengthAdded value: +20000
- Changed
rechtspraak_cijfers1 field changed- added
Input schema / additionalPropertiesAdded value: +false
- Changed
voorspel_uitkomst5 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / rechtsgebied / descriptionAdded value: +"Optioneel. Zet de rechtsgebied-feature van het model. Laat weg om het rechtsgebied door de server uit de tekst te laten afleiden; dat valt terug op 'onbekend' als de tekst te weinig houvast geeft." - added
Input schema / properties / rechtsgebied / enumAdded value: +[ + "bestuursrecht", + "civiel recht", + "strafrecht", + "overig", + "onbekend" +] - added
Input schema / properties / tekst / descriptionAdded value: +"Volledige Nederlandse zaak- of procestekst. Platte tekst, max 20.000 tekens; langere tekst wordt hier geweigerd voordat er iets wordt verstuurd. Dictum en uitkomst-zinnen mogen erin blijven staan: die worden er aan de serverkant uitgeknipt." - added
Input schema / properties / tekst / maxLengthAdded value: +20000
3 tool updates
v0.1.0- First observed
lekkage_check - First observed
rechtspraak_cijfers - First observed
voorspel_uitkomst
TDQS
Scored across 3 tools
Each tool targets a clearly distinct concern: predicting case outcomes, retrieving benchmark metrics, and checking for verdict leakage. The descriptions explicitly cross-reference each other to prevent misselection.
All names are lowercase snake_case, but the pattern is mixed: voorspel_uitkomst is verb+object, while rechtspraak_cijfers and lekkage_check are noun phrases. The names are readable and coherent within the domain, but there is no consistent verb_noun convention.
Three tools form a tight, well-scoped set: predict an outcome, retrieve the underlying benchmark metrics, and measure leakage. Each tool earns its place with minimal overlap.
The surface covers the full workflow for this specialized purpose: model prediction, benchmark verification, and data-quality/leakage checking. No obvious missing operations for the stated domain.
Maintenance
Related MCP Connectors
Slovak court decisions as MCP tools. 12,000+ decisions, GDPR-compliant, pseudonymized, SLA-backed.
Source-anchored search and civic intelligence for Dutch municipal council records.
Deterministic prompt-injection detector; signed, offline-verifiable verdicts. Not an LLM.
Explainable ambiguity signals for text. Submitted text is processed but never retained.
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceAn MCP server that provides comprehensive US legislation.1337MIT

Vaquill-AI/vaquill-mcpofficial
AlicenseAqualityAmaintenanceMCP server for Vaquill legal research API. Covers US federal + 50-state law (USC, CFR, state legislation, CourtListener case law)257MIT
conformi-searchofficial
AlicenseAqualityAmaintenanceInstallable MCP server for EU legal research with verifiable CELEX citations from the EUR-Lex corpus (DE/EN/FR).21MIT- FlicenseNot gradedqualityDmaintenanceEnables semantic and keyword search over legal documents, conflict detection, and document overview, supporting Indonesian and English texts.-