precis-mcp
A local data-quality validation server that exposes 4 tools for checking Excel/CSV/TSV tabular data against a Precis project manifest.
validate_data— Run validation against aproject.precis.yamlmanifest and get contract JSON (is_valid/errors/summary); optionally scope to a single table or a custom data directory.check_config— Inspect project configuration loading (loading_errors/ load counts) without executing any validation.describe_constraints— List all available constraint types with theirrefs/paramsdocumentation (no inputs required).infer_schema— Infer column types from a CSV/Excel/JSON file and return a draft schema YAML, with optional table id, display name, sample row count, andsource.path.
Precis
本地优先的可视化数据质量工具 / Local-First Visual Data Quality Tool
可视化建模 · 全链路校验 · 数据不出本机
Alpha — 核心功能已实现,接口与配置格式可能调整,暂不建议生产环境使用。 当前阶段暂不接受外部 Pull Request,欢迎 Issues 与 Discussions。
Alpha stage. Core features are implemented; interfaces and config formats may change. Not recommended for production. No external PRs accepted at this stage. Issues and Discussions welcome.
项目简介
Precis 是一款针对 Excel / CSV / TSV 表格数据的质量校验工具。校验流程在可视化画布上以节点和连线的方式编排——从数据接入、清洗转换到多维度质量检查,全程无需编写代码。所有数据均在本地处理,不上传至任何外部服务。
提供三种使用方式,按需选择:
桌面应用(推荐):图形界面,开箱即用
命令行:适合批量执行,或交由 AI 编程助手(Kimi Code / Claude Code 等)调用
HTTP 接口:用于集成到自有系统

Related MCP server: dq-mcp
功能特性
可视化编排 — 在画布上拖拽节点、连接流程,即可完成校验建模
10 种检查规则 — 必填校验、唯一性、引用完整性、允许值清单、数值范围、条件判断、自定义脚本、字符集、日期逻辑、多规则组合
22 种数据转换 — 字符串拆分、模式提取、数学计算、分组聚合、过滤、排序等
大文件支持 — 超大文件自动分块处理
本地运行 — 数据不出本机;中英文双语界面;桌面应用内置运行环境,安装即可使用
快速开始
环境要求
工具 | 版本 | 说明 |
Node.js |
| 含 npm |
Python |
| 3.12 或 3.13 |
Rust | stable | 可选,仅构建终端界面时需要 |
node --version # 应 ≥ 20.19.0 或 ≥ 22.12.0
python3 --version # 应为 3.12.x 或 3.13.x若系统默认 python3 低于 3.12(macOS 自带 3.9),可通过 Homebrew 安装:
brew install python@3.12 # 安装后以 python3.12 调用安装
git clone https://github.com/AirSaiga/Precis.git
cd Precis
npm run setup:mac # macOS / Linux
# npm run setup:win # Windows一键脚本将依次完成:Python 环境检测、虚拟环境创建、依赖安装、前端与桌面应用构建。
无需图形界面、仅需在命令行或 AI 助手中执行校验,可直接
pip install precis-cli获取precis命令;接入 AI 编程助手的方式见integrations/。
git clone https://github.com/AirSaiga/Precis.git
cd Precis
# 1. 安装前端依赖(根目录一次安装,覆盖全部子项目)
npm run install:all
# 2. 配置后端环境
cd backend
python3.12 -m venv .venv # 若命令为 python3.13 则相应替换
source .venv/bin/activate # Windows PowerShell: .venv\Scripts\Activate.ps1;Git Bash: source .venv/Scripts/activate
pip install --upgrade pip
pip install -e ".[dev]" # 含 pytest / ruff / mypy 等开发工具
cd ..注意:凡调用
python的 npm 脚本(dev/backend:dev/cli/electron:dev),执行前须先运行source backend/.venv/bin/activate,否则将使用系统自带的低版本 Python。
启动
npm run electron:dev # 桌面应用(推荐):自动启动后端与前端
npm run dev # 开发模式:前后端分离运行,支持热更新
npm run cli # 纯命令行,不启动图形界面
npm run start:tui # 终端界面(实验性,不随发布版本提供)更多启动方式与脚本说明见
scripts/README.md与tui-rust/README.md。
验证安装
npm run cli:validate # 使用内置示例数据(qa_test/qa_simple/)执行一次校验正常输出校验结果即表示环境就绪。
MCP Server(AI 助手直连)
本仓库实现了一个 Model Context Protocol (MCP) server(stdio 传输,基于官方 MCP Python SDK)。支持 MCP 的 AI 编程助手(Claude Code、Cursor、Kimi Code 等)配置后可直接调用校验引擎。
pip install "precis-cli[mcp]" # 安装后获得 precis-mcp 命令在 MCP 客户端中配置(mcpServers 格式,Claude Code / Cursor 等通用):
{
"mcpServers": {
"precis": {
"command": "precis-mcp"
}
}
}提供 4 个工具:
工具 | 作用 |
| 执行校验,返回结构化错误报告(与 CLI |
| 检查项目配置加载情况 |
| 列出全部约束类型与参数说明 |
| 从数据文件推断 schema 草稿 |
更多 AI 助手接入方式(skill、插件包等)见 integrations/。
项目结构
Precis/
├── backend/ # 校验引擎与命令行
├── frontend/ # 可视化编辑器
├── electron/ # 桌面应用外壳
├── tui-rust/ # 终端界面
├── e2e/ # 自动化测试
├── qa_test/ # 内置示例数据
├── scripts/ # 构建与部署脚本
└── docs/ # 架构文档更多文档
开发者:开发命令与打包说明见 CONTRIBUTING.md;架构细节见 AGENTS.md 与 docs/ARCHITECTURE.md
CHANGELOG.md — 变更日志
SECURITY.md — 安全说明
各子项目详情:backend · frontend · electron · tui-rust · scripts · e2e
许可证
Apache-2.0 — 详见 LICENSE_NOTICE.md。
Overview
Precis is a data quality tool for Excel / CSV / TSV tabular data. Validation workflows are composed on a visual canvas using nodes and connections — from data ingestion and transformation to multi-dimensional quality checks, without writing any code. All data is processed locally and never uploaded to any external service.
Three usage modes, choose as needed:
Desktop app (recommended): graphical interface, ready out of the box
Command line: suitable for batch execution, or for invocation by AI coding assistants (Kimi Code / Claude Code, etc.)
HTTP API: for integration into your own systems

Features
Visual composition — build validation workflows by arranging nodes and connections on a canvas
10 check rules — required values, uniqueness, referential integrity, allowed value lists, numeric ranges, conditional checks, custom scripts, character sets, date logic, and multi-rule combinations
22 data transforms — string splitting, pattern extraction, math evaluation, grouping and aggregation, filtering, sorting, and more
Large file support — oversized files are automatically processed in chunks
Runs locally — data never leaves your machine; bilingual Chinese/English interface; the desktop app bundles its runtime and works immediately after installation
Quick Start
Prerequisites
Tool | Version | Notes |
Node.js |
| includes npm |
Python |
| 3.12 or 3.13 |
Rust | stable | optional, only required to build the terminal UI |
node --version # should be ≥ 20.19.0 or ≥ 22.12.0
python3 --version # should be 3.12.x or 3.13.xIf the system default python3 is older than 3.12 (macOS ships 3.9), install one via Homebrew:
brew install python@3.12 # invoke as python3.12 after installationInstallation
git clone https://github.com/AirSaiga/Precis.git
cd Precis
npm run setup:mac # macOS / Linux
# npm run setup:win # WindowsThe setup script performs, in order: Python detection, virtual environment creation, dependency installation, and frontend / desktop app build.
If you only need validation from the command line or an AI assistant, run
pip install precis-clito get thepreciscommand; seeintegrations/for AI coding assistant integration.
git clone https://github.com/AirSaiga/Precis.git
cd Precis
# 1. Install frontend dependencies (single root install covering all sub-projects)
npm run install:all
# 2. Set up the backend environment
cd backend
python3.12 -m venv .venv # substitute python3.13 if that is your command
source .venv/bin/activate # Windows PowerShell: .venv\Scripts\Activate.ps1; Git Bash: source .venv/Scripts/activate
pip install --upgrade pip
pip install -e ".[dev]" # includes pytest / ruff / mypy and other dev tools
cd ..Note: any npm script that invokes
python(dev/backend:dev/cli/electron:dev) requiressource backend/.venv/bin/activatefirst; otherwise the system's older Python will be used.
Launch
npm run electron:dev # Desktop app (recommended): starts backend and frontend automatically
npm run dev # Dev mode: backend and frontend run separately with hot reload
npm run cli # Command line only, no graphical interface
npm run start:tui # Terminal UI (experimental, not included in releases)Additional launch options and script details:
scripts/README.mdandtui-rust/README.md.
Verify the Installation
npm run cli:validate # runs a validation pass on the bundled sample data (qa_test/qa_simple/)Successful validation output indicates the environment is ready.
MCP Server (for AI assistants)
This repository implements a Model Context Protocol (MCP) server over stdio, built on the official MCP Python SDK. MCP-capable AI coding assistants (Claude Code, Cursor, Kimi Code, etc.) can call the validation engine directly once configured.
pip install "precis-cli[mcp]" # provides the precis-mcp commandClient configuration (mcpServers format, works with Claude Code / Cursor and others):
{
"mcpServers": {
"precis": {
"command": "precis-mcp"
}
}
}Four tools are exposed:
Tool | Purpose |
| Run validation, returning a structured error report (same contract as CLI |
| Check project configuration loading |
| List all constraint types and their parameters |
| Infer a schema draft from a data file |
More AI assistant integration options (skills, plugin packages): integrations/.
Project Structure
Precis/
├── backend/ # Validation engine and command line
├── frontend/ # Visual editor
├── electron/ # Desktop app shell
├── tui-rust/ # Terminal interface
├── e2e/ # Automated tests
├── qa_test/ # Bundled sample data
├── scripts/ # Build and deployment scripts
└── docs/ # Architecture documentationMore Documentation
Developers: dev commands and packaging notes in CONTRIBUTING.md; architecture details in AGENTS.md and docs/ARCHITECTURE.md
CHANGELOG.md — Changelog
SECURITY.md — Security notes
Per-project details: backend · frontend · electron · tui-rust · scripts · e2e
License
Apache-2.0 — See LICENSE_NOTICE.md for details.
Available Tools
4 toolscheck_configA
Load and inspect a Precis project configuration file without validating data. Use it when validate_data fails, or to diagnose configuration problems first (unsupported version, missing files, dangling references); read-only. Returns: manifest_path, version_ok (whether the manifest version is supported), schemas_loaded/constraints_loaded (counts of schema/constraint files loaded successfully), loading_errors (one entry per load error), warnings. The manifest must be inside the server working directory and must exist, otherwise the call fails.
| Name | Required | Description | Default |
|---|---|---|---|
| manifest | Yes | Absolute path to project.precis.yaml (must be inside the server working directory; paths outside it are rejected) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it declares read-only semantics, that no validation is performed, the prerequisite that the manifest must exist inside the server working directory, and the hard failure mode otherwise. It stops short of describing idempotency or performance/IO characteristics, but covers the safety-relevant behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then usage routing, then prerequisites and returns. The return-field enumeration is long but earns its place because no output schema exists. Slightly dense, but no filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema means the description must enumerate return values, and it does (manifest_path, version_ok, loaded counts, loading_errors, warnings). Combined with the stated prerequisites and failure conditions, an agent has everything needed to call it correctly, which is unusual for a tool with no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One parameter at 100% schema coverage, so the schema already documents the absolute-path requirement and the working-directory rejection. The description restates the same constraint rather than adding format, default, or resolution detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Load and inspect a Precis project configuration file') plus a scope qualifier ('without validating data') that immediately separates it from the sibling validate_data. An agent can distinguish it from describe_constraints and infer_schema without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('when validate_data fails') and a diagnostic purpose (unsupported version, missing files, dangling references), naming the alternative sibling outright. This is exactly the when/when-not/alternative routing an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_constraintsA
List all 10 constraint types supported by Precis (NotNull/Unique/AllowedValues/Range/ForeignKey/Conditional/Scripted/Charset/DateLogic/Composite) with their refs/params documentation. Use it when writing or editing *.constraint.yaml files, or when unsure which parameters a constraint type accepts; takes no arguments, read-only. Returns a types array whose entries contain: type (constraint type name), refs (documentation of the referenced schema table/column IDs), params (parameter keys, allowed values, defaults). The content is derived from the actions registry as the single source of truth, so it matches the validation engine behavior.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well: it declares 'takes no arguments, read-only', discloses the return shape, and notes the content is derived from the actions registry as the single source of truth matching the validation engine. This is real behavioral context beyond a bare description, though it omits things like caching or versioning.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then usage, then return shape. The parenthetical enumeration of all 10 type names is slightly verbose but genuinely informative since no output schema exists. No wasted filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description compensates fully: it explains what is returned (types array with type/refs/params), the provenance of the data, and the read-only no-arg nature. An agent has everything needed to call and interpret this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description reinforces this with 'takes no arguments' and instead documents the semantics of the returned refs/params fields, which is useful but not parameter guidance per se.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: list the 10 constraint types Precis supports, and even enumerates them (NotNull/Unique/.../Composite). An agent immediately knows this is a reference-lookup tool rather than a validator or schema inferrer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use it 'when writing or editing *.constraint.yaml files, or when unsure which parameters a constraint type accepts', which is a clear triggering context. It does not name or contrast with the sibling tools (validate_data, check_config, infer_schema), but those have distinct purposes so the risk of confusion is low.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
infer_schemaA
Infer column types by sampling the head of a data file (CSV/Excel/JSON/JSONL) and generate a draft V2 schema structure. Use it when creating a schema for a new data source or rebuilding an existing schema; read-only, nothing is written to disk - the caller decides whether to save the returned draft. Returns a schema dictionary with V2 fields such as id, name and columns (each column carries its inferred type: string/integer/float/decimal/boolean/date). Note: type inference is based on a sample (1000 rows by default), so extreme values outside the sample may change the actual type; review the draft by hand before verifying it with validate_data. The file path must be inside the server working directory, otherwise the call fails.
| Name | Required | Description | Default |
|---|---|---|---|
| table_id | No | Table ID. Pass the existing id when replacing a schema to preserve references; a new id is generated otherwise | |
| data_file | Yes | Path to a CSV/Excel/JSON/JSONL data file (must be inside the server working directory; paths outside it are rejected) | |
| table_name | No | Table display name; defaults to the data file name | |
| sample_rows | No | Number of rows to sample (default 1000; must be a positive integer). Larger samples infer types more accurately but run slower | |
| source_path | No | source.path written into the schema (path to the data file, relative); omitted when not provided |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it declares read-only behavior, that nothing is written to disk and the caller decides whether to save, the approximate sample size driving inference, the failure mode for paths outside the working directory, and the caveat that out-of-sample extreme values may change the real type.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, followed by usage, return shape, caveats and constraints in a logical order. It is on the long side and repeats the sample-size default that the schema already states, but every sentence carries usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description justifiably describes the returned schema dictionary and its V2 fields. Combined with the usage trigger, the sampling caveat and the path restriction, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters like table_id, sample_rows and source_path are already documented in the schema; the description largely restates the default 1000-row sample and the working-directory path constraint. It adds the semantic link between the draft's column types and the sample, but little beyond structured data.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('infer column types by sampling the head of a data file') and names the artifact produced ('generate a draft V2 schema structure'). It is clearly distinguishable from validate_data, which the description positions as a separate verification step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('when creating a schema for a new data source or rebuilding an existing schema') and points to validate_data as the downstream verification step. It stops short of stating when *not* to use it or naming a true alternative tool for the same job.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_dataA
Validate Precis project data against all constraint rules (10 types: NotNull, Unique, AllowedValues, Range, ForeignKey, Conditional, Scripted, Charset, DateLogic, Composite). Use it to check data quality or to re-validate data after editing constraint configuration; read-only, it modifies no files. Returns the contract JSON: is_valid (overall pass/fail), errors (one entry per violation, with table name, column name, row number, error_code and details), summary (violation counts); full field definitions in docs/contracts/validate-json-v1.md. Preconditions: manifest points to an existing project.precis.yaml inside the server working directory; a path outside the working directory or a missing file returns an isError result.
| Name | Required | Description | Default |
|---|---|---|---|
| table | No | Validate only this table (schema id or table display name); defaults to validating all tables | |
| manifest | Yes | Absolute path to project.precis.yaml (must be inside the server working directory; paths outside it are rejected) | |
| data_directory | No | Root directory of the data files. Relative data-source paths declared in the schema resolve against this directory; defaults to the directory containing the manifest |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it declares read-only behavior ('modifies no files'), spells out the exact return contract (is_valid, errors shape, summary), and names the failure mode (path outside working dir or missing file returns isError). Preconditions are stated up front. This is unusually complete disclosure for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then usage, return contract, and preconditions in a logical order; every sentence carries information. It is dense and runs long for a single paragraph, which costs it a point on readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description documents the return shape inline and points to docs/contracts/validate-json-v1.md for full field definitions. Preconditions and error behavior are covered, so an agent has everything needed to invoke and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all three parameters, giving a baseline of 3. The description reinforces the manifest constraint (must point to an existing project.precis.yaml inside the working directory) but adds nothing about the table or data_directory parameters beyond what the schema already says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Validate Precis project data against all constraint rules') and even enumerates the 10 constraint types, which is far more precise than the sibling names check_config, describe_constraints, or infer_schema imply. An agent can tell this is the data-validation tool, not a schema/constraint introspection tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to reach for it: 'check data quality or to re-validate data after editing constraint configuration'. That covers the two main use contexts, but it never contrasts itself against the sibling tools (e.g. describe_constraints) so the routing is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.11- Changed
check_config1 field changed- changed
Input schema / properties / manifest / descriptionPrevious value: -"project.precis.yaml 路径"New value: +"Absolute path to project.precis.yaml (must be inside the server working directory; paths outside it are rejected)"
- Changed
infer_schema5 fields changed- changed
Input schema / properties / data_file / descriptionPrevious value: -"CSV/Excel/JSON 数据文件路径"New value: +"Path to a CSV/Excel/JSON/JSONL data file (must be inside the server working directory; paths outside it are rejected)" - changed
Input schema / properties / sample_rows / descriptionPrevious value: -"采样行数(默认 1000)"New value: +"Number of rows to sample (default 1000; must be a positive integer). Larger samples infer types more accurately but run slower" - changed
Input schema / properties / source_path / descriptionPrevious value: -"写入 schema 的 source.path"New value: +"source.path written into the schema (path to the data file, relative); omitted when not provided" - changed
Input schema / properties / table_id / descriptionPrevious value: -"表 ID(替换既有 schema 时传原 id)"New value: +"Table ID. Pass the existing id when replacing a schema to preserve references; a new id is generated otherwise" - changed
Input schema / properties / table_name / descriptionPrevious value: -"表显示名"New value: +"Table display name; defaults to the data file name"
- Changed
validate_data3 fields changed- changed
Input schema / properties / data_directory / descriptionPrevious value: -"数据目录(缺省为 manifest 所在目录)"New value: +"Root directory of the data files. Relative data-source paths declared in the schema resolve against this directory; defaults to the directory containing the manifest" - changed
Input schema / properties / manifest / descriptionPrevious value: -"project.precis.yaml 路径(须在工作目录内)"New value: +"Absolute path to project.precis.yaml (must be inside the server working directory; paths outside it are rejected)" - changed
Input schema / properties / table / descriptionPrevious value: -"只校验指定表(缺省校验全部)"New value: +"Validate only this table (schema id or table display name); defaults to validating all tables"
4 tool updates
v0.1.0- First observed
check_config - First observed
describe_constraints - First observed
infer_schema - First observed
validate_data
TDQS
Scored across 4 tools
Each tool has a largely distinct purpose: validate_data checks data, check_config inspects a project's config files, describe_constraints serves as generic reference documentation, and infer_schema produces draft schemas. The main overlap is between validate_data and check_config, since both are used for diagnosing failures, but the descriptions explicitly frame check_config as the follow-up to a validate_data failure, which mitigates most confusion.
All four tools follow a clean verb_noun snake_case convention: validate_data, check_config, describe_constraints, infer_schema. There are no deviations or mixed conventions.
Four tools is a good fit for a focused validation/diagnostics server; each tool serves a clear purpose (validate, diagnose config, reference docs, infer schema). It is on the lighter side, so slightly under-scoped but not thin enough to be a problem.
The read-only validation lifecycle is well covered: validate, diagnose, reference constraint types, and draft schemas. Minor gaps exist, such as no tool to inspect an existing saved schema or enumerate a project's declared constraints, but these are workarounds rather than blockers.
Maintenance
Related MCP Connectors
Validate and clean CSV before import. Find duplicates, missing values and invalid dates with row-level reports. Apply only the cleanup you request; files are not retained. Run view_csv_demo free without a key. Custom CSV: EUR 9 for 100 operations, valid 90 days, no subscription. Setup: https://check.orvel.dev/docs/
Auto-discover validation rules from data — scan, profile, health-score. No rules to write.
Validate JSON, YAML, XML and CSV with exact line/column errors and silent-corruption warnings.
Open, inspect, filter, edit and convert xlsx and csv files from your AI chat. Processing is local.
Related MCP Servers
- AlicenseAqualityAmaintenanceZero-config data quality monitoring as MCP tools. Profiles a warehouse (Postgres, BigQuery, Snowflake, MySQL, DuckDB), detects anomalies, and gates CI — read-only with the connection resolved server-side, never via the model.652 PyPI11MIT
- AlicenseNot gradedqualityCmaintenanceEnables language models to run data-quality checks and profiling on local files, using dbt-style assertions like not_null, unique, relationships, and accepted_values.MIT
- AlicenseNot gradedqualityAmaintenanceLocal-first MCP server for data quality that finds suspicious data, explains findings with evidence, tracks drift, and supports human-approved, reversible repair workflows. Deterministic by default, with AI optional.108 PyPI2Apache 2.0
- AlicenseAqualityCmaintenanceEnables validating CSV structure, checking simple schemas, converting between CSV and JSON, sampling rows, and finding duplicate keys on local files. All processing stays local, so no user data is ever uploaded.6MIT