misata-mcp
1-line summary: The Misata MCP server turns plain-English stories or declared schemas into relational synthetic databases — previewing, generating, seeding, and auditing data for AI agents and developers.
Discover domains (
list_domains) — list the 18 built-in business domains with trigger keywords and sample stories.Preview interpretation (
preview_story) — see the detected domain, confidence, locale, scale, and tables without generating rows.Inspect schema (
inspect_schema) — return full tables, columns, relationships, and distribution params for a story.Generate from a sentence (
generate_dataset) — write a multi-table relational dataset as CSVs with seedable, reproducible output.Generate from a schema (
generate_from_schema) — the primary tool: design tables, columns, FKs, constraints, correlations, distributions, rollups, time series, state machines, outcome curves, rate curves, waterfalls, stock flows, lifecycles, and deliberate dirty data with exact-count defects.Seed a live database (
seed_database) — introspect a real Postgres/SQLite DB, plan the insert order, then apply with integrity verification (plans by default; requires explicit apply/truncate/append).Validate YAML (
validate_yaml) — structurally and semantically check a misata.yaml before generation.Audit a dataset (
audit_dataset) — score existing CSVs 0–100 for reader-visible contradictions (backwards timestamps, non-reconciling derived columns, geography mismatches, constant columns).Validate domain rules (
validate_domain) — flag physiologically or financially impossible values against clinical, clinical_trial, financial, or fintech ranges.
Misata
The relational synthetic data engine that satisfies exact business outcomes.
Generate complete multi-table databases with foreign keys that resolve, math that reconciles, and domain-authentic prose with zero Lorem Ipsum. From a single sentence, YAML schema, or live database.
Prefer a visual interface? Try Misata Studio to design schemas on an interactive canvas and generate datasets directly in your browser.
The Paradigm Shift
Most synthetic data tools take existing data and imitate it. But in modern engineering, you rarely have clean data to start with—or you need data designed around a specific target outcome:
"Monthly revenue rises from $50k to $200k with a Q3 slump."
"Fraud rate starts at 1% in Q1 and climbs to 6% by Q4."
"Every customer's
total_spentstrictly equals the sum of their order line items."
Misata works in reverse: you declare the outcome, and Misata solves for the individual rows that hit it to $0.00 error, guaranteed by a closed-form Gamma conditional-sum mechanism (arXiv:2606.08736). No machine learning models, no real data required, and zero hallucinated foreign keys.
⚡ 30-Second Quickstart
Install via pip:
pip install misataCLI
Generate a complete relational dataset from a single description:
misata generate \
--story "Brazilian fintech with R$ payments, CPF verification, and 3% fraud" \
--rows 1000 \
--output-dir ./demo_dataOutputs clean CSVs and an oracle_report.json verifying zero foreign key orphans, constraint satisfaction, and statistical fidelity.
Python
import misata
# Generate multi-table DataFrames from plain English
data = misata.generate("SaaS startup with 500 users, monthly subscriptions, and 12% churn")
users_df = data["users"]
subscriptions_df = data["subscriptions"]Enrich Existing Data in 1 Line
Replace boring or blank columns in an existing DataFrame with domain-authentic prose:
import misata
import pandas as pd
df = pd.read_csv("support_cases.csv")
# Automatically detects semantic columns (subjects, resolution notes, memos, error traces)
df_enriched = misata.enrich_text(df, seed=42)Related MCP server: MCP Data Visualization Server
🚀 The 6 Unique Capabilities of Misata
What separates Misata from legacy libraries like Faker and imitation models like SDV:
1. Exact Outcome Conformance ($0.00 Error)
Declare the aggregate outcome (revenue curve, churn rate, seasonal surge, default curve), and Misata generates micro-rows whose monthly or annual sums match your target to $0.00 error. While off-the-shelf synthesizers miss aggregate targets by 74–86%, Misata hits them provably (read the research paper).
2. Instant Sandbox Oracle for AI Coding Agents
Misata includes a native Model Context Protocol (MCP) server. Cursor, Claude Code, and Windsurf can call create_sandbox to spin up an isolated SQLite database seeded with realistic relational data in 2 seconds. AI agents can test SQL queries and test code against real tables instead of guessing schema.
pip install "misata[mcp]" && misata mcp install --client all3. Topological DAG Relational Integrity (0 Orphan Foreign Keys)
Parent-child relationships across 10+ tables are resolved via topological rank ordering. Self-referential hierarchies, composite unique constraints, and temporal causality (signup_date <= order_date <= payment_date <= refund_date) are strictly enforced.
4. Cross-Vertical Text Realism (Zero Lorem Ipsum)
No Latin placeholder gibberish. Combinatorial microtext generators provide authentic text across 15+ real-world industries:
E-Commerce: Category-conditioned product descriptions (electronics, apparel, home, beauty), authentic return reasons, delivery notes.
B2B SaaS: Support ticket subjects, multi-line issue descriptions, agent resolution notes, competitor churn reasons.
FinTech: Bank statement descriptors, ACH/wire remittance memos, AML audit overrides.
Healthcare: Clinical SOAP progress notes, chief complaints, discharge instructions, medical dosage schedules (
TID,QID,PRN).Customer Reviews: Sentiment calibrated provably to 1-to-5 star ratings.
5. Vectorized Speed (100x–500x Faster Than Faker)
While Faker iterates row-by-row in pure Python (5,000–15,000 rows/s), Misata uses vectorized NumPy array operations. It generates 500,000 to 16,000,000 rows/second on a single CPU core.
6. Introspect & Seed Live Databases (misata seed)
Point Misata directly at a PostgreSQL, MySQL, or SQLite database. Misata introspects foreign keys, check constraints, enums, and column types, generates a topological insertion plan, and seeds production-like data directly back into your database without writing schema code.
🎯 Dozens of Real-World Use Cases
Misata is built for engineers, testers, data teams, and founders across dozens of everyday workloads:
🛠️ Data Engineering & ETL Pipelines
Known-Answer Pipeline Testing: Declare exact KPI targets (e.g. $1.2M Q4 revenue), generate synthetic raw tables, and verify your dbt, Spark, or SQL transforms return the exact expected figure.
High-Throughput Stress Testing: Churn 10,000,000+ rows in seconds to test partition boundaries, shuffle performance, and warehouse scaling.
Schema Migration Dry Runs: Rehearse destructive column backfills and foreign key additions against realistic data before deploying to production.
Deterministic CI/CD Fixtures: Reproducible seeds ensure test assertions never flake across continuous integration runs.
🤖 AI Coding Agents & LLM Development
Agent SQL Sandboxes: Provide Cursor, Claude Code, and Windsurf with isolated, pre-seeded databases to validate SQL queries without production access.
Text-to-SQL Benchmark Generation: Build complex relational schema benches with diverse joins to evaluate fine-tuned coding models.
Evaluation Database Packs (Evalpacks): Create verified eval datasets with independent DuckDB answer keys where the ground truth cannot be wrong.
🗄️ Application Development & Database Seeding
Local & Staging Environment Seeding: Seed staging Postgres, MySQL, or SQLite databases with realistic customer histories and 0 broken foreign keys (
misata seed).ORM Model Fixtures: Generate rich fixtures matching your Prisma schema or SQLAlchemy declarative models.
Multi-Tenant Isolation Verification: Test row-level security (RLS) and tenant isolation rules without cross-tenant key leakage.
Incremental Data Growth (
generate_diff): Add 5,000 new rows to an existing dataset while auto-offsetting IDs and preserving referential integrity.
📊 BI, Product Demos & Sales Engineering
Board-Ready Dashboard Demos: Populate Tableau, PowerBI, and Metabase dashboards with convincing seasonality (Black Friday spikes, summer dips) rather than flat random noise.
Sales Engineering Prototypes: Demo customer-facing analytics with authentic company names, human names, and transaction histories with zero PII exposure.
Feature Previews: Preview upcoming charts, cohorts, and metrics before production customer data accumulates.
💳 FinTech, Banking & Payments
Double-Entry Ledger Balancing: Generate accounting transactions where total debits strictly equal credits across every ledger account.
Credit Risk & Loan Tapes: Calibrate delinquency curves, credit score distributions, and default rates for credit portfolio testing.
AML & Fraud Detection Testing: Inforce exact fraud incidence rates (e.g. 2.4%) with authentic transaction memos and AML audit trails.
Payment Remittance: Simulate SWIFT, ACH, and card transactions with valid routing numbers, CVVs, and statement descriptors.
🏥 Healthcare & Clinical Informatics
HIPAA Safe-Harbor Synthetic Cohorts: Generate realistic patient populations, vital signs, and encounter histories with zero PHI liability.
Clinical NLP Model Evaluation: Evaluate healthcare LLMs against authentic SOAP notes, chief complaints, and discharge summaries.
Ward & Scheduling Simulation: Simulate hospital appointment grids with realistic 15-minute intervals, business hours, and weekend dips.
📦 E-Commerce & Supply Chain Logistics
Multi-Category Catalog Modeling: Produce realistic item specs and descriptions conditioned on category (electronics, apparel, home, industrial).
Return & Refund Workflows: Simulate return logistics with authentic return reasons, restocking milestones, and customer refund dates.
Route & Fleet Optimization: Compute realistic routes with Haversine distance calculations and valid secondary addresses (Apt, Suite, Bldg).
🛡️ Cybersecurity & IT Infrastructure
Network Intrusion Datasets: Generate netflow logs, port scans, and DDoS traffic patterns for security tool benchmarking.
System Exception & Error Analysis: Populate observability dashboards with realistic deadlocks, HTTP 504 timeouts, and connection pool exhaustion logs.
Compliance Audit Logging: Simulate SOC2/HIPAA access logs with documented managerial access override justifications.
🔬 Machine Learning & Statistical Research
Synthetic Twins from CSV (
misata.mimic): Clone distributions and correlations from sensitive CSVs without copying a single original row.Hierarchical Cluster Modeling (ICC): Generate multi-site data with specified Intraclass Correlation Coefficients for mixed-effects regression.
Time-Series Autocorrelation (AR1): Generate longitudinal entity trajectories that maintain realistic temporal memory.
⚡ Why You Should Never Use Faker Again
Faker was built over a decade ago for single-attribute mock values. For modern applications, it introduces critical failure modes:
Problem in 2026 | Faker Reality | Misata 0.9.6.60 Advantage |
Relational Topology | ✗ 0 concept of databases or FKs; manual glue code required | ✓ Strict topological DAG; 0 orphan FKs guaranteed |
Cross-Column Coherence | ✗ Incoherent (e.g. "Male" name, mismatched email, invalid city) | ✓ Coherent identities, addresses, and causality |
Text Realism | ✗ 2,000-year-old Latin "Lorem Ipsum" or robotic templates | ✓ 15+ domain microtext pools (SOAP notes, tickets, memos) |
Mathematical Consistency | ✗ Violates basic accounting ( | ✓ Exact mathematical formulas and balanced ledgers |
Performance | ✗ ~10k rows/s (single-threaded Python loops) | ✓ 500k to 16M rows/s (Vectorized NumPy engine) |
Aggregate Targets | ✗ Impossible (uniform random noise) | ✓ Exact closed-form outcome conformance ($0.00 error) |
Database Seeding | ✗ Manual SQL scripts or ORM boilerplate | ✓ One-command introspection and seeding ( |
Read the complete Faker vs SDV vs Misata Guide for full benchmarks and code comparisons.
🛠️ Eight Ways to Generate Data
Misata fits whatever workflow you already use:
Input Mode | Best For | Learn More |
1. Plain English Story | Rapid prototyping, zero configuration | |
2. YAML Schema-as-Code | Committing versioned data definitions to git | |
3. Live Database Seeding | Introspecting and populating Postgres, MySQL, SQLite | |
4. Python Dict Schema | Programmatic in-memory generation in Python scripts | |
5. dbt Project Schemas | Generating fixtures directly from | |
6. Prisma Schema | Next.js and Node.js developers seeding full-stack apps | |
7. Multi-Provider LLMs | Groq, OpenAI, Claude, Gemini, or Ollama-driven schemas | |
8. Incremental Growth | Appending rows with offset IDs and preserved FKs |
🌐 20+ Built-in Industry Domains
Generate domain-complete schemas with tuned statistical distributions out of the box:
SaaS · E-Commerce · FinTech · Healthcare · Logistics · Credit Risk · HR & People · Streaming Media · Insurance · CRM & Sales · Food Delivery · Travel & Hospitality · Gaming · Crypto & DeFi · Predictive Maintenance · Network Intrusion · Islamic Finance · EdTech · Real Estate · Contact Centers · Manufacturing SPC
See the Complete Domain Catalog.
⚡ Performance
Measured on standard Apple M-series hardware (single CPU core, no GPU):
Workload | Row Count | Generation Time | Throughput |
Single table (lognormal distribution) | 1,000,000 | 0.06 s | ~16M rows/s |
Star schema (5 tables, 4 FK dependencies) | 1,055,030 | 1.54 s | ~687k rows/s |
Multi-table enterprise database | 100,000 | 0.42 s | ~240k rows/s |
📚 Documentation Index
For in-depth guides, API references, and architecture deep dives:
Getting Started & Quickstart: 5-minute tutorial from installation to your first dataset.
Text Realism &
enrich_text: Deep dive into all 15 microtext pools and DataFrame text enrichment.AI Agent MCP Server Guide: Setup and tool reference for Cursor, Claude Code, and Windsurf.
Outcome Curves & Seasonality: Mathematical specification of revenue and growth curves.
Database Seeding in Python: Introspecting and seeding production databases.
Mimic Mode Guide: Generating privacy-safe synthetic twins from CSVs.
Export Formats: Exporting to DuckDB, Apache Parquet, Arrow IPC, and ANSI/Postgres SQL.
Apache Spark & Databricks: Scaling synthetic generation across distributed Spark clusters.
Full Schema Declarations Reference: Complete parameter reference for every column type and constraint.
📄 Research & Citation
The closed-form exact-outcome conformance engine is formalised in arXiv preprint 2606.08736:
@article{rasin2026declarative,
title = {Declarative Outcome-Conformant Synthesis: Exact, Closed-Form
Specification Satisfaction and a Conformance Benchmark},
author = {Rasin, Muhammed},
year = {2026},
url = {https://arxiv.org/abs/2606.08736v1}
}🤝 Contributing & Development Status
Misata is currently under massive, rapid development to push synthetic realism to its absolute limit: expanding real-world domain knowledge, deepening seed pool vocabularies, elevating textual column realism, and advancing statistical fidelity across every industry vertical.
Contributions from domain experts, data engineers, and researchers are warmly welcomed! Whether you want to:
Enrich Text & Vocabulary Pools: Add authentic seeds and grammar rules for specialized domains in
misata/vocab_seeds.pyandmisata/microtext.py.Contribute a Domain Capsule: Expand built-in industry templates (healthcare, legal, banking, engineering, supply chain).
Advance Statistical Fidelity: Improve multi-variate copulas, time-series dynamics, or outcome-curve solvers.
Report Edge Cases & Realism Flaws: Open an issue or discussion whenever generated values don't look 100% human-authentic.
git clone https://github.com/rasinmuhammed/misata
cd Misata
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest -qMisata is open-source under the MIT License.
Available Tools
11 toolsaudit_datasetAudit a dataset for reader-visible contradictionsARead-onlyIdempotent
Score a folder of CSVs for the contradictions a human reader would catch.
This is Misata's coherence audit run on data that already exists — data an agent generated in an earlier step, data a user built by hand, or the output of some other tool. It checks, among other things:
timestamps that run backwards (shipped before ordered, resolved before opened),
derived columns that do not reconcile with their inputs (
total!=quantity * unit_price),geographic fields that disagree (city / state / postcode / country),
near-constant columns (98% one value — a distribution tell),
filler text and out-of-scale numerics.
A score of 100 is clean. Below ~85 usually means the schema is missing
realism structure, not that individual rows need patching: add
__correlations__, profiles, time_series, a __state_machine__,
or an __outcome_curves__ declaration and regenerate.
Args:
dataset_dir: Directory containing one CSV per table (e.g. the
output_dir returned by generate_from_schema).
top_findings: Max findings to include in the response (default 20).
Returns:
{"score": 0-100, "clean": bool, "counts": {...}, "findings": [...]}.
| Name | Required | Description | Default |
|---|---|---|---|
| dataset_dir | Yes | ||
| top_findings | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnly, openWorld, idempotent, non-destructive) already cover the safety profileestr; the description adds real behavioral context on top: what kinds of contradictions it detects, that scores below ~85 indicate schema issues rather than data issues, and the recommended remediation path. It also specifies the top_findings cap and the return shape. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but each section earns its place: the opening sentence states the operation, the bulleted list gives the agent concrete detection categories, the score-interpretation paragraph is actionable, and the parameter/return details are compact. The multi-line structure with bullets and monospace examples is readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains what this audit is for, when it is meant to be run (on already-generated data), how to interpret the 0-100 score, and what to do next (add realism structures). It avoids duplicating the output schema but explains the score semantics and the folder format. Minor gap: it does not explicitly say that a clean score means no action is needed, but the score interpretation and remediation advice cover the essential behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is effectively 0% for both parameters (one has only a type and a default, the other only a type), so the description carries the burden. It adequately explains 'dataset_dir' (one CSV per table, the output_dir of generate_from_schema) and 'top_findings' (max findings, default 20). That covers both parameters even though the schema itself has descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Score a folder of CSVs' with a specific verb and resource. It explicitly enumerates the types of contradictions detected (timestamps, reconciliation, geography, near-constant columns, filler text), which gives the agent a precise picture of what this tool does and distinguishes it from generation/schema tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool: on data that already exists, whether agent-generated or hand-built. It also tells the agent what the findings mean and how to respond — add structure declarations instead of patching data — which acts as an implicit exclusion for fix-data workflows. Sibling context (generate_from_schema) reinforces the routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_sandboxCreate an instant sandbox databaseAIdempotent
Create and seed a real, isolated SQLite sandbox database for testing.
Perfect for AI coding agents: spins up a queryable database with realistic, referentially intact tables (0 orphan foreign keys) so you can test SQL queries, analytics logic, or application code immediately without touching a real database.
Args: domain: Pre-built domain template ('saas', 'ecommerce', 'fintech', 'healthcare'). story: Plain-English description (e.g. 'A restaurant with tables, bookings, and guests'). schema: Explicit schema dict if you want exact column control. rows: Base row count (default: 100). Transactions scale proportionally. seed: Random seed for 100% reproducible results (default: 42). db_path: Target SQLite file path (default: '.misata/sandbox.db').
Returns: Dictionary with db_url, db_path, table summaries, DDL, and sample verification queries.
| Name | Required | Description | Default |
|---|---|---|---|
| rows | No | ||
| seed | No | ||
| story | No | ||
| domain | No | ||
| schema | No | ||
| db_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, openWorldHint=false, and idempotentHint=true, so safety and repeatability are partly covered. The description usefully adds that seeded data has '0 orphan foreign keys' (referential integrity) and that the result is reproducible via a default seed, and discloses the default write path '.misata/sandbox.db'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose followed by Args and Returns sections, which is well organized and every element is informative. There is mild promotional padding ('Perfect for AI coding agents') and the narrative benefit paragraph could be trimmed, but nothing is truly wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter creation tool with annotations covering the safety profile and an output schema covering return values, the description supplies purpose, all parameter meanings, and a return summary. The remaining gap is guidance on how it differs from the many sibling generation/seed tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the schema provides only titles and defaults. The description compensates fully: it names the four built-in domains, explains 'story' as plain English, clarifies 'schema' as an explicit dict for column control, notes rows scaling, and explains seed reproducibility (default 42) and db_path default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create and seed a ... SQLite sandbox database') and adds the distinguishing traits of being real, isolated, and for testing. It does not, however, differentiate itself from close siblings like seed_database, generate_dataset, or create_from_schema, so an agent cannot route confidently on description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear usage context: 'test SQL queries, analytics logic, or application code immediately without touching a real database,' which tells the agent when this is the right tool. There are no explicit exclusions or named alternatives despite several overlapping siblings (seed_database, generate_dataset, generate_from_schema).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_datasetGenerate a dataset from a sentenceA
Generate a synthetic dataset from a story and write it to disk as CSV files.
Returns the output directory, file paths, row counts per table, and a small sample of rows for each table so the agent can show the user what was produced without loading every row into context.
Args: story: Plain-English description of the dataset. rows: Default row count for the primary table. seed: Optional random seed (same seed → byte-identical output). output_dir: Where to write CSVs. Defaults to a fresh temp dir. sample_rows: Number of rows from each table to include in the response (max 50).
| Name | Required | Description | Default |
|---|---|---|---|
| rows | No | ||
| seed | No | ||
| story | Yes | ||
| output_dir | No | ||
| sample_rows | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavioral context beyond the annotations. It states that the tool writes CSV files to disk, returns output directory and file paths, and includes row counts and sample rows. It also discloses determinism: 'same seed → byte-identical output.' These details are not present in the annotations (readOnlyHint false, destructiveHint false) and valuable for the agent to anticipate side effects. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured. The opening sentence states purpose, followed by a clear list of return values, then a bulleted Args section. Every sentence earns its place, and the use of a list for parameters improves readability. There is no redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 5 parameters, an output schema, and complex behavior (file writing, determinism). The description explicitly mentions return values (output directory, file paths, row counts, sample rows), addresses parameter defaults, and explains the sample_rows cap. It is complete for an agent to use effectively without further documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has no property descriptions (coverage 0%), so the description carries the full burden. It includes an Args section that explains each of the 5 parameters: story, rows, seed, output_dir, and sample_rows, adding meaning like 'Defaults to a fresh temp dir' and 'max 50'. This fully compensates for the schema's lack of descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: 'Generate a synthetic dataset from a story and write it to disk as CSV files.' It specifies the resource (dataset from a story) and distinguishes it from siblings like generate_from_schema, which focuses on schema-driven generation. The verb 'generate' and the context 'from a story' are specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: when you have a plain-English story to convert into a dataset. It also explains that it returns sample rows and file paths, which is helpful for the agent to show results. However, it does not explicitly mention alternatives (e.g., 'use generate_from_schema if you have a schema') or state when NOT to use it, leaving a slight gap in decision-making guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_from_schemaGenerate a dataset from a schemaAIdempotent
Generate a dataset from a schema you design. This is the primary Misata tool.
DIVISION OF LABOUR You (the agent) design what the data should look like — the tables, columns, business rules, and declared targets. Misata handles the hard guarantees: every FK resolves, every rollup reconciles to the cent, every declared outcome curve is hit exactly, every seed produces byte-identical output. The response includes a per-relationship integrity proof.
SCHEMA FORMAT {table_name: {column_name: spec, ...}, ...}
TABLE-LEVEL KEYS (inside a table dict) "rows": 5000 Per-table row count; overrides the global rows arg. Always set this — one global count rarely fits all tables. "constraints" List of row-level business rules (see CONSTRAINTS below). "correlations" List of pairwise Pearson targets (see CORRELATIONS below). "state_machine" Markov terminal-state assignment (see STATE MACHINE below).
COLUMN TYPES integer, float, decimal Numeric. string Short categorical or free text. text Long free text (descriptions, notes). email, phone, url, uuid Semantic strings; always valid format. date, datetime Temporal; realistic granularity applied automatically. boolean True/False with declared probability.
COLUMN SPEC KEYS (inside a column dict) primary_key: true Auto-incremented PK; column excluded from CSV output. foreign_key: {table, column} Child FK; referential integrity guaranteed + verified. min / max Numeric or date bounds. decimals: 2 Decimal places for float output. unique: true All values in the column are distinct. nullable: true Allow nulls (default true). enum: [...] Categorical choices. Add probabilities: [...] for weights; omit for realistic Zipf-shaped rank frequencies. probabilities: [...] Weights for enum choices; must sum to 1.0.
DISTRIBUTIONS (float / integer columns) distribution: normal Also: lognormal, uniform, exponential, beta, poisson, power_law, gamma. mean / std Normal params. Can be a scalar OR a per-row parent-entity lookup: mean: {formula: "@patients.hba1c_baseline"} The FK is resolved per-row so each child row's distribution is anchored to its parent's value. Use this for longitudinal data where within-entity variation should be modelled separately from between-entity variation. mu / sigma Lognormal params (mu and sigma are of log(x)). min / max Hard clamps applied after sampling. Use lognormal for money, file sizes, session durations — anything right-skewed and strictly positive. Use normal for measurements.
DERIVED COLUMNS formula: "quantity * unit_price" Row-level arithmetic; pandas eval syntax. formula: "hours * @employees.rate" Cross-table via FK: @parent_table.column. rollup: {from_table, fk, agg, column} Parent column that EXACTLY reconciles with child rows under JOIN. agg: sum/count/mean/ max/min. Add where: {col: val} to filter. RULE: use rollup (not formula) for any parent column that summarises child rows. Rollups are closed-form exact; formulas cannot cross the FK boundary correctly.
CODE-STYLE STRINGS pattern: "SKU-\d{5}" Single pattern expanded per row. pattern: ["A/\d{5}", "\d{6}"] List: one shape drawn per row. pattern_weights: [0.7, 0.3] Weights for pattern list (optional). Supported tokens: \d (digit), [A-Z] (uppercase letter), [a-z] (lowercase), literal chars, {n} repeat count. Example: "[A-Z]{2}-\d{4}" → "AB-3721".
TEXT SEMANTICS text_type: person_name Always beats column-name inference. Options: person_name, email, company, city, country, postal_code, phone, url, description, username, product_name, review_text, address, job_title. Dates: appointment times snap to 15-min business-hours grids; signups follow waking-hour rhythms; machine events keep sub-second precision. Names, genders, and emails are generated jointly and always agree.
STRATIFIED DISTRIBUTIONS (profiles) Use when different subgroups need different distributions for the same column. profiles: [ {when: "arm == 'placebo'", distribution: normal, mean: -0.35, std: 0.50}, {when: "arm == 'high_dose'", distribution: normal, mean: -1.25, std: 0.55}, ] Rows that match no profile get the column's top-level distribution. The when expression is a pandas eval string; reference any already-generated column in the same table. Always list profiles after the columns they reference.
INFORMATIVE MISSINGNESS (MAR) null_when: "dropout == False" Null this column when expression is true. missing_if: Missing-At-Random tied to a predictor column. predictor: hba1c_baseline relationship: higher_increases_probability # or lower_increases_probability base_rate: 0.05 # null probability at predictor median max_rate: 0.40 # null probability at predictor extreme Use null_when for status-conditional nulls (dropout_visit is null when not dropped out). Use missing_if when missingness is correlated with an observed variable.
EXACT INCIDENCE CONTROL exact_incidence: Hit the declared count exactly (not approximately). mode: exact rate: 0.22 # exactly floor(n * 0.22) rows become True group_by: arm # optional: apply per group rates: {placebo: 0.15, high_dose: 0.55} # per-group exact rates Use exact_incidence instead of probability on boolean columns when the user states a precise rate that must hold in the data, not just on average.
WITHIN-ENTITY TIME SERIES (longitudinal autocorrelation) time_series: Re-writes a column to have AR1 autocorrelation entity_id: patient_id within each entity group. order_by: visit_number model: AR1 # AR1 | linear_trend | random_walk | mean_reversion phi: 0.72 # autocorrelation coefficient (AR1 only) noise_std: 0.30 anchor_column: hba1c_baseline # starting value (column in the same table) trend: slope_mean: -0.08 # mean drift per step slope_std: 0.02 # per-entity slope variability Required for any longitudinal dataset (clinical visits, IoT sensors, user sessions). Without it every row is independent and the data fails any time-series test.
CONSTRAINTS (table-level constraints list) {"type": "inequality", "column_a": "visit_date", "operator": ">=", "column_b": "enroll_date", "action": "cap"} Enforces column_a OP column_b. action: "cap" (snap column_a to column_b) or "drop" (remove violating rows). Works on dates and numerics. {"type": "col_range", "low_column": "min_price", "column": "price", "high_column": "max_price", "action": "cap"} Keeps low_column <= column <= high_column. {"type": "max_per_group", "group_by": "user_id", "max_count": 3} Limits rows per group value. {"type": "unique_combination", "columns": ["user_id", "product_id"]} No duplicate (col_a, col_b) pairs. Use constraints for any business rule that must hold on every row: visit_date >= enrollment_date, price > cost, resolution_day > onset_day.
CORRELATIONS (table-level correlations list) [{"col_a": "bmi", "col_b": "systolic_bp", "r": 0.41}] Enforced via Iman-Conover (rank reordering): preserves each column's marginal distribution while hitting the declared Pearson r exactly. Declare correlations for any pair of measurements that co-vary in the real domain (bmi/bp, income/spending, tenure/salary). Also supports full matrix syntax: correlations: matrix: columns: [hba1c, glucose, bmi] values: hba1c: [1.00, 0.65, 0.28] glucose: [0.65, 1.00, 0.22] bmi: [0.28, 0.22, 1.00]
ICC CLUSTER EFFECTS (parent table cluster_effect) cluster_effect: affects_table: visits affects_columns: hba1c: icc: 0.18 # intraclass correlation coefficient sd_total: 1.5 # total standard deviation; sd_between = sqrt(icc)*sd_total systolic_bp: sd_between: 8.0 # supply sd_between directly if preferred Applies per-parent-entity random intercepts to the named child columns. Required for multi-site or multi-centre designs — without it all sites look identical and any ICC statistical test will detect the synthetic origin. icc: 0.10-0.30 is typical for clinical measurements across sites.
STATE MACHINE (table-level state_machine) state_machine: state_column: patient_status initial_state: enrolled transitions: enrolled: {on_treatment: 0.97, screen_failure: 0.03} on_treatment: {completed: 0.77, dropout: 0.23} Assigns one terminal state to every row by following the Markov chain. States with no outgoing transitions are terminal. Use for any process with defined states: clinical trial statuses, customer lifecycle, order fulfilment stages, support ticket resolution.
SCHEMA-LEVEL DIRECTIVES (top-level keys, siblings of the tables)
outcome_curves Declare aggregate targets the engine hits EXACTLY. [{"table": "orders", "column": "amount", "time_column": "order_date", "time_unit": "month", "value_mode": "absolute", "start_date": "2024-01-01", "avg_transaction_value": 120.0, "curve_points": [ {"month": 1, "target_value": 50000.0}, {"month": 6, "target_value": 110000.0}, {"month": 12, "target_value": 200000.0} ]}] ALWAYS use this when the user states what a number should sum to per period: "revenue grows from $50k to $200k", "Q4 spike", "10x growth". avg_transaction_value drives row count per period; set it to roughly the median row value for that column.
rate_curves Per-period rate targets for boolean/categorical columns. [{"table": "transactions", "column": "is_fraud", "time_column": "transaction_date", "rate_points": [ {"period": "2024-01", "rate": 0.02}, {"period": "2024-Q4", "rate": 0.05} ]}] Use when fraud rate, churn rate, or conversion rate changes over time.
group_shares Exact shares of a measure across a categorical column. [{"table": "orders", "measure": "amount", "group_column": "plan", "shares": {"Starter": 0.2, "Pro": 0.5, "Enterprise": 0.3}}] The measure sums to those proportions per group, exactly. Paired with an outcome_curves on the same table+measure, the split holds inside every declared period. Use for any "A is 40% of revenue, B is 35%..." statement.
waterfalls Movements that reconcile to declared running balances. [{"table": "mrr_movements", "starting_value": 100000, "points": [{"period": "2026-01", "ending_value": 106000}, ...], "inflow_shares": {"new": 0.7, "expansion": 0.3}, "outflow_shares": {"churn": 1.0}}] Use for MRR bridges, cash-flow statements, any "opening + inflows - outflows = closing" ledger that has to tie out.
stock_flows Per-unit inventory identity: closing = opening + received
shipped, enforced for every SKU across every period. Use for warehouse / inventory data where stock levels must be internally consistent.
lifecycles A state machine with legal transitions (stricter than a table-level state_machine: illegal jumps are impossible, not just unlikely). [{"name": "order_flow", "table": "orders", "state_column": "status", "start_column": "placed_at", "initial": "placed", "states": [{"name": "placed"}, {"name": "paid"}, {"name": "shipped"}, {"name": "delivered"}, {"name": "refunded", "terminal": true}], "transitions": [["placed","paid"],["paid","shipped"], ["shipped","delivered"],["delivered","refunded"]]}]
missingness Why a value is missing, conditionally (schema-level MNAR). [{"table": "contacts", "column": "notes", "rate": 0.75, "else_rate": 0.05, "when_column": "is_active", "when_op": "==", "when_value": false}] The null rate is exact per branch. Use when missingness itself carries signal a cleaning step should be tested against.
DIRTY DATA ON PURPOSE (exact counts, so a test has a known number to find) duplicates [{"table": "contacts", "count": 60}] typos [{"table": "contacts", "column": "city", "count": 120}] outliers [{"table": "orders", "column": "amount", "count": 40}] Each injects exactly that many defects, leaving primary/unique/foreign keys intact. Use when the user is building or testing a data-quality or cleaning pipeline.
domain Domain hint for post-generation validation.
"domain": "clinical_trial" # or "clinical", "financial", "fintech"
When set, generate_from_schema runs domain validation automatically and
returns it as domain_validation in the response — no second call
needed. Built-in ranges: HbA1c 4-14 %, BMI 10-80, systolic BP 60-260,
age 0-130, glucose 2-40, cholesterol 1-20, hemoglobin 3-25 for clinical;
price >= 0, discount 0-1, rate -1 to 100 for financial.
There are ~24 schema-level declarations in total. The ones above cover the common cases; for the rest (retention cohorts, DAG edges, closure tables, graph motifs, time grids, bitemporal history) fetch https://misata.studio/docs/reference/declarations.md and check the exact key before inventing one.
DESIGN RULES — follow these to get the best result in one pass
Always set rows per table. A fintech schema with customers=2000, accounts=4000, transactions=50000 is far better than 1000 everywhere.
Every child table needs a FK column pointing to its parent PK. Without it orphan rows are generated and the integrity proof will fail.
Use lognormal for money, file sizes, response times (right-skewed, strictly positive). Use normal for measurements (height, score, temp).
Declare correlations for any pair that co-varies in the real domain. Generated data with an identity correlation matrix is the clearest synthetic-data tell there is.
Use exact_incidence instead of probability when the user states a precise rate. "3% fraud" with probability: 0.03 gives ~3% on average; exact_incidence gives exactly 3%.
Use rollup (not formula) for any parent column that must reconcile with child rows. customers.total_spent generated independently of orders will never match; a rollup makes it exact.
Use outcome_curves any time the user mentions a revenue shape, a growth trajectory, a seasonal pattern, or a specific period total. It is the single feature most likely to be forgotten and most visible when absent.
Use profiles when two groups need different distributions. A clinical trial where all arms share one HbA1c distribution is statistically wrong and the difference will be caught by any summary table.
For longitudinal data (visits, sessions, sensor readings), add time_series to the key measurement columns. Independent rows fail every autocorrelation test and are visually obvious when plotted.
Add state_machine to any entity that moves through a process. An order table with no status progression is not realistic order data.
Args: schema: Dict of table defs plus optional schema-level directives. rows: Default row count for tables without rows. seed: Random seed (same seed → byte-identical output on any machine). output_dir: Where to write CSVs. Omit to use a fresh temp dir. sample_rows: Rows per table to include in the JSON response (max 50).
| Name | Required | Description | Default |
|---|---|---|---|
| rows | No | ||
| seed | No | ||
| schema | Yes | ||
| output_dir | No | ||
| sample_rows | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral context beyond annotations: same seed produces byte-identical output, CSV output goes to output_dir or a temp dir, domain validation may run automatically, and the response includes an integrity proof. It does not contradict the annotations and substantially helps an agent anticipate side effects and outputs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very long, but the length is largely justified by the complexity of the schema DSL it defines. It is organized with clear section headers, front-loads the primary purpose, and includes practical design rules, making the length readable rather than a hazard.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is effectively a spec: it covers the schema structure, column types, constraints, correlations, state machines, time-series handling, top-level directives, and it even tells the agent where to fetch documentation for unsupported declarations. Given the tool's complexity and the otherwise empty schema descriptions, this is complete enough to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the entire burden of parameter meaning. It documents every argument in the Args section and uses a very detailed, nested DSL to explain the schema parameter—covers row counts, constraints, distributions, rollups, outcome curves, duplicates, and more.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb-and-resource statement: 'Generate a dataset from a schema you design' and reinforces it as 'the primary Misata tool.' However, it does not explicitly disambiguate from the sibling generate_dataset tool, so it is clear but not fully sibling-differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is abundant design guidance and feature-selection advice within the description, such as when to prefer rollup over formula and when to use exact_incidence. However, it does not explicitly say when to use this tool instead of generate_dataset or the validation siblings, so usage vs. alternatives is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_schemaInspect the schema behind a storyARead-only
Return the full schema (tables, columns, relationships) for a story without generating data.
Heavier than preview_story — includes every column with its type and
distribution params. Use when the user wants to see the structure they'll
get, or to author a misata.yaml file from a natural-language seed.
Args: story: Plain-English description of the dataset. rows: Default row count for the primary table.
| Name | Required | Description | Default |
|---|---|---|---|
| rows | No | ||
| story | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as read-only and non-destructive; the description builds on this by noting it does not generate data and by explaining the type of output (every column with type and distribution params). It stops short of detailing performance implications beyond 'heavier' or clarifying idempotent behavior, but adds meaningful context beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly written, starting with a clear one-sentence purpose, followed by a comparative detail, use cases, and a compact Args list. Every sentence earns its place with no redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only inspection tool with two parameters and a documented output schema, the description covers the what, when, and parameters, and even notes the lack of data generation. It could go further by describing error scenarios or output format specifics, but these are already captured by the output schema and annotations, making it adequately complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate, and it does via an Args section that explains `story` as 'Plain-English description of the dataset' and `rows` as 'Default row count for the primary table.' This adds clear meaning beyond the bare types, though it could elaborate on how changing `rows` affects its behavior. The high coverage shift justifies a score above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Return the full schema (tables, columns, relationships) for a story without generating data,' using a specific verb and resource scope. It goes beyond a generic statement by explicitly contrasting with `preview_story` ('Heavier than') and providing concrete use cases, making its purpose unmistakable and distinctive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly notes this tool is 'Heavier than `preview_story`' and provides two direct use cases: showing the structure the user will get and authoring a `misata.yaml` file from a seed. This gives an agent clear when-to-use guidance and signals an alternative, even if it doesn't mention explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_domainsList built-in domainsARead-onlyIdempotent
List the 18 built-in business domains Misata can generate from natural language.
Each domain has trigger keywords and a sample story you can pass to
preview_story or generate_dataset. Use this when the user asks
"what kinds of data can you generate?" or to suggest a story format.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds useful context beyond annotations by revealing that the result contains 18 domains, each with trigger keywords and a sample story, which the agent can pass to preview_story or generate_dataset.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core action in the first sentence. The subsequent sentences add useful behavioral and usage context without unnecessary fluff, though it could have been slightly tighter by merging the first two sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is sufficient for a parameterless list tool: it names the exact output scope, explains what each returned domain contains, gives intended usage triggers, and references downstream consumer tools. The output schema and annotations cover the remaining structural and safety aspects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, and the schema coverage is effectively 100% because there is no input to document. There is nothing for the description to add about parameter syntax or meaning, so the baseline score for a zero-parameter tool applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb+resource: 'List the 18 built-in business domains Misata can generate from natural language.' This clearly distinguishes the tool from the generative siblings by indicating it is a read-only catalog of available domains.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use this tool: 'Use this when the user asks "what kinds of data can you generate?" or to suggest a story format.' It also mentions downstream tools like preview_story or generate_dataset, but it does not explicitly address when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_storyPreview how a story is interpretedARead-only
Inspect what Misata would generate from a story — without generating any rows.
Returns the detected domain, confidence, near-misses, locale, scale, and a preview of the tables that would be produced. Use this to confirm interpretation before committing to a (potentially large) generation.
Args: story: Plain-English description of the dataset. rows: Default row count for the primary table (affects preview only).
| Name | Required | Description | Default |
|---|---|---|---|
| rows | No | ||
| story | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds context: it clarifies no rows are generated, that rows affects preview only, and lists returned fields (domain, confidence, near-misses, locale, scale, tables). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: a one-sentence purpose, a list of returned items, a usage note, and an Args section. It is concise with no redundant text and front-loads the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 params, output schema exists), the description covers purpose, usage, parameters, and return values. It provides sufficient context for an agent to understand when and how to invoke it, without needing to rely on the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must clarify parameters. It provides Args: 'story: Plain-English description of the dataset' and 'rows: Default row count for the primary table (affects preview only).' This fully compensates for missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Inspect what Misata would generate from a story — without generating any rows.' This uses a specific verb (inspect) and resource (story interpretation) and differentiates from sibling generation tools like generate_dataset and seed_database.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use: 'Use this to confirm interpretation before committing to a (potentially large) generation.' It implies the alternative is to generate, but does not name the exact sibling tools. Still, the guidance is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_sandboxQuery the sandbox databaseARead-onlyIdempotent
Execute a read-only SQL query against the sandbox database.
Allows AI coding agents to verify SQL queries, test aggregations, or inspect data in the sandbox directly without opening external terminals.
Args: query: SQL SELECT query to execute. db_path: Optional path to the database (defaults to '.misata/sandbox.db'). limit: Max rows to return (default: 100).
Returns: Dict with columns, rows (as list of dicts), row_count, and execution status.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| db_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false, and destructiveHint=false, and the description's 'read-only' claim is consistent with them. It adds value beyond annotations by disclosing the default database location and the default row cap, plus the fact that no external terminal is needed. It does not say what happens if a non-SELECT statement is submitted, which is the remaining behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then a rationale sentence, then clearly labeled Args and Returns sections. Every element is relevant, though the rationale sentence is slightly verbose and the Returns block partly duplicates the output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-param, read-only sandbox query tool with annotations and a declared output schema, the definition covers purpose, use cases, all parameters, defaults, and return shape — so an agent can call it correctly. The only omission is error/enforcement behavior for invalid or non-read-only SQL, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and largely does so: it defines query as the SQL SELECT to execute, db_path as an optional path defaulting to '.misata/sandbox.db', and limit as the max rows with a default of 100. This covers all three parameters with meaning and defaults, though it omits any syntax note such as whether multi-statement or mutating SQL is rejected.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource — 'Execute a read-only SQL query against the sandbox database' — plus the scope constraint (read-only, sandbox). This is immediately distinguishable from siblings like inspect_schema, seed_database, or create_sandbox.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete use cases ('verify SQL queries, test aggregations, or inspect data in the sandbox directly without opening external terminals'), which tells the agent when this tool is the right pick over running an external shell. It stops short of naming alternatives (e.g., inspect_schema for structure) or stating exclusions, so it is clear context without explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
seed_databaseSeed a real databaseADestructive
Fill a live Postgres or SQLite database with realistic, connected data, read from the database's own schema.
Reads the tables, columns, and foreign keys directly from the target database, generates data that respects them, inserts parents before children, then queries the database back to confirm every foreign key resolves. No schema file and no ORM are needed: a connection string is enough.
SAFETY — this is the only Misata tool that writes to a user's database:
It plans by default. With
apply=False(the default) nothing is written; you get the table list, insert order, existing row counts, and what would be inserted. Show that plan to the user.Only call again with
apply=Trueafter the user has seen the plan and agreed. Never passapply=Trueon a first call.If any target table already has rows, the write is refused unless the user chooses
truncate=True(wipe and reseed) orappend=True(keep existing rows, seed only empty tables, and draw foreign keys from the rows already there). Never guess between these.truncate=TrueDESTROYS existing data. Only use it on a throwaway development database and only when the user explicitly asks.
Args:
db_url: Connection string, e.g. postgresql://localhost/myapp_dev
or sqlite:///dev.db.
rows: Base row count; reference and transaction tables scale from it.
apply: False (default) plans only. True performs the write.
truncate: Wipe target tables (children first) before seeding.
append: Keep populated tables and seed only the empty ones.
tables: Optional allow-list of table names to seed.
skip_tables: Tables to leave untouched (migrations, auth, etc.).
seed: Random seed; the same seed reproduces the same data.
Returns:
A plan (applied: false) or a result with per-table row counts and
a per-relationship integrity proof (integrity.verified).
| Name | Required | Description | Default |
|---|---|---|---|
| rows | No | ||
| seed | No | ||
| apply | No | ||
| append | No | ||
| db_url | Yes | ||
| tables | No | ||
| truncate | No | ||
| skip_tables | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the annotations by detailing the destructive nature (truncate=True destroys data), the write behavior, and the safety flow (plan first, then apply). It transparently discloses side effects and the conditions under which writes are allowed, adding crucial context that annotations alone do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Although lengthy, the description is well-structured with clear sections (summary, safety, args) and every sentence adds value. It avoids redundancy and uses formatting (bold, bullet points) to improve readability, making it efficient despite its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 8 parameters and output schema, the description thoroughly covers all aspects: parameter usage, safety considerations, and the return value (plan vs result with integrity verification). It is complete and self-contained, leaving no critical gaps for an agent to misuse the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, but the description's Args section explains every parameter in detail, including examples, defaults, and how they interact (e.g., rows scaling, tables/skip_tables filtering, apply vs truncate/append). This fully compensates for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: filling a live Postgres or SQLite database with realistic, connected data based on the database's own schema. It distinguishes itself from sibling tools like generate_from_schema by emphasizing that no schema file or ORM is needed, and it is the only tool that writes to a database.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains when to use this tool (to populate a database) and differentiates it from alternatives by noting it is the only one that writes. It also provides clear instructions on the safe default (apply=False) and when to use truncate/append, making it obvious when to call it versus other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_domainValidate a dataset against domain rulesARead-onlyIdempotent
Flag values that are physiologically or financially impossible for a domain.
Where audit_dataset checks internal consistency, this checks values
against what the outside world allows. Built-in ranges include, for
clinical / clinical_trial: HbA1c 4-14%, BMI 10-80, systolic BP
60-260, age 0-130, glucose 2-40, cholesterol 1-20, hemoglobin 3-25; for
financial / fintech: price >= 0, discount 0-1, rate -1 to 100.
Use after generating a dataset in a regulated or measurement-heavy domain, or when the user asks "is this data plausible for a real clinic / bank?".
Args:
dataset_dir: Directory containing one CSV per table.
domain: One of clinical_trial, clinical, financial,
fintech.
Returns:
{"passed": bool, "errors": [...], "warnings": [...]}. passed
is True when there are no ERROR-level findings.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | ||
| dataset_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, covering safety. The description adds significant behavioral context beyond that: the exact output structure (passed, errors, warnings), the meaning of 'passed', and the built-in domain ranges. It also clarifies that it does not check internal consistency, preventing misinterpretation. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: a one-sentence summary, a clarifying contrast, domain-specific ranges, usage guidance, and a clean Args/Returns section. It is front-loaded with the core purpose and every sentence serves a distinct function. No fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema, the description still details the return format and the logic for 'passed'. It covers parameter semantics, domain values, file structure, and usage timing. For a tool with only two parameters and a well-defined output schema, this is complete and leaves no ambiguity for the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden. It explains dataset_dir as 'Directory containing one CSV per table' and domain as one of four enumerated values. This fully compensates for the schema's lack of descriptions, giving the agent everything needed to populate both parameters correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Flag values that are physiologically or financially impossible for a domain.' It immediately contrasts with the sibling audit_dataset, distinguishing its purpose clearly. The specific built-in ranges reinforce the scope, making the tool's function unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use: 'after generating a dataset in a regulated or measurement-heavy domain' and when the user asks about plausibility. It also names the alternative audit_dataset and differentiates by checking 'internal consistency' vs 'outside world' values. This gives the agent clear routing criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_yamlValidate a misata.yamlARead-onlyIdempotent
Validate a misata.yaml document at two levels.
Runs both checks in sequence:
Structural — the published JSON Schema (correct field types, required fields, enum values). Catches typos and shape errors.
Semantic —
misata.validate_schema(probabilities sum to 1.0, every foreign_key has a matching Relationship, no cycles, outcome curves reference real columns, etc.). These are the rules that would crash generation; the error messages include suggested fixes.
Use this when an agent has authored or edited a misata.yaml on the
user's behalf and wants to confirm it parses and will actually
generate before invoking generate_dataset.
Args: yaml_text: The full contents of a misata.yaml file as a string.
Returns:
{"valid": true} if both checks pass; otherwise
{"valid": false, "errors": [...], "stage": "structural"|"semantic"}
with the layer that failed first.
| Name | Required | Description | Default |
|---|---|---|---|
| yaml_text | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description significantly enriches the annotations (readOnlyHint, idempotentHint) by detailing the two-step execution (structural then semantic), giving examples of semantic rules, and specifying the exact return format including the 'stage' of failure. It doesn't mention any rate limits or auth needs, hence not a 5, but it provides strong context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely well-structured and concise for the complexity it handles. It uses markdown headers for the two validation levels, a 'Use this when' section for guidance, and clearly formatted Args and Returns blocks. Every sentence provides valuable information without unnecessary verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema and detailed annotations, the description perfectly complements them. It explains the two types of validation, the order of execution, and what happens on failure. This is complete for an agent to understand the tool's behavior, when to use it, and what to expect in return.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 0% (no 'description' field in the schema for the parameter), the description explicitly documents the 'yaml_text' parameter: 'The full contents of a misata.yaml file as a string.' This fully compensates for the schema's lack of description, so while it increases transparency, the schema handles the variable type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool validates a 'misata.yaml' document at two distinct levels (structural and semantic). It uses a specific verb+resource pattern and the detailed explanation of the two checks distinguishes it from sibling tools like 'validate' on other resources or 'inspect_schema'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use the tool: 'Use this when an agent has authored or edited a misata.yaml on the user's behalf...'. It also clearly defines the alternative by mentioning 'before invoking generate_dataset', which is a sibling tool, and contrasts parsing with semantic validation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.11- Added
create_sandbox - Added
query_sandbox
2 tool updates
v0.1.9- Added
audit_dataset - Added
validate_domain
7 tool updates
v0.1.0- First observed
generate_dataset - First observed
generate_from_schema - First observed
inspect_schema - First observed
list_domains - First observed
preview_story - First observed
seed_database - First observed
validate_yaml
TDQS
Scored across 11 tools
Most tools have clearly distinct purposes: preview_story vs inspect_schema differ in depth, generate_dataset (story-driven) vs generate_from_schema (schema-driven) are separated by input, and audit_dataset (internal consistency) vs validate_domain (external plausibility ranges) are explicitly contrasted. The only mild overlap is preview_story and inspect_schema, which both take a story and surface structure, but the descriptions clarify the lighter-vs-heavier distinction well.
Every tool follows a clean snake_case verb_noun pattern: preview_story, validate_yaml, create_sandbox, query_sandbox, list_domains, inspect_schema, generate_dataset, generate_from_schema, seed_database, validate_domain, audit_dataset. No camelCase mixing or vague verbs, and the naming maps predictably to resource plus action.
11 tools is well-scoped for a synthetic-data platform spanning preview, generation, validation, sandboxing, and DB seeding. Each tool earns its place and there is no redundancy or filler.
The surface covers the full lifecycle: discovery (list_domains), preview (preview_story/inspect_schema), authoring validation (validate_yaml), generation (generate_dataset/generate_from_schema), post-hoc QA (audit_dataset/validate_domain), and live deployment (create_sandbox/query_sandbox/seed_database). No obvious gaps or dead ends for the stated domain.
Maintenance
Related MCP Connectors
Generate realistic, FK-consistent synthetic test data for your databases from your AI assistant.
Generate realistic relational test data — 156 field types, 22 locales, JSON/CSV/SQL, free previews.
Ask data questions in natural language. Get SQL, insights, and charts from your databases.
Privacy-safe synthetic financial data for LatAm fintech, AI agents, testing and ML.
Related MCP Servers
- AlicenseBqualityDmaintenanceGenerates realistic mock data using Faker.js for database seeding, API testing, and development environments. Supports person/company data, custom patterns, multi-locale generation, and structured datasets with referential integrity.4233 npm7MIT
- AlicenseNot gradedqualityDmaintenanceEnables creating interactive data visualizations from natural language queries using DuckDB for local databases or Databricks for enterprise data warehouses. Supports multiple chart types, CSV imports, SQL queries, and automatic statistical analysis through Claude Desktop.19MIT
- FlicenseNot gradedqualityDmaintenanceGenerates realistic, context-aware synthetic data for AI agents to populate databases, mock APIs, and create test scenarios without exposing real PII.9 npm3-
- AlicenseAqualityDmaintenanceGenerates realistic, referentially-coherent test data (SQL INSERTs, JSON, or CSV) from your database schema, resolving foreign keys and respecting constraints. Paste CREATE TABLE DDL or a JSON schema and get ready-to-run seed data with valid relationships.250 npmMIT