Skip to main content
Glama

Misata Studio: verified synthetic data

Generate a verified dataset

generate_dataset
Idempotent
Make a verified relational dataset and wait for it: every foreign key checked, dates correctly
ordered, declared aggregates and rates exact, before anything is returned. Deterministic for a seed.

Use this when you give a `schema` or `ddl` (seconds). For a plain-English `request` the engine
designs the tables with a model, which can take several minutes: call start_generation instead and
poll get_status, so the call does not sit open and time out.

Args:
    request:  Plain-English description (needs an LLM key unless `schema`/`ddl` is also given).
    schema:   A Misata schema dict for exact structural control. No key needed for structure.
    ddl:      CREATE TABLE statements. No key needed for structure.
    seed:     Reproducibility seed — the same request/schema and seed always produce the same rows.
    research: Ground realistic numbers (prices, growth rates) in real published facts via a web
              search. Off by default (an anonymous caller's request should not trigger external
              calls unless asked for); needs a key regardless of `schema`/`ddl`.
    blueprint: The engine's full design language (blueprint_guide has the reference): readings around
              a parent's nominal, drift, autocorrelation, cause-and-effect, event windows, exact
              aggregates. Use it whenever the data must behave like the real process. No key needed;
              run validate_blueprint on it first.

Returns:
    dataset_id:   pass this to get_certificate / query_dataset / export_dataset.
    passed:       whether every check held. A dataset that did not pass is still returned, with
                  `certificate.findings` saying what failed — inspect before trusting it.
    verification: a short, plain-language account of what was checked and what held — show it to the
                  person as the proof, in place of asking them to take the data on trust.
    tables:       {name: {columns, preview_rows (first 20), total_rows}} — the preview only; every
                  row is in the stored dataset, reachable by query_dataset/export_dataset.
    certificate:  the short form (claims, requirements, findings). get_certificate returns the rest
                  (patterns, realism scorecard, every proof chart).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
ddlNo
seedNo
schemaNo
requestNo
researchNo
blueprintNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover safety (not read-only, open-world, idempotent, non-destructive), and the description adds substantial context beyond them: deterministic seeding, LLM-key requirements per argument, external web calls for research, blocking/wait behavior and timeout risk, and the important disclosure that a dataset failing verification is still returned with findings.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then routing guidance, then Args and Returns. Though long, the length is justified by six undocumented params and no output schema, and no sentence is filler — even the note to show verification 'to the person as the proof' is actionable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully specifies return fields (dataset_id, passed, verification, tables previews, certificate) and the follow-up tools to use them with. It also covers failure semantics and polling alternatives, leaving nothing an agent needs in order to call and interpret it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden, and it does: each of the six params (request, schema, ddl, seed, research, blueprint) is explained with meaning beyond its type, including key requirements, default behavior, and preconditions like validating the blueprint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource with scope: 'Make a verified relational dataset' where every foreign key, date ordering, aggregate and rate is checked before return. It explicitly distinguishes the fast path (schema/ddl, seconds) from start_generation for plain-English requests, so an agent can route without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use: call this when supplying schema or ddl, but use start_generation plus get_status polling for plain-English requests so the call doesn't time out. It also flags that research is off by default and that validate_blueprint should be run first when using a blueprint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources