Skip to main content
Glama

Misata Studio: verified synthetic data

Server Details

Hosted: plan and build multi-table test and demo data, export 25 formats, with checks shown.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

A4.2/5.0

Scored across 12 tools

Disambiguation4/5

Tools mostly target distinct operations (plan, generate, validate, query, export, certify, job-control). The one real overlap is generate_dataset vs start_generation, which share arguments and do the same work; the descriptions clearly explain the sync/async split, and get_certificate vs the inline certificate is likewise clarified as full vs short form.

Naming Consistency4/5

Nearly all names follow a clean verb_noun pattern (generate_dataset, export_dataset, get_status, validate_blueprint, cancel_generation, etc.). Minor deviations are blueprint_guide (noun_noun) and whoami, but the convention is otherwise predictable and readable.

Tool Count5/5

12 tools is well-scoped for a dataset-generation service, with each tool earning its place: planning, validation, sync/async generation, querying, exporting, certification, ready-made lookup, job control, and account info.

Completeness4/5

The surface covers the full lifecycle from planning through validation, generation, verification, querying, and export, plus ready-made dataset discovery. Minor gaps exist (no explicit list_datasets or delete/cleanup dataset), but these are workarounds rather than dead ends.

Available Tools

12 tools
blueprint_guideBlueprint referenceA
Read-onlyIdempotent
Inspect
The reference for the engine's full design language (the `blueprint` argument of generate_dataset,
start_generation and validate_blueprint): roles, distributions, formulas that read parent columns
and draw noise, per-group sequences (seq, lag, cumsum, ar1 drift, random walk), aggregates, event
windows, cause-and-effect, lifecycles, with patterns. Read it before writing a blueprint.
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, openWorldHint=false), so the burden is low. The description adds useful scope context by listing the covered topic areas and the three sibling tools whose `blueprint` argument it explains, making it clear this is a static reference rather than something with side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One long but information-dense sentence listing covered topics, followed immediately by the front-loaded imperative 'Read it before writing a blueprint.' The topic enumeration is long, though every item earns its place by telling the agent whether the construct it needs is documented.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return values need no explanation, and the description covers scope, coverage areas, and the call condition. For a zero-parameter static reference tool, nothing an agent needs is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the rubric baseline is 4. The description goes slightly beyond by identifying the `blueprint` argument it exists to document and the tools that accept it, but there are no parameters for it to clarify.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific resource — the engine's full design language — and enumerates the exact constructs it documents (roles, distributions, formulas, sequences, aggregates, windows, lifecycles). It explicitly ties itself to the sibling tools that consume the blueprint argument (generate_dataset, start_generation, validate_blueprint), so an agent can distinguish it from them without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The closing directive 'Read it before writing a blueprint' gives clear, actionable usage context and implicitly covers the when-not case (you don't need it if you aren't authoring a blueprint). It stops short of naming an explicit alternative or exclusion, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_generationCancel a generationA
Idempotent
Inspect

Stop a running start_generation job. It stops at the next stage boundary and keeps nothing.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-destructive, idempotent, non-read-only behavior, so the bar is lower; the description adds genuinely useful semantics by noting cancellation is not immediate ('stops at the next stage boundary') and that no partial output is retained ('keeps nothing'). It still omits whether the job_id remains queryable afterward and what the response contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action, with no filler. Every clause carries information about scope or timing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema and annotations covering safety, the description is nearly complete on behavior, but it never tells the agent where to obtain job_id or what happens to the job record after cancellation — the one real gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter with 0% schema description coverage, and the description never explains job_id — where it comes from (presumably start_generation) or its format. It only implicitly links the param to a 'running job', so it does not compensate for the documentation gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Stop a running start_generation job') and names the sibling tool that creates the job, so an agent can distinguish it from start_generation, get_status, and the other siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the pairing with start_generation, but there is no explicit statement of when to cancel versus checking status (get_status) or letting the job finish. No exclusions or prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_datasetExport a datasetA
Idempotent
Inspect
Export a generated dataset as a file. Returns a `download_url` the person can open (or you can
fetch, e.g. with curl) for as long as the dataset is held, about 2 hours. Give the person the link
rather than pasting file contents into the chat.

Args:
    dataset_id: from a prior generate_dataset call.
    format:     data: csv, parquet, jsonl, json, avro, xlsx, feather, orc, sqlite, duckdb, sql.
                code and docs: dbt, notebook, dictionary, dbml, mermaid, prisma, sqlalchemy,
                typescript, jsonschema, expectations, django, openapi, mockapi, demo.
                `sql` is schema.sql (DDL with keys) + data.sql (COPY/INSERT) — the way to seed a
                real database: run the returned SQL through your own database connection, since
                this server never holds a database credential itself.
    dialect:    for `sql` only: postgres, mysql, sqlite, mssql, oracle, bigquery, snowflake.
    inline:     also return the file itself as `base64` (only for files under a few MB). Use it
                when you must write the file yourself and cannot fetch a URL.

Returns:
    filename, content_type, bytes, download_url, expires_at (unix seconds), and `base64` when
    `inline` and small enough.
ParametersJSON Schema
NameRequiredDescriptionDefault
formatNocsv
inlineNo
dialectNopostgres
dataset_idYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (which only cover safety/idempotency) by disclosing the ~2 hour link lifetime, the `expires_at` field, the size ceiling on `inline`/base64, and the fact that the server never holds a database credential. These are exactly the operational traits an agent needs and that structured fields do not carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then Args, then Returns, and every sentence is load-bearing. The exhaustive format enumeration makes the block long, but since the schema supplies no enums that length is largely justified rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description documents the return payload (filename, content_type, bytes, download_url, expires_at, conditional base64), and it covers the auth/credential boundary and lifetime. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden and does so: it enumerates every `format` value grouped by data vs code/docs, explains what `sql` produces and how to consume it, restricts `dialect` to `sql` with its valid values, and scopes `inline` to small files. This fully compensates for the undocumented schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Export a generated dataset as a file') and ties the artifact back to a named sibling ('from a prior generate_dataset call'), so an agent can place it in the workflow without opening the schema. It also distinguishes the deliverable (a link) from what the agent should do with it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use direction: hand the person the link rather than pasting contents, use `inline` only when you must write the file yourself and cannot fetch a URL, and for `sql` run the returned DDL/DML through your own connection. The `sql` note effectively rules out an alternative (expecting the server to hold credentials).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_ready_datasetFind a ready-made datasetA
Read-onlyIdempotent
Inspect
The ready-made datasets Misata publishes: free sample databases (direct download, public domain)
and premium datasets with a full answer key (a free preview, then a one-off price). Check this
first when someone wants sample, demo, practice or teaching data for a common scenario (retail,
e-commerce, SaaS, manufacturing SPC, predictive maintenance, insurance claims, fraud/AML,
clinical, network security): handing over a dataset that already exists is instant. If none fits
their tables, generate one instead.

Args:
    query: what the person needs, in their words. Only orders the list (closest first); every
           dataset is still returned, so judge the fit yourself from the tables and summary.

Returns:
    datasets: [{slug, kind (free|premium), title, summary, rows, tables, page_url, and
    download_url (free) or free_preview_url + buy_url + price_usd (premium)}].
ParametersJSON Schema
NameRequiredDescriptionDefault
queryNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds real behavioral context beyond them: the query does not filter (every dataset is always returned, only ordered), free datasets are public-domain direct downloads, premium ones are preview-then-one-off-price, plus the exact returned fields including download/preview/buy URLs and price. No contradiction with readOnlyHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded (purpose, then when-to-use, then Args/Returns) and uses clear sections. Slightly padded by the long parenthetical scenario list and the 'instant' aside, but nothing is misleading or wasted enough to harm selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape (datasets array with slug, kind, title, summary, rows, tables, page_url, and the free/premium download fields), plus the non-filtering semantics. An agent has everything needed to call it and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% (no description on `query`), so the description carries the full burden and does it well: it is free-text in the user's words, it only orders results closest-first, and it does not filter — telling the agent to judge fit itself. That is meaning no schema field conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: publishes/finds ready-made datasets, split into free sample DBs and premium datasets with answer keys. It explicitly distinguishes itself from the generate path ('If none fits their tables, generate one instead'), so an agent can separate it from generate_dataset without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Check this first when someone wants sample, demo, practice or teaching data for a common scenario' gives an explicit trigger, and the fallback ('generate one instead') names the alternative action with its condition. The named scenario list further pins down when this applies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_datasetGenerate a verified datasetA
Idempotent
Inspect
Make a verified relational dataset and wait for it: every foreign key checked, dates correctly
ordered, declared aggregates and rates exact, before anything is returned. Deterministic for a seed.

Use this when you give a `schema` or `ddl` (seconds). For a plain-English `request` the engine
designs the tables with a model, which can take several minutes: call start_generation instead and
poll get_status, so the call does not sit open and time out.

Args:
    request:  Plain-English description (needs an LLM key unless `schema`/`ddl` is also given).
    schema:   A Misata schema dict for exact structural control. No key needed for structure.
    ddl:      CREATE TABLE statements. No key needed for structure.
    seed:     Reproducibility seed — the same request/schema and seed always produce the same rows.
    research: Ground realistic numbers (prices, growth rates) in real published facts via a web
              search. Off by default (an anonymous caller's request should not trigger external
              calls unless asked for); needs a key regardless of `schema`/`ddl`.
    blueprint: The engine's full design language (blueprint_guide has the reference): readings around
              a parent's nominal, drift, autocorrelation, cause-and-effect, event windows, exact
              aggregates. Use it whenever the data must behave like the real process. No key needed;
              run validate_blueprint on it first.

Returns:
    dataset_id:   pass this to get_certificate / query_dataset / export_dataset.
    passed:       whether every check held. A dataset that did not pass is still returned, with
                  `certificate.findings` saying what failed — inspect before trusting it.
    verification: a short, plain-language account of what was checked and what held — show it to the
                  person as the proof, in place of asking them to take the data on trust.
    tables:       {name: {columns, preview_rows (first 20), total_rows}} — the preview only; every
                  row is in the stored dataset, reachable by query_dataset/export_dataset.
    certificate:  the short form (claims, requirements, findings). get_certificate returns the rest
                  (patterns, realism scorecard, every proof chart).
ParametersJSON Schema
NameRequiredDescriptionDefault
ddlNo
seedNo
schemaNo
requestNo
researchNo
blueprintNo

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover safety (not read-only, open-world, idempotent, non-destructive), and the description adds substantial context beyond them: deterministic seeding, LLM-key requirements per argument, external web calls for research, blocking/wait behavior and timeout risk, and the important disclosure that a dataset failing verification is still returned with findings.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then routing guidance, then Args and Returns. Though long, the length is justified by six undocumented params and no output schema, and no sentence is filler — even the note to show verification 'to the person as the proof' is actionable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully specifies return fields (dataset_id, passed, verification, tables previews, certificate) and the follow-up tools to use them with. It also covers failure semantics and polling alternatives, leaving nothing an agent needs in order to call and interpret it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden, and it does: each of the six params (request, schema, ddl, seed, research, blueprint) is explained with meaning beyond its type, including key requirements, default behavior, and preconditions like validating the blueprint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource with scope: 'Make a verified relational dataset' where every foreign key, date ordering, aggregate and rate is checked before return. It explicitly distinguishes the fast path (schema/ddl, seconds) from start_generation for plain-English requests, so an agent can route without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use: call this when supplying schema or ddl, but use start_generation plus get_status polling for plain-English requests so the call doesn't time out. It also flags that research is off by default and that validate_blueprint should be run first when using a blueprint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_certificateDataset certificateA
Read-onlyIdempotent
Inspect
The full certificate for a dataset made by generate_dataset: every claim (stated vs actual),
every requirement's status and evidence, realism findings, planted defects/anomalies if any, and
the pattern charts. This is the whole answer key, not the trimmed form generate_dataset returns.
ParametersJSON Schema
NameRequiredDescriptionDefault
dataset_idYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered structurally. The description earns credit by disclosing the payload's depth and contents (full vs trimmed), which is meaningful behavioral context the annotations cannot convey. It does not mention rate limits, auth requirements, or pagination, keeping it below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler: the first establishes what the tool returns and enumerates contents, the second front-loads the key differentiation from the trimmed generate_dataset output. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must describe the return surface — and it does so thoroughly, listing every category of certificate content. With annotations covering the safety profile and a single obvious parameter, nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the single parameter dataset_id is undocumented in the schema. The description implicitly scopes it by tying the tool to 'a dataset made by generate_dataset,' which gives the identifier useful provenance, but it adds no format, sourcing, or lookup detail beyond that. Adequate but with a clear gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('the full certificate for a dataset') and enumerates exactly what it contains: claims stated vs actual, requirement statuses and evidence, realism findings, planted defects, and pattern charts. It also distinguishes itself from the sibling generate_dataset by noting this is the untrimmed form, so an agent can select it without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly situates usage: this is the 'whole answer key' versus the trimmed certificate generate_dataset returns, which tells the agent when to prefer this tool. It stops short of an explicit when-not-to-use statement or naming a distinct alternative tool by name, so it is clear context rather than a full routing rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_statusGeneration statusA
Read-onlyIdempotent
Inspect
Where a start_generation job is. While running: the stage it is in and how long it has run. When
done: the same answer generate_dataset returns (dataset_id, verification, tables, certificate).
ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnlyHint, idempotentHint and destructiveHint=false, lowering the bar, yet the description adds useful state-dependent return detail: running returns stage and elapsed time, done returns dataset_id, verification, tables and certificate. It omits polling frequency or rate-limit behavior, which would complete the picture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three telegraphic sentences, front-loaded with the core purpose and then the state-specific returns. Every sentence carries information; nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotation-level return info, the description does the heavy lifting by describing what comes back in each state. It is nearly complete, missing only polling/rate guidance and explicit job_id provenance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single job_id parameter is undescribed in the schema, so the description must compensate. It does so only indirectly, implying job_id identifies a job produced by start_generation, but gives no format or sourcing detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (a start_generation job) and what it reports ('Where a start_generation job is'), with the return shape spelled out per state. It also anchors itself to the start_generation sibling, so an agent can distinguish it from generate_dataset without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implicitly establishes the polling pattern after start_generation and clarifies the two states (running vs done), noting that the done-case result matches generate_dataset. It never states an explicit 'use this when' or a 'do not call if' exclusion, so guidance is clear but not fully prescriptive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plan_datasetPlan a datasetA
Read-onlyIdempotent
Inspect
See the tables, sizes and relationships the engine would build, before any rows exist. Free (no
rows are made), so it is worth calling before generate_dataset on anything non-trivial: review
what it understood and assumed, then adjust your request or schema before spending a real call.

Args:
    request: Plain-English description (needs an LLM key, unless `schema`/`ddl` is also given —
             then it is still used to ground realism, e.g. locale and what columns mean).
    schema:  A Misata schema dict (see the server instructions for the format). No key needed for
             structure.
    ddl:     CREATE TABLE statements. No key needed for structure.

Returns:
    route:         "chat" (not a dataset request — see `reply`), "design" (the model designed the
                   tables) or "pack" (matched a built-in shape).
    tables:        name, estimated rows, columns, foreign keys for each table the engine would build.
    understanding: what the engine read the request as (business, archetype, assumptions).
    requirements:  every specific thing the request asked for, so you can see what was understood.
ParametersJSON Schema
NameRequiredDescriptionDefault
ddlNo
schemaNo
requestNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnly, idempotent, non-destructive), yet the description still adds real behavioral context: the call is free and creates no rows, and an LLM key is required for `request` unless `schema`/`ddl` are supplied. The three possible `route` outcomes further disclose how the call actually resolves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and cost argument in the first sentence, then structured Args/Returns sections. Slightly verbose in the Args block, but every line conveys non-obvious information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the Returns block is genuinely necessary and enumerates route, tables, understanding, and requirements. Combined with annotations and param docs, an agent has everything needed to call and interpret this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description carries the burden and does so well: it explains each of the three params (plain-English request needing a key, Misata schema dict, CREATE TABLE ddl) and their key requirements. Minor deduction because the schema dict format is deferred to 'server instructions' rather than summarized.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete outcome — 'See the tables, sizes and relationships the engine would build, before any rows exist' — which is a dry-run preview distinct from every sibling. An agent can instantly distinguish it from generate_dataset without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to call it 'before generate_dataset on anything non-trivial' and names the alternative it precedes. It also tells the agent the follow-up action: review assumptions, then adjust the request or schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_datasetQuery a dataset with SQLA
Read-onlyIdempotent
Inspect
Read-only SQL over a dataset's own tables (one SELECT/WITH statement; every table name is a view
over that dataset's own files — nothing else on the server is reachable this way).

Args:
    dataset_id: from a prior generate_dataset call.
    sql:        a single SELECT (or WITH ... SELECT) statement. No semicolons, file paths, or
                statements that write or read outside the dataset (checked before running).
    limit:      rows returned (capped at 5000).

Returns:
    columns, rows, truncated (whether more rows existed than `limit`).
ParametersJSON Schema
NameRequiredDescriptionDefault
sqlYes
limitNo
dataset_idYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/no-open-world, so the safety profile is covered; the description nonetheless adds substantial non-obvious behavior: single-statement-only, no semicolons or file paths, validation '(checked before running),' a hard cap of 5000 rows, and the 'truncated' signal for overflow. These are enforcement details an agent cannot infer from annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core behavior, then cleanly partitioned into Args and Returns sections. Every sentence earns its place; the parenthetical about views and reachability eliminates a real ambiguity rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the return shape (columns, rows, truncated) and explaining what truncated means. All three parameters are documented, the validation and cap behavior are stated, and annotations cover safety — nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the load, and it does: it documents the provenance of dataset_id, the exact SQL grammar accepted, the prohibition on writes/outside reads, and that limit is capped at 5000. It omits the schema's default of 1000 for limit, a small gap against a 0%-coverage schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource+scope: 'Read-only SQL over a dataset's own tables,' and clarifies the isolation boundary ('every table name is a view over that dataset's own files'). No sibling tool (generate_dataset, export_dataset, plan_dataset, etc.) competes for this slot, so an agent can immediately tell what it does and what it does not reach.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It supplies a prerequisite that routes the agent correctly: dataset_id 'from a prior generate_dataset call,' which tells the agent this only works on an already-materialized dataset. It does not explicitly state when NOT to use it (e.g. use export_dataset instead when you want the full result downloaded), so it stops short of naming alternatives and exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_generationStart a generation (background)A
Idempotent
Inspect
Start a generation in the background and return a job_id at once. Use it for any plain-English
`request` (a model designs the tables, which takes minutes) — then call get_status(job_id) every
20-30 seconds until it says done. Arguments are the same as generate_dataset.

Returns:
    job_id, status "running". get_status gives the stage and, when finished, the dataset.
ParametersJSON Schema
NameRequiredDescriptionDefault
ddlNo
seedNo
schemaNo
requestNo
researchNo
blueprintNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=false, idempotentHint=true, destructiveHint=false), and the description adds behavior annotations do not: the call is asynchronous, returns immediately with a running job, and the model-design step takes minutes. The polling cadence and expected job state are genuinely useful operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the verb, the async nature, and the polling workflow; no sentence is padding. Since there is no output schema, the short Returns section earns its place, though the text is slightly loose in formatting.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly supplies the return contract (job_id, status 'running', and what get_status eventually yields). The remaining gap is the undocumented parameter set, which is mitigated but not eliminated by the reference to generate_dataset.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 6 parameters, so the description must carry the meaning. It adds real context for `request` (plain-English, model designs the tables) but handles the other five only by deferring to generate_dataset's arguments, forcing a cross-tool schema lookup for ddl, seed, schema, blueprint and research.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start a generation') plus the distinguishing scope ('in the background'), and names the immediate return value (job_id). This makes it clearly separable from the sibling generate_dataset, which is the synchronous counterpart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use guidance ('for any plain-English request') and a full follow-up workflow: poll get_status(job_id) every 20-30 seconds until done. It does not state when NOT to use it (e.g. when you want the dataset inline rather than a job), which keeps it short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_blueprintValidate a blueprintA
Read-onlyIdempotent
Inspect
Check a blueprint before generating it: every error said as what to change, the design traps it
falls into (a column that will round away, a copy of a value not made yet, a window that counts
nothing), its size against your row cap, and a small preview run (a few thousand rows) with the
verifier's findings and sample rows, so you can see the data behave before the real call. Free.

Returns:
    valid, errors (fix these), warnings (read these), estimated_rows, row_cap,
    preview: {passed, findings, tables: {name: first rows}} when `preview`.
ParametersJSON Schema
NameRequiredDescriptionDefault
previewNo
blueprintYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive, and the description adds genuinely useful context beyond them: it is 'Free,' the preview runs a few thousand rows, and it describes the verifier's findings and sample rows. It does not cover permissions or rate limits, so 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and benefits, with the Returns block cleanly separated. The parenthetical trap examples are verbose but informative rather than redundant, so it is appropriately sized though slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so: it lists valid, errors, warnings, estimated_rows, row_cap, and the preview shape. Safety is covered by annotations. The remaining gap is the unspecified blueprint object structure, but overall an agent has enough to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the `preview` parameter well ('a small preview run (a few thousand rows)'), but the required `blueprint` object (with additionalProperties and nested objects) gets no structural description beyond the general behavior of what is checked. Partial compensation warrants a 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Check a blueprint before generating it') and enumerates exactly what gets checked: errors, design traps, size against row cap, and a preview run. This clearly distinguishes it from generation siblings like generate_dataset or start_generation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'before generating it' places it clearly in the workflow ahead of the real generation call. It does not explicitly name alternative siblings (e.g., blueprint_guide) or state when not to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

whoamiWho am I connected asA
Read-onlyIdempotent
Inspect

Which account this connection is, and what it may do: signed in with a Studio MCP key or anonymous, the row cap, and whether a model key is available for plain-English requests.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, non-destructive, and closed-world behavior. The description adds useful behavioral context: it reports authentication mode and two capability details (row cap and model-key availability), telling the agent what the call surfaces beyond safety.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with the core identity answer and then the capabilities. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-argument read-only introspection tool with no output schema, the description supplies enough of the return content (auth mode, row cap, model key availability) to know what the call yields. Annotations cover safety.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No input parameters exist, so the schema is trivially complete. The description does not need to explain parameter semantics; baseline 4 applies for a zero-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific introspection purpose: account identity, auth mode (Studio MCP key or anonymous), row cap, and model key availability. An agent can immediately tell this is a connection-context probe, distinct from sibling dataset and generation tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than explicit: an agent can infer it should call whoami to inspect current identity and capabilities, but the description offers no when/when-not guidance or alternatives among siblings. No explicit routing is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 12 tool updates
    • First observedblueprint_guide
    • First observedcancel_generation
    • First observedexport_dataset
    • First observedfind_ready_dataset
    • First observedgenerate_dataset
    • First observedget_certificate
    • First observedget_status
    • First observedplan_dataset
    • First observedquery_dataset
    • First observedstart_generation
    • First observedvalidate_blueprint
    • First observedwhoami

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Generates realistic, referentially-coherent test data (SQL INSERTs, JSON, or CSV) from your database schema, resolving foreign keys and respecting constraints. Paste CREATE TABLE DDL or a JSON schema and get ready-to-run seed data with valid relationships.
    2
    37 npm
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Generate realistic multi-table synthetic datasets from plain English. Supports 18 industry domains, narrative growth curves (Black Friday, Q4 spike, 10x MRR), and 15 locale packs. Works with Claude Desktop, Cursor, Windsurf, Zed, and Continue.
    11
    289 PyPI
    69
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Runs SQL statements in memory against real PostgreSQL 18 and SQLite 3.49 engines after loading your schema and sample rows, returning each statement's own verdict: errors with line, column, and hints, constraint violations, or the resulting rows. It also compares both dialects to expose portability differences, all with no database server, network, or API key.
    1
    MIT
  • F
    license
    Not graded
    quality
    A
    maintenance
    Visual no-code generator that turns any database into multiple scoped MCP servers — one per access group, with PII masking and fail-closed query scoping built in.
    5 npm
    5
    -
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources