Skip to main content
Glama

Misata Studio: verified synthetic data

Query a dataset with SQL

query_dataset
Read-onlyIdempotent
Read-only SQL over a dataset's own tables (one SELECT/WITH statement; every table name is a view
over that dataset's own files — nothing else on the server is reachable this way).

Args:
    dataset_id: from a prior generate_dataset call.
    sql:        a single SELECT (or WITH ... SELECT) statement. No semicolons, file paths, or
                statements that write or read outside the dataset (checked before running).
    limit:      rows returned (capped at 5000).

Returns:
    columns, rows, truncated (whether more rows existed than `limit`).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
sqlYes
limitNo
dataset_idYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/no-open-world, so the safety profile is covered; the description nonetheless adds substantial non-obvious behavior: single-statement-only, no semicolons or file paths, validation '(checked before running),' a hard cap of 5000 rows, and the 'truncated' signal for overflow. These are enforcement details an agent cannot infer from annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core behavior, then cleanly partitioned into Args and Returns sections. Every sentence earns its place; the parenthetical about views and reachability eliminates a real ambiguity rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the return shape (columns, rows, truncated) and explaining what truncated means. All three parameters are documented, the validation and cap behavior are stated, and annotations cover safety — nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the load, and it does: it documents the provenance of dataset_id, the exact SQL grammar accepted, the prohibition on writes/outside reads, and that limit is capped at 5000. It omits the schema's default of 1000 for limit, a small gap against a 0%-coverage schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource+scope: 'Read-only SQL over a dataset's own tables,' and clarifies the isolation boundary ('every table name is a view over that dataset's own files'). No sibling tool (generate_dataset, export_dataset, plan_dataset, etc.) competes for this slot, so an agent can immediately tell what it does and what it does not reach.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It supplies a prerequisite that routes the agent correctly: dataset_id 'from a prior generate_dataset call,' which tells the agent this only works on an already-materialized dataset. It does not explicitly state when NOT to use it (e.g. use export_dataset instead when you want the full result downloaded), so it stops short of naming alternatives and exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources