Skip to main content
Glama
getsimba-ai

Simba MCP Server

Official
by getsimba-ai

Get Data Report

get_data_report
Read-onlyIdempotent

Report actual stored dataset values over any date range or granularity, grouped by brand, channel, or dimension, using declared metric roles instead of model training windows.

Instructions

Report actual data from a stored dataset: any window, any grain, by brand, channel or dimension.

Reads the dataset itself (every column, any date range) rather than a fitted model's training window. Use it for "sales and TV spend in the North region for August, by week".

Roles are DECLARED, never guessed from column names. Declare them with upload_data(roles=...) or per request with roles. Without a declaration only the schema's own naming rules apply: a column named date, {channel}_spend and {channel}_activity. Every other column is reported as role "unknown" and is not aggregated — declare the KPI and hierarchy columns.

Role vocabulary (aggregation, unit) — also in get_data_schema under x-simba-roles:

  • kpi (sum), spend (sum, currency), activity (sum), multiplier (mean)

  • outcome:online_sales|store_sales|margin (sum, currency), outcome:orders|new_customers (sum)

  • media:impressions|clicks|grps (sum; give a channel: {"role": "media:grps", "channel": "tv"})

  • control:price|rate|index (mean), control:stock (each brand's last value, summed)

  • hierarchy, dimension:market|product|campaign (keys for filtering and group_by)

  • date

Buckets: week = ISO week from Monday; month/quarter = calendar. A weekly row counts in the month of its week-start date. The response's meta.aggregation states every rule applied.

Args: dataset_id: The uploaded file id (from upload_data or list_uploads). Registered pipeline outputs are uploaded files too. start, end: Optional ISO dates (YYYY-MM-DD), inclusive. granularity: "native" (default), "week", "month" or "quarter". group_by: "hierarchy", "channel", or a dimension role such as "dimension:market". hierarchy: Keep only this brand/region value. metrics: Roles or role families to include, e.g. ["kpi", "spend", "outcome:orders"] or ["control"]. Default: every metric role present. roles: {column: role | {"role", "channel"}} overriding roles stored at upload.

Returns {dataset: {id, name, source, version, sha256, data_through}, granularity, rows: [{period_start, period_end, group, metric, value, unit}], meta: {basis: "dataset", aggregation, roles, channels}}. Errors carry a code: dataset_not_found (404), invalid_report_request (400), report_too_large (413, over 10,000 rows — narrow the window or coarsen the granularity).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
endNo
rolesNo
startNo
metricsNo
group_byNo
hierarchyNo
dataset_idYes
granularityNonative

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.12.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (readOnly/idempotent/non-destructive): it documents aggregation rules, bucket semantics (ISO week from Monday, weekly rows counted in the month of week-start), the meta.aggregation disclosure, and specific error codes with statuses (404, 400, 413 with the 10,000-row cap and remediation). This is exactly the extra behavioral context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose and routes the reader through clearly labeled sections (roles, buckets, args, returns). It is dense and long, but given 0% schema coverage most of the length is load-bearing; only minor tightening is possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the full lifecycle an agent needs: correct invocation, role declaration fallback, aggregation semantics, return shape, and error handling. Since an output schema exists it wisely only summarizes the return structure rather than fully restating it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden, and it does: all 8 parameters are explained in an Args block (dataset_id origin, inclusive ISO start/end, granularity values, group_by options, hierarchy filter, metrics defaults, roles override format). It additionally supplies the role vocabulary and channel-qualified role syntax that the schema does not encode.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Report actual data from a stored dataset') with explicit scope ('any window, any grain, by brand, channel or dimension'). It also distinguishes itself from the model-based alternative by noting it 'Reads the dataset itself ... rather than a fitted model's training window', which separates it from siblings like get_model_results and get_data_schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete usage example ('sales and TV spend in the North region for August, by week') and implies the read-raw-data-vs-model distinction, plus the prerequisite that roles must be declared via upload_data(roles=...) or per request. It never explicitly names a sibling to avoid the way a 5 would, so guidance is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools