Skip to main content
Glama
getsimba-ai

Simba MCP Server

Official
by getsimba-ai

Upload Data

upload_data

Import a CSV dataset into Simba to prepare it for Bayesian marketing mix model building and receive the file ID needed for model creation.

Instructions

Upload a CSV dataset to Simba for use in model building.

Provide EXACTLY ONE of csv_content (raw CSV text) or csv_path (a file path on the machine running this MCP server). Prefer csv_path for anything beyond trivial size — it avoids passing megabytes of CSV through the conversation.

The CSV should follow the canonical schema: one row per time period with date, KPI, multiplier, hierarchy, media activity/spend columns, and optional control variables.

IMPORTANT:

  • CSV only (not Excel). Maximum file size: 10 MB (API-enforced).

  • Row minimum: check get_data_schema -> x-simba-constraints.min_rows for the declared minimum; enforcement may be more permissive, and the upload response's warnings field is authoritative. More rows = tighter posteriors (104+ weekly rows recommended).

  • Media columns must follow naming: {channel}_activity and {channel}_spend.

  • Use 0 for inactive periods, not blank or NA.

  • csv_path is only available when the server runs locally (stdio). On HTTP/SSE deployments it is disabled unless SIMBA_MCP_ALLOW_LOCAL_FILES=1.

Args: csv_content: The full CSV text content (not base64, just raw CSV text). csv_path: Path to a .csv file readable by the MCP server process. name: Optional dataset name for identification. Defaults to the file stem when csv_path is used. filename: Optional original filename to record alongside the dataset. roles: Optional column roles for get_data_report, stored with the dataset: {column: role} or {column: {"role": role, "channel": name}}. Roles are declared, never guessed; see get_data_report for the vocabulary. An unknown role or a column the CSV lacks is refused.

Returns the uploaded file ID (needed for create_model), row/column counts, and any validation warnings.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
nameNo
rolesNo
csv_pathNo
filenameNo
csv_contentNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.12.0
    • addedInput schema / properties / roles
      Added value: +{
      +  "anyOf": [
      +    {
      +      "additionalProperties": true,
      +      "type": "object"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Roles"
      +}
  2. Changed1 schema field changedv0.5.0
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "additionalProperties": true,
      +  "title": "upload_dataDictOutput",
      +  "type": "object"
      +}
  3. First observedv0.3.2

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations, which only flag a non-readonly, non-idempotent, open-world write. The description adds hard operational constraints the agent cannot get elsewhere: CSV-only (not Excel), 10 MB API-enforced cap, row-minimum sourcing from get_data_schema with warnings as authoritative, naming conventions, and the stdio-vs-HTTP availability of csv_path. This is exactly the kind of context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the one-line purpose and mutual-exclusion rule, then groups constraints under an IMPORTANT block. Dense but every bullet (size cap, row minimum, naming, 0-fill, csv_path availability) is actionable. Slightly verbose, but the length is justified by the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite an output schema already existing, the description still summarizes the return (file ID for create_model, counts, warnings), and covers constraints, deployment caveats, and role semantics. For a 5-param upload tool with no schema descriptions, nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the full burden, and it does: it documents all five parameters in an Args block, including the mutual exclusion of csv_content/csv_path, defaulting behavior of name, and the shape and refusal semantics of roles. This adds real meaning beyond the bare schema types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (upload), resource (CSV dataset), destination (Simba), and immediate downstream purpose (model building). This clearly distinguishes it from sibling read tools like get_upload or list_uploads.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit rule for choosing between the two mutually exclusive input modes: 'Provide EXACTLY ONE... Prefer csv_path for anything beyond trivial size.' It also explains the deployment condition under which csv_path is available. No alternative tool is named, but the routing guidance for its own arguments is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools