Skip to main content
Glama

CatchAll (by NewsCatcher)

Create Dataset From Csv

create_dataset_from_csv

Create a new dataset by uploading a CSV file.

The CSV must have at least a name column. For meaningful entity enrichment each row should also include a domain column or a description column (or both) — a row with only a name is accepted but produces lower-quality enrichment. Additional columns are mapped to entity attributes. Max file size is plan-dependent. To add CSV rows to an existing dataset, use append_csv_to_dataset instead.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fileYesCSV content (required) — raw CSV text or standard base64-encoded CSV, capped at 10 MB after decoding. Server-side file paths are not accepted.
nameYesHuman-readable dataset name (required).
api_keyNoCatchAll API key. Optional if provided via x-api-key header or CATCHALL_API_KEY env var.
project_idNoOptional project ID to associate this dataset with (new in 1.6.1).
descriptionNoOptional dataset description.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It discloses that rows with only a name produce lower-quality enrichment, that max file size is plan-dependent, and that additional columns are mapped to entity attributes. However, it does not mention any failure modes, idempotency, or side effects beyond creation, which are minor gaps for a create operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured. It leads with the primary action, then provides necessary constraints and an alternative in a clear, scannable format. Every sentence contributes value, with no redundant fluff. The two-paragraph layout keeps it easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists and parameters are fully described, the description covers the essential usage context: what the tool does, required CSV format, enrichment caveats, file size constraints, and the alternative tool. It does not mention authentication or prerequisites, but these are implied by the api_key parameter and standard API usage. The description is sufficiently complete for an agent to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds meaningful semantics about the 'file' parameter by explaining the required columns and the enrichment behavior, which goes beyond the schema's simple 'CSV content' description. It also clarifies the role of additional columns. This added context justifies a score above the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear statement: 'Create a new dataset by uploading a CSV file.' This specifies the verb (create), resource (dataset), and method (uploading CSV). It also explicitly distinguishes itself from the sibling 'append_csv_to_dataset' by noting that the alternative is for adding rows to an existing dataset, which prevents confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage guidance: it states the required CSV structure (must have a 'name' column, and recommends 'domain' or 'description' for enrichment) and warns about file size limits. It directly names the alternative tool and the condition for using it ('To add CSV rows to an existing dataset, use append_csv_to_dataset instead'), leaving no ambiguity about when to choose this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources