Skip to main content
Glama

paleobiology-mcp-server: query staged occurrences with SQL

paleobiology_dataframe_query
Read-onlyIdempotent

Run a read-only SQL SELECT against occurrence result sets staged on a DataCanvas by paleobiology_search_occurrences. This is how you analyze a large fossil set without re-fetching it: count occurrences by early_interval, group by formation, country (cc), or accepted_name, or filter by a paleo/modern coordinate range. The classification column is JSON — roll up by rank with json_extract_string(classification, '$.family') (also $.phylum, $.class, $.order, $.genus). Staged rows are occurrences, so collection-only fields such as lithology are not present. Reference tables by the table_name that search_occurrences returned — call paleobiology_dataframe_describe first if you do not know the table or column names. SELECT only; writes and file-reading functions are rejected.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
sqlYesA read-only SQL SELECT. Reference tables by the names paleobiology_search_occurrences / _describe returned.
canvas_idYesCanvas id returned by paleobiology_search_occurrences when its result spilled.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
rowsNoResult rows (capped at the canvas row limit). Keys are the selected column names.
errorNoPresent when the call failed. Absent on success.
row_countNoNumber of rows in the full result before any row cap.
truncatedNoTrue when the result exceeded the canvas row cap and rows were trimmed.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / properties / canvas_id / pattern
      Added value: +"^[A-Za-z0-9_-]{10}$"
  2. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, lowering the bar. The description adds meaningful behavioral constraints beyond that: only SELECT is allowed, writes and file-reading functions are rejected, rows are occurrences only so collection-only fields like lithology are absent, and classification is JSON requiring json_extract_string. This is substantial added transparency with no contradiction of the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat long but each section earns its place: purpose, use case, JSON extraction pattern, row-scope caveat, schema-discovery pointer, and access restriction. It is front-loaded with the core purpose and then layers important operational details. A small redundancy exists between 'read-only SQL SELECT' and the final 'SELECT only' restriction, but overall it remains tight for a SQL tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and annotations already cover read-only/idempotent behavior, the description is complete enough for correct invocation. It covers parameter semantics, the staging flow from search_occurrences, the JSON quirk, collection-field absence, and the safety boundary. An agent can use this tool without external documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds real semantic value by explaining what SQL forms are acceptable, how to reference table names from search_occurrences, and how to query JSON classification columns. These details go beyond the schema's parameter descriptions and help the agent form correct queries.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource: run a read-only SQL SELECT against occurrence result sets staged on a DataCanvas. It clearly ties the tool to paleobiology_search_occurrences and distinguishes it from sibling tools like paleobiology_dataframe_describe or search_collections. An agent can immediately understand this is the query/analysis step for staged occurrence data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use context: analyzing a large fossil set without re-fetching data, and provides examples of valid query patterns. It also steers the agent to paleobiology_dataframe_describe when table/column names are unknown. It stops short of explicitly enumerating all alternative sibling tools and when not to use them, but the staging and read-only constraints make the boundary clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.