Skip to main content
Glama

Create schema from sample

create_schema_from_sample

Generate and auto-save a schema from reviewed samples, returning schema_id, schema content and record links. Supply entity_samples (or samples_csv), sample_record_id, or both; explicit samples override stored JSON while attachment inheritance is preserved. Samples must describe one entity type in one language. Generation combines their fields and annotates relationships; it does not redesign the approved structure. Resolve consequential modeling choices and whether to generate semantic IDs before calling; semantic IDs need an organization embedding model and add cost. Requires editor; synchronous and billed. Review returned suggestions before applying edits with property tools. Sample review and canonicalizations: enricher://docs/schema-from-sample; schema format: enricher://docs/schema-reference.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modelNoModel composite key, or auto (default) for the organization task selection. Explicit keys are discovered through list_models.auto
languageNoLanguage of schema type names/descriptions and annotations; defaults to the sample-key language. Sample property names are not translated. Separate from enrichment output languages.
samples_csvNoSamples as CSV text instead of entity_samples (never both): the first row is ALWAYS the header, each data row one sample. Delimiter ',', ';' or tab; headers become identifier keys ('Author Name' -> author_name); each column gets one type (integer, number, boolean or text; empty cell = null; decimal commas read in ';'/tab text). Over 20 rows, 20 are kept, covering every column. Read the result's csv_import and relay the kept rows and renamed headers.
attachment_idsNoUUIDs used as source context and persisted for regeneration. With sample_record_id, omit to inherit its linked attachments; pass [] to deliberately use none, or a non-empty list to override.
entity_samplesNo1..20 reviewed instances of one entity type. Required without sample_record_id. Explicit samples override stored JSON while keeping attachment inheritance. Fields are unioned; missing/null observations become nullable. Use consistent names across samples and array items. Single-language data only; set multilingual flags after generation with update_schema_property.
timeout_secondsNoWall-clock cap; returns a timeout error past this.
sample_record_idNoUUID of a successful sample_generation record. Its stored sample(s) are used when entity_samples is omitted, and its linked attachments are inherited when attachment_ids is omitted.
generate_semantic_idsNoAdd semantic IDs to eligible keyed objects. Requires an organization embedding model and adds resolution cost. Recommend for reusable identities without stable machine keys; obtain agreement before enabling unless already authorized. Defaults false. Modeling guidance: enricher://docs/schema-from-sample.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / properties / samples_csv
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "string"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "description": "Samples as CSV text instead of entity_samples (never both): the first row is ALWAYS the header, each data row one sample. Delimiter ',', ';' or tab; headers become identifier keys ('Author Name' -> author_name); each column gets one type (integer, number, boolean or text; empty cell = null; decimal commas read in ';'/tab text). Over 20 rows, 20 are kept, covering every column. Read the result's csv_import and relay the kept rows and renamed headers.",
      +  "title": "Samples Csv"
      +}
  2. First observed

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well past the annotations: discloses that it auto-saves, is synchronous and billed, requires editor role, that semantic IDs need an org embedding model and add cost, and that generation will not redesign the approved structure. The mutation ('auto-save') is consistent with readOnlyHint=false and destructiveHint=false. This is exactly the extra context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, inputs, and constraints, and every sentence is actionable (inputs, override rules, cost, role, docs link). It is somewhat long and partially overlaps the schema parameter descriptions, but there is little true filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and annotations cover the safety profile. The description still supplies the workflow, prerequisites, billing/cost caveats, and doc pointers an agent needs to call this correctly in one shot.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter already has a detailed inline description, so the schema carries most semantic load; baseline 3 applies. The description restates override/inheritance behavior and the single-entity/single-language constraint, but mostly mirrors what the schema descriptions already say rather than adding new meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource pair ('Generate and auto-save a schema from reviewed samples') and names the concrete outputs (schema_id, schema content, record links). It is clearly distinguishable from siblings like analyze_sample, save_schema, and create_benchmark_scenario. An agent knows exactly what this produces without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real when-to-use guidance: supply entity_samples or samples_csv (never both), or sample_record_id, and resolve modeling/semantic-ID decisions before calling. It also directs the agent to the docs link for review/canonicalization. It stops short of naming explicit sibling alternatives (e.g. save_schema, add_schema_property) as substitutes, so it is clear context but not full routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.