Talonic - Data Extraction MCP
Server Details
Extract structured, schema-validated data from any document, including PDFs, scans, images, spreadsheets, and forms, directly inside the chat.
- Status
- Healthy
- Uptime
- 3.5% over 55 days
- OAuth
- Works in Glama
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Score is being calculated.
Available Tools
36 toolstalonic_askAsk a question over the workspaceAInspect
Ask a natural-language question over the workspace's documents and get a cited, verified answer (markdown). The Talonic agent plans over the structured field plane, runs read-only SQL over extracted cells, reads document text, and grounds every load-bearing claim in a source span. Consumes credits.
USE WHEN: the user asks an open question about their documents ('which vendors invoiced us twice in May?'), wants a summary across documents, or the answer needs reasoning over several fields.
NOT FOR: reading a known field's values (talonic_field_values, free) or filtering documents by a known value (talonic_filter, free); locating which field holds a concept (talonic_find_data).
ARGS: question; optional scope { document_ids[], schema_id, pipeline_id, data_product_id, document_type, source_id, tags[], ingested_after, ingested_before } (ANDed), conversation_id (continue a thread), output_format { instruction, template }, wait_seconds (0–55, default 45).
RETURNS: { ask_id, status ('completed'|'processing'|'error'), conversation_id, answer (markdown), citations[] { quote, document_id, kind, filename, app_url }, verification { verdict, checks_total, checks_unsupported, correction }, usage { tokens, credits_charged }, tool_calls, artifacts[], waited_ms }. If status is still 'processing' after the wait, call talonic_get_answer with the ask_id.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | Restrict the question to a slice of the workspace; present fields are ANDed. | |
| question | Yes | The question, in the user's words. | |
| wait_seconds | No | Seconds to wait for the answer before returning 'processing' (default 45, max 55). | |
| output_format | No | Shape the answer (form only, never grounding). | |
| conversation_id | No | Continue this conversation; the agent sees prior turns. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses internal mechanics: planning over the structured field plane, running read-only SQL over extracted cells, reading document text, and grounding claims in source spans. It also reveals the cost ('Consumes credits') and the asynchronous behavior with instructions to call talonic_get_answer if status is still 'processing'. This goes well beyond the annotation hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is organized into clear sections: purpose, when to use, when not to use, args, and returns. The purpose is front-loaded, and every sentence carries routing, invocation, or behavior information. Despite being long, the length is justified by the tool's complexity and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and the absence of an output schema, the description is self-sufficient: it covers parameters, scoping, return shape, cost, verification, and the async continuation path. The explicit instruction to call talonic_get_answer when status is 'processing' closes the loop for correct use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents all parameters including ANDed scope semantics, wait_seconds default, and output_format behavior. The description's ARGS section provides a useful at-a-glance summary but little new semantic information beyond the schema, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Ask a natural-language question over the workspace's documents') and names the deliverable: a cited, verified markdown answer. It clearly distinguishes this tool from sibling tools by describing it as for open, multi-field reasoning questions. There is no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description has explicit USE WHEN and NOT FOR sections that name the sibling alternatives (talonic_field_values, talonic_filter, talonic_find_data) and even note that they are free. It states concrete conditions for choosing this tool over those alternatives, so an agent can route correctly without guessing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_claim_agent_taskClaim Agent TaskAInspect
Claim an available Agent-stage task, or reclaim it after its lease expires.
USE WHEN: ready to process a task. Save the returned execution_epoch and lease_expires_at. NOT FOR: merely inspecting work (use talonic_get_agent_task) or extending an active lease (use talonic_heartbeat_agent_task). A conflicting live claim returns HTTP 409. A successful claim returns the task payload.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | Agent task UUID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations (readOnlyHint=false, destructiveHint=false) are present and align with the description. The description adds behavioral context beyond annotations, such as lease expiry, conflicting live claim returning 409, and the need to save execution_epoch and lease_expires_at. This makes the tool's behavior transparent and expected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise yet comprehensive, using structured sections 'USE WHEN', 'NOT FOR', and explicit behavioral notes. It avoids redundancy and packs essential information without being overly verbose. The structure enhances readability and preserves information density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers all critical aspects: the action (claim/reclaim), when to use, what to save from the return, error handling (409 conflict), and the output (task payload). It also differentiates from related tools, ensuring the agent knows exactly when this tool is appropriate. The description is self-sufficient for safe and correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, task_id, is fully described in the schema with 'Agent task UUID.' The description also hints at the return value containing execution_epoch and lease_expires_at, which provides additional semantic context. Since schema coverage is 100% and the description adds useful detail, parameter semantics are clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Claim an available Agent-stage task, or reclaim it after its lease expires.' It specifies the verb 'claim', the resource 'Agent-stage task', and additional nuance about reclaiming. It also distinguishes from sibling tools like talonic_get_agent_task and talonic_heartbeat_agent_task.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit usage guidance is provided: 'USE WHEN: ready to process a task' and 'NOT FOR: merely inspecting work (use talonic_get_agent_task) or extending an active lease (use talonic_heartbeat_agent_task)'. Behavioral details about conflicting claims (HTTP 409) and successful claims (returns task payload) further clarify when and how to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_claim_decision_taskClaim Decision TaskAInspect
Claim an available decision task (or reclaim one whose lease expired) and receive the decision bundle: task metadata with the new execution_epoch, output_contract, precedents, and a package descriptor { package_kind, record_count, page_size, first_cursor, documents }.
USE WHEN: ready to decide a listed task. There is no separate get: claiming IS the payload fetch. Save execution_epoch and lease_expires_at; then read the records with talonic_read_decision_package starting at first_cursor (null means the package has no records). NOT FOR: extending a live lease (talonic_heartbeat_decision_task) or returning a decision (talonic_submit_decision_task). A conflicting live claim returns HTTP 409. AUTH: a tlnc_ key needs a per-app 'decide' grant; an OAuth connector session needs the apps:decide scope (consented at connect) and a live workspace role of senior_member or above. A 403 (decide_grant_required, insufficient_scope, insufficient_tier) names what is missing — tell the user, do not retry.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | Decision task UUID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even though annotations already mark this as not read-only and not destructive, the description adds significant behavioral context: claiming is the payload fetch, leases can expire and be reclaimed, a conflicting live claim returns HTTP 409, and auth failures return 403 with specific error codes (decide_grant_required, insufficient_scope, insufficient_tier) and instruct the agent not to retry. This goes well beyond what annotations alone convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but tightly structured with clear labels: the core behavior, USE WHEN, NOT FOR, and AUTH. Every sentence carries operational value—there is no filler or repetition. The most critical facts (claiming fetches the bundle, save the lease fields, use first_cursor) are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there is no output schema, the description compensates by explicitly naming the decision bundle fields, the package descriptor fields, the meaning of null first_cursor, the follow-up tool to use, conflict behavior, and authentication requirements. This is complete enough for an agent to invoke the tool correctly without further inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers 100% of the parameter information: task_id is described as 'Decision task UUID' with format and pattern constraints. The description adds workflow context but does not add new parameter-level semantics beyond the schema. With full schema coverage, the baseline of 3 is appropriate; there is no gap to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Claim an available decision task'), the target resource, and a concrete outcome ('receive the decision bundle') with a detailed list of returned fields. It explicitly distinguishes itself from siblings by noting that claiming is the payload fetch and there is no separate get. An agent can clearly tell this apart from talonic_read_decision_package, talonic_heartbeat_decision_task, and talonic_submit_decision_task.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit USE WHEN guidance ('ready to decide a listed task') and a NOT FOR section naming the exact alternatives: talonic_heartbeat_decision_task and talonic_submit_decision_task. It also provides workflow direction: save execution_epoch and lease_expires_at, then read with talonic_read_decision_package starting at first_cursor. This leaves no ambiguity about when to invoke this tool versus its siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_extractExtract Data from DocumentInspect
Extract structured, schema-validated JSON from a document: PDF, scan, image, DOCX, email, or photo. Returns the requested fields with per-field confidence scores.
USE WHEN: the user asks to extract data from a document, turn a PDF into JSON, pull fields from a file, or parse a scan / form / statement / receipt / report, for common (invoice, contract) or unusual document types.
NOT FOR: full plain text (use talonic_to_markdown) · finding documents (use talonic_search / talonic_filter).
BY NAME: if the user names a file, call talonic_search first to get its document_id, then call this.
ARGS: define the fields you want with inline schema (JSON Schema, e.g. {type:'object',properties:{vendor_name:{type:'string'}}}) OR a saved schema_id, not both. Don't know the fields yet? Set auto_schema:true to let Talonic discover them (open capture) and return a suggested schema you can refine. Provide EXACTLY ONE document source: document_id (cheapest, a workspace doc), file_url (public URL), or file_data+filename (small local files only).
COST: each call uses credits; talonic_get_balance shows the remaining balance.
RETURNS: data (the JSON), confidence.overall and confidence.fields (treat <0.7 as needs review), document metadata, extraction_id.
| Name | Required | Description | Default |
|---|---|---|---|
| schema | No | Inline schema definition. REQUIRED unless `schema_id` is provided. Recommended: full JSON Schema {type:'object', properties:{...}}. Also accepted: flat key-type map {field_name:'string', amount:'number'}. Mutually exclusive with `schema_id`. | |
| file_url | No | URL to a document file. The Talonic API fetches it server-side. Use this for documents already on the public web. | |
| filename | No | Original filename including extension, e.g. 'invoice.pdf'. Used to infer MIME type when uploading via `file_data`. Required when `file_data` is provided. | |
| file_data | No | Base64-encoded file bytes. Recommended path when the agent already has the file in memory (e.g., the user attached a PDF to the conversation). Pair with `filename` so MIME type can be inferred. Works regardless of where the file lives on disk. | |
| file_path | No | Local path to a document file. Only works if the MCP server has read access to that path. In sandboxed chat clients (Claude Desktop, Cowork) where uploads land in a host-owned directory, use `file_data` instead. | |
| schema_id | No | ID of a saved schema. REQUIRED unless `schema` is provided. Accepts UUID or SCH-XXXXXXXX short id from talonic_list_schemas. Mutually exclusive with `schema`. | |
| auto_schema | No | Open capture: when true, extract WITHOUT providing a schema — Talonic discovers the document's fields and returns them plus a suggested schema you can refine and reuse. Use this when you don't yet know the fields. Mutually exclusive with `schema` and `schema_id`. | |
| document_id | No | ID of a document already in the workspace, to re-extract with a new schema. | |
| instructions | No | Natural-language guidance for the extractor, e.g. 'Focus on the billing section. Amounts are in EUR.' | |
| include_markdown | No | Include OCR-converted markdown in the response alongside structured data. | |
| include_provenance | No | Include per-field provenance (source_text, section, page) showing where each value was found in the document. |
Output Schema
| Name | Required | Description |
|---|---|---|
| cost | No | Per-call cost and post-call balance, parsed from the X-Talonic-* response headers. `null` for non-extract calls; not always present on legacy clients. |
| data | Yes | The extracted structured data, shape determined by the schema. |
| links | No | URLs for self, document, and human-readable dashboard view. |
| schema | No | Schema metadata: which schema was used and how it can be saved. |
| status | Yes | Extraction status (e.g. 'complete'). |
| document | Yes | Metadata about the ingested document. |
| markdown | No | OCR-converted markdown. Present only when `include_markdown: true`. |
| confidence | No | Extraction confidence. Treat fields below ~0.7 as needing human review. |
| processing | No | Processing metadata: duration, pages processed, region. |
| provenance | No | Per-field source evidence (source_text, section, page). Present only when `include_provenance: true`. |
| request_id | No | Server-assigned request ID for support and debugging. |
| extraction_id | Yes | Stable identifier for this extraction. |
talonic_fail_decision_taskFail Decision TaskDestructiveInspect
Report that the claimed decision task cannot be decided: raises a Human Review with your reason AND applies the app's declared fallback policy (rules decide, hold for review, or fail the run).
USE WHEN: the package is insufficient or contradictory and no claimant could decide it. This ends the task (status 'failed'). NOT FOR: temporary give-backs (talonic_release_decision_task) or a decision you can make with low confidence (submit it with confidence set). ARGS: task_id, execution_epoch, reason (up to 2,000 characters). Stale epoch is HTTP 409. AUTH: a tlnc_ key needs a per-app 'decide' grant; an OAuth connector session needs the apps:decide scope (consented at connect) and a live workspace role of senior_member or above. A 403 (decide_grant_required, insufficient_scope, insufficient_tier) names what is missing — tell the user, do not retry.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | Why the task cannot be decided, up to 2,000 characters. Recorded on the Human Review the fail raises. | |
| task_id | Yes | Decision task UUID. | |
| execution_epoch | Yes | Execution epoch returned by the successful claim. Stale epochs are rejected with 409. |
talonic_field_valuesRead Field ValuesARead-onlyIdempotentInspect
Read a field's CURRENT VALUES across documents, with provenance — one row per bound occurrence: document id + filename + type, the value, confidence, the raw name it was captured under, the verbatim source text, and the resolution band that bound it.
USE WHEN: the user asks 'what are all the X across my documents', you need to tabulate or aggregate one concept across the corpus, or you want the evidence (source text + document) behind a value. NOT FOR: multi-field row-shaped queries over documents (talonic_filter) or one document's full field set (talonic_get_document).
ARGS: exactly one of field_id or name; optional document_id (one document), value (case-insensitive contains filter), limit (max 100), cursor.
RETURNS: { field_id, canonical_name, concept_ids, data[] of { occurrence_id, document_id, document_filename, value, confidence, provenance{ raw_field_name, source_text, resolved_by, needs_confirmation, via_redirect }, links }, pagination }. Rows are Sources-IAM filtered for the caller.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Field NAME instead of an id — resolved through canonical name, synonyms, merge aliases and the registry's spelling fold, then followed to the live concept. Use the user's wording ('Invoice No', 'Vertragsnummer'). | |
| limit | No | Page size (default 20, max 100). | |
| value | No | Case-insensitive contains filter on the value text. | |
| cursor | No | Opaque cursor from pagination.next_cursor. | |
| field_id | No | Field UUID (from talonic_list_fields / talonic_search). | |
| document_id | No | Only occurrences on this document. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, and the description reinforces a read-only data-retrieval behavior. Beyond annotations, it discloses provenance details, one-row-per-occurrence semantics, the return envelope, pagination, and that rows are Sources-IAM filtered for the caller. This is rich behavioral context with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear USE WHEN, NOT FOR, ARGS, and RETURNS sections, and the core behavior is stated in the first sentence. Every sentence adds decision-relevant information; there is no filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining the return shape, and it does so in useful detail: field metadata, data array fields, nested provenance fields, links, and pagination. It also covers all parameters, scoping behavior, and IAM filtering, making it complete enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents each parameter. The description adds meaningful selection-semantics beyond the schema by stating 'exactly one of field_id or name' and summarizing the optional filters, value as a case-insensitive contains filter, limit cap, and cursor. This compensates for the schema's lack of required-parameter constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with a specific verb and resource: 'Read a field's CURRENT VALUES across documents, with provenance — one row per bound occurrence.' It names the exact output granularity and distinguishes itself from siblings in NOT FOR, explicitly calling out talonic_filter and talonic_get_document. An agent can understand what this tool does and how it differs from nearby tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Contains explicit USE WHEN guidance ('what are all the X across my documents', tabulate/aggregate, need evidence behind a value) and NOT FOR exclusions with sibling names. This gives the agent clear decision criteria for when to select this tool over alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_filterFilter Talonic DocumentsARead-onlyInspect
Find documents by their extracted field VALUES using composable conditions (e.g. 'invoices where total > 1000').
USE WHEN: value-based criteria on extracted fields — numeric/date/text comparisons or presence checks.
NOT FOR: free-text / concept search (use talonic_search) · a single document by id (use talonic_get_document).
ARGS: conditions[] (AND-ed). Each = EXACTLY ONE of field (canonical name) or field_id (UUID), an operator, and usually a value. Operators: eq, neq, gt, gte, lt, lte, between (needs value AND value_to), contains, is_empty / is_not_empty (no value). value/value_to are string|number|boolean matching the field type (ISO YYYY-MM-DD for dates).
TEXT FILTERS: for eq/contains/is_not_empty on a text field, just TRY a natural field name ('currency', 'vendor_name') — names resolve server-side and an unresolved field surfaces in warnings[] rather than erroring. Do NOT block on discovering the field first; search-first is only required for numeric operators.
NUMERIC GUARD: gt/gte/lt/lte/between only work when the field's dataType is 'number'. Call talonic_search first and check dataType; a numeric op on a string field returns zero matches. If the response has warnings[], surface them to the user — do not silently retry.
RETURNS: data[] (matching documents with field values), total, warnings[].
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number for pagination. | |
| sort | No | Optional sort by a field. | |
| limit | No | Results per page. Default 50 server-side. | |
| search | No | Optional free-text search applied alongside the filters. | |
| conditions | Yes | One or more filter conditions, AND-ed together. | |
| source_connection_id | No | Optionally scope to a specific source connection. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | Yes | Documents matching the filter conditions, with their extracted field values. |
| page | No | Current page number. |
| total | No | Total documents matching across all pages. |
| warnings | No | API warnings surfaced by the Talonic filter endpoint. Most commonly raised when a numeric operator is applied to a string-typed field, in which case the warning explains the lexicographic-comparison trap and suggests a schema-design change. Agents should surface these to the user rather than silently retrying. |
| pagination | No | Cursor-based pagination metadata. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations give readOnlyHint and destructiveHint. Description adds: conditions are AND-ed, operator-specific behaviors (is_empty takes no value), numeric guard (check dataType first), text field name resolution with warnings, and instruction to surface warnings. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (USE WHEN, NOT FOR, ARGS, TEXT FILTERS, NUMERIC GUARD, RETURNS) and front-loaded. It's somewhat long but each sentence adds value, appropriate for the complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 params, nested objects, many operators) and existence of output schema, the description covers all critical aspects: purpose, usage, parameter details, edge cases, warnings, return structure. Nothing missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but description adds deep semantics: explains field/field_id exclusivity, operator list, value/value_to for between, data types, and text filter resolution. This goes well beyond schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Find documents by their extracted field VALUES using composable conditions' and gives an example. It distinguishes from siblings talonic_search and talonic_get_document in the 'NOT FOR' section.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'USE WHEN: value-based criteria on extracted fields' and 'NOT FOR: free-text / concept search (use talonic_search); a single document by id (use talonic_get_document).' Also provides guidance on numeric guard and warning handling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_find_dataFind Data by ConceptARead-onlyIdempotentInspect
Locate the REAL data behind a natural-language concept before querying anything: semantic + lexical retrieval that resolves a phrase ('payment volume per transaction', 'counterparty', 'Vertragslaufzeit') to the registry fields, values, documents and text passages that carry it — even when the field is captured under a different name.
USE WHEN: the user asks about a concept and you are not sure which field holds it, when talonic_list_fields / talonic_search came back empty or ambiguous, or when the answer may live in document prose rather than a captured cell. NOT FOR: reading a known field's values (talonic_field_values) or filtering by a known field (talonic_filter).
ARGS: query (the concept, in the user's words), optional top_k (1–25, default 10), document_ids (hard scope).
RETURNS: ranked planes — FIELDS (canonical_name, field ids/keys, maturity/tier, occurrence_count, sample values with their documents), VALUES, DOCUMENTS and PASSAGES — every item a ready handle for the next call. Read-only, no LLM cost.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The natural-language concept to locate. | |
| top_k | No | Max results per plane (default 10). | |
| document_ids | No | Restrict to these document ids (hard filter, enforced server-side). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false. The description goes beyond that by specifying the return structure (ranked planes: FIELDS, VALUES, DOCUMENTS, PASSAGES) and adding operational context: 'Read-only, no LLM cost.' It also notes that results are 'ready handles for the next call,' which tells the agent the output is directly reusable. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place, organized with bolded section headers (USE WHEN, NOT FOR, ARGS, RETURNS) that make scanning easy. The core purpose is front-loaded in the first sentence, and the rest provides structured, non-redundant detail. No filler or tautology.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is complex (semantic retrieval across multiple planes, no output schema), and the description fully compensates. It explains what each return plane contains, how to interpret the results (ranked, with handles for subsequent calls), and the hard scope semantics of document_ids. With no output schema, this description carries the full burden and meets it. Also covers safety via annotations and adds the 'no LLM cost' note, which is operationally important.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: each parameter already has a description in the schema (query, top_k, document_ids). The description reinforces these by restating query as 'the concept, in the user's words,' top_k default as 10, and document_ids as a hard scope. This adds minor nuance beyond the schema (e.g., 'hard scope' and default) but largely repeats it, so it stays above baseline 3 without reaching 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('locate') and resource ('real data behind a natural-language concept'), then immediately explains the mechanism (semantic + lexical retrieval) and the scope (fields, values, documents, passages). It explicitly distinguishes itself from siblings by naming talonic_list_fields / talonic_search and stating when those fall short, so an agent can select it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a dedicated 'USE WHEN' section that lists three concrete triggers (unsure which field holds a concept, sibling tools returned empty/ambiguous, answer may live in prose) and a 'NOT FOR' section naming two alternatives (talonic_field_values, talonic_filter). This is explicit, actionable routing guidance with no room for inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_get_agent_taskGet Agent TaskARead-onlyInspect
Fetch one Agent-stage task's immutable input snapshot, instructions, and declared output contract. This disclosure is audited.
USE WHEN: inspecting a listed task before deciding whether to process it. NOT FOR: acquiring the task (use talonic_claim_agent_task) or returning results (use talonic_submit_agent_task). ARGS: task_id. RETURNS: metadata, input_snapshot, output_contract, instructions, and timeout_fallthrough.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | Agent task UUID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety profile is covered. The description adds value by stating the data is immutable and audited, and mentions the timeout_fallthrough return field, which is non-annotated behavioral context. A small deduction because it doesn't elaborate on what timeout_fallthrough means or other edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with purpose, and uses clear labeled sections (USE WHEN / NOT FOR / ARGS / RETURNS). Every sentence provides actionable information with no repetition or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-based tool with rich annotations and 100% schema coverage, the description covers the purpose, usage, and return fields. The only gap is not explaining the 'timeout_fallthrough' field, but that's an output detail rather than a functional requirment. It doesn't need an output schema since it lists return keys.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers task_id fully with format and description, so baseline is 3. The description just says 'ARGS: task_id' which adds no new semantics. That's acceptable given the schema's completeness.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches an Agent-stage task's immutable input snapshot, instructions, and output contract. It names the specific resource ('Agent-stage task') and distinguishes it from siblings like claim/submit by explicitly naming those alternatives. The verb 'Fetch' is concrete and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit USE WHEN and NOT FOR guidance, naming tailonic_claim_agent_task and tailonic_submit_agent_task as alternatives. This fully clarifies the tool's position in the workflow and prevents misuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_get_answerPoll an ask for its answerARead-onlyIdempotentInspect
Poll an ask started by talonic_ask that was still processing when the wait ended.
USE WHEN: talonic_ask returned status 'processing' with an ask_id — poll every few seconds until 'completed' or 'error'.
NOT FOR: asking a new question (talonic_ask).
ARGS: ask_id.
RETURNS: the same answer envelope as talonic_ask (answer, citations[], verification, usage) or { status: 'processing', poll_hint }.
| Name | Required | Description | Default |
|---|---|---|---|
| ask_id | Yes | From talonic_ask. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context beyond annotations: polling cadence, expected terminal statuses ('completed' or 'error'), and the processing response shape with poll_hint. This is meaningful supplementary disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core purpose, and uses labeled sections (USE WHEN, NOT FOR, ARGS, RETURNS) that make scanning easy. Every sentence contributes useful information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter polling tool, the description is complete: it explains when to use it, what argument to pass, what the return envelope looks like, and how to handle the processing state. Annotations cover safety and idempotency, and the schema covers the parameter format. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents ask_id as 'From talonic_ask.' The description reinforces this by saying ARGS: ask_id and explaining it comes from a talonic_ask response, but it does not add substantial new meaning beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Poll') and resource ('an ask for its answer'), and explicitly distinguishes itself from talonic_ask by stating it is for asks that were still processing. This makes the tool's purpose unmistakable and separates it from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit USE WHEN conditions (talonic_ask returned status 'processing' with an ask_id) and a NOT FOR exclusion (asking a new question via talonic_ask). This gives an agent clear decision criteria for when to select this tool over alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_get_balanceGet Talonic Credit BalanceARead-onlyInspect
Read the workspace's Talonic credit balance, EUR value, tier, 30-day burn, and projected runway.
USE WHEN: the user asks about credits/budget, or before a large batch when you want to confirm headroom. NOT FOR: the per-call cost of a single extraction (that is on the talonic_extract response). ARGS: none. RETURNS: balance_credits, balance_eur, tier, burn_rate_30d_credits, projected_runway_days (-1 = no recent usage), tier_resets_at.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| tier | Yes | API tier of the workspace. |
| balance_eur | Yes | Current balance in EUR (two decimals). |
| tier_resets_at | Yes | ISO 8601 timestamp of the next monthly tier reset. |
| balance_credits | Yes | Current credit balance. |
| burn_rate_30d_credits | Yes | Total credits consumed in the trailing 30 days. |
| projected_runway_days | Yes | Projected days of runway at the current 30-day average burn. `-1` when burn is zero (cannot compute). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it read-only. The description adds value by detailing returned fields and noting that projected runway is -1 when no recent usage. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is concise: a single introductory sentence followed by structured USE WHEN, NOT FOR, ARGS, and RETURNS sections. Every sentence is informative and earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (read-only, no parameters, documented output schema), the description covers all necessary context: what is returned, when to use, and what not to use for. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 0 parameters (100% coverage). Description explicitly states 'ARGS: none.' Baseline for 0 parameters is 4, and the tool meets that standard without needing further parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it reads the workspace's Talonic credit balance, EUR value, tier, 30-day burn, and projected runway. The verb 'Read' and specific resource details make purpose unambiguous. It naturally distinguishes from credit-related queries that would go to talonic_extract.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit USE WHEN and NOT FOR sections provide clear guidance: use when user asks about credits/budget or before a large batch; not for per-call cost (pointing to talonic_extract). This effectively differentiates from sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_get_documentGet Talonic DocumentARead-onlyInspect
Fetch a single document's metadata and processing status from the workspace.
USE WHEN: 'tell me about document X', or to poll status after talonic_request_upload until the file is ready.
NOT FOR: full text (use talonic_to_markdown) · extracted fields (use talonic_extract).
BY NAME: if the user names a file, call talonic_search first to get its document_id, then call this.
ARGS: document_id.
RETURNS: filename, pages, type_detected, language, and status. Status lifecycle: pending_upload -> uploading -> queued -> extracting -> completed. Wait for completed before calling talonic_extract on a freshly uploaded doc. Terminal failure statuses: ocr_failed, extraction_failed, error — stop polling and report the failure to the user if any of these appear. To read the document's text, call talonic_to_markdown with this id.
| Name | Required | Description | Default |
|---|---|---|---|
| document_id | Yes | The Talonic document ID. Get this from a previous talonic_extract or talonic_search response. |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| links | No | |
| pages | No | |
| source | No | |
| status | No | |
| triage | No | |
| filename | No | |
| mime_type | No | |
| created_at | No | |
| size_bytes | No | |
| original_path | No | |
| type_detected | No | |
| processing_log | No | |
| extraction_count | No | |
| language_detected | No | |
| latest_extraction_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true. The description adds valuable behavioral details: the status lifecycle, terminal failure statuses, and instructions to stop polling and report failures. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured with clear sections (USE WHEN, NOT FOR, BY NAME, ARGS, RETURNS). Every sentence adds value, and the formatting aids readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (1 parameter, output schema exists), the description covers purpose, usage, behavior, parameter origin, return values, and status lifecycle. It also integrates with siblings for complete workflow guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and describes the document_id parameter well. The description adds context about getting the ID from previous responses, but does not provide new semantic meaning beyond what the schema offers. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Fetch a single document's metadata and processing status from the workspace.' It uses a specific verb and resource, and distinguishes itself from sibling tools (talonic_to_markdown, talonic_extract) by stating what it is NOT for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'USE WHEN' and 'NOT FOR' sections provide clear context. It specifies when to poll after upload, when to use alternatives, and even instructs to call talonic_search first if user names a file.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_get_fieldGet Field Concept CardARead-onlyIdempotentInspect
Get the CONCEPT CARD for one Field Registry field: what it means (curated description + extraction instruction), its synonyms and aliases, maturity, where it occurs (document/occurrence counts, first/last seen, document-type spread), its value distribution (top values with counts, distinct count, examples), schema usage, and identity links (superseded_by, absorbed concepts).
USE WHEN: you must decide whether a field is the right concept for a question, need example values or the value shape before writing a filter, or hold a field NAME from the user and need the live concept behind it. NOT FOR: listing many fields (talonic_list_fields) or reading every value (talonic_field_values).
ARGS: exactly one of field_id or name. Names are resolved through canonical name → spelling fold → merge aliases → synonyms (then case-insensitive fallbacks) and followed to the live concept; the response says which arm matched. include_history: true appends the curation trail (merges, renames, maturity moves).
RETURNS: the card { id, canonical_name, maturity, data_type, definition, identity, occurrence, values, usage, links } plus resolution when a name was given and history when requested.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Field NAME instead of an id — resolved through canonical name, synonyms, merge aliases and the registry's spelling fold, then followed to the live concept. Use the user's wording ('Invoice No', 'Vertragsnummer'). | |
| field_id | No | Field UUID (from talonic_list_fields / talonic_search). | |
| include_history | No | Append the concept's curation history (tier changes, merges, renames), newest first. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds real behavioral context beyond that: the name-resolution pipeline (canonical name → spelling fold → merge aliases → synonyms → case-insensitive fallbacks), the fact that the response reports which resolution arm matched, and that include_history appends the curation trail. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized with labeled sections: definition, USE WHEN, NOT FOR, ARGS, and RETURNS. The most important usage information is front-loaded, and each section earns its place. The small overlap between the opening summary and the RETURNS key list improves scannability rather than adding bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description provides the full return shape and the conditional presence of resolution and history fields. It also explains the resolution behavior and the exact-one-argument rule. The main gap is missing failure semantics, such as what happens when a supplied name or field_id cannot be resolved, which would help an agent handle the not-found case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds a meaningful nuance the schema does not encode: exactly one of field_id or name must be provided, even though the schema marks none as required. It also clarifies resolution order and the conditional return fields, going slightly beyond the per-parameter schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get'), a specific resource ('CONCEPT CARD for one Field Registry field'), and an explicit inventory of what the card contains. It further distinguishes itself from siblings by naming talonic_list_fields and talonic_field_values in the NOT FOR section. An agent can confidently tell this apart from the other field-related tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The USE WHEN section gives three concrete triggering scenarios, and the NOT FOR section explicitly names alternatives and what they are for. It also adds the crucial invocation constraint that exactly one of field_id or name should be supplied. This is explicit, actionable guidance with clear exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_get_pricingGet Talonic PricingARead-onlyInspect
Read Talonic's machine-readable credit pricing catalog: fixed per-unit rates so you can predict spend BEFORE running anything.
USE WHEN: estimating the cost of a planned extraction/structuring/matching job, or answering a pricing question. Public — works without spending credits. NOT FOR: the workspace's current balance (use talonic_get_balance) or what it has already spent (use talonic_get_usage). ARGS: none. RETURNS: currency, credits_per_eur, multipliers (e.g. batch 0.5x), and units[] — each { unit, label, credits, eur, free }.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| units | Yes | The per-unit pricing catalog. |
| currency | Yes | Billing currency (always EUR). |
| multipliers | Yes | Processing-mode multipliers applied on top of per-unit cost (e.g. { realtime: 1, batch: 0.5 }). |
| credits_per_eur | Yes | Credits per EUR (e.g. 1000 = €1). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds value by stating it is public and does not spend credits, which is consistent with the annotations. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured with clear headings (USE WHEN, NOT FOR, ARGS, RETURNS). Every sentence is informative and earns its place. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and an output schema (though not shown), the description fully covers what the agent needs: purpose, usage, and return format. The annotations provide safety context, making it complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameters and 100% schema coverage, baseline is 4. The description mentions 'ARGS: none' and explains the return structure (currency, credits_per_eur, etc.), which adds meaning beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Read' and identifies the resource 'Talonic's machine-readable credit pricing catalog'. It clearly distinguishes from sibling tools like talonic_get_balance and talonic_get_usage by stating its purpose is for estimating costs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly provides use cases ('USE WHEN: estimating cost'), non-use cases ('NOT FOR: balance/usage'), and names alternative tools. This leaves no ambiguity about when to invoke this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_get_runPoll a Spec runARead-onlyIdempotentInspect
Poll a Spec run started by talonic_run_spec: normalised status plus document-level progress and, for pipelines, per-phase progress.
USE WHEN: waiting for a run to finish — poll every 5–10 s; stop on completed or failed.
NOT FOR: reading the structured rows (talonic_get_run_results) or starting a run (talonic_run_spec).
ARGS: exactly one of pipeline_id (run_kind 'pipeline') or run_id (run_kind 'run'), from the RunEnvelope.
RETURNS: { run_kind, run_id, pipeline_id, spec_id, status, raw_status, input_count?, progress { total_documents, completed_documents, error_documents, phases?[] }, documents?[], error_message?, created_at, updated_at }.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No | From a run_kind 'run' envelope (/v1/run). | |
| pipeline_id | No | From a run_kind 'pipeline' envelope (/v1/pipelines). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly, idempotent, and non-destructive behavior, so the description does not need to repeat those. It adds useful context by explaining that status is normalised and that progress includes documents and pipeline phases, plus the terminal states to watch for. It does not add much on failure modes or rate limits, but the existing annotations cover the safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organised with clear labels for use, exclusions, arguments, and return shape. Every section contributes directly to correct selection and invocation, with no wasted words. The structured format makes it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description provides a detailed return shape including progress, optional fields, and error_message. It also explains the polling loop and terminal statuses, giving an agent everything needed to drive the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only lists two optional UUID parameters, so the description materially clarifies them by stating 'exactly one' and mapping each parameter to the right run_kind and envelope. This is critical usage information the schema itself does not encode, making the parameter semantics substantially clearer.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action, 'Poll a Spec run', and enriches it with the actual value returned: normalised status, document-level progress, and per-phase progress for pipelines. It also names sibling alternatives it is not, such as talonic_get_run_results and talonic_run_spec, so the agent can distinguish it immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit 'USE WHEN' condition with concrete polling cadence and terminal statuses, and an explicit 'NOT FOR' section naming the relevant alternatives. This leaves no ambiguity about when to choose this tool versus its siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_get_run_resultsRead a Spec run's rowsARead-onlyIdempotentInspect
Read a Spec run's structured rows — one row per document with the Spec's fields as clean values (held/pending-review cells serialize null), plus the column definitions.
USE WHEN: talonic_get_run reports completed (partial rows are also readable while processing).
NOT FOR: progress (talonic_get_run) or per-field provenance of a single value (include: ['provenance'] here, or talonic_field_values).
ARGS: exactly one of pipeline_id (run_kind 'pipeline', optionally with the envelope's run_id to scope to that submission) or run_id alone (run_kind 'run'); optional document_id (one document), include (['cells','provenance'] — heavier payload), limit (1–200, default 50), cursor.
RETURNS: { run_kind, status, columns[] of { field_key, display_name, data_type }, data[] of { document_id, filename, record_id, status ('complete'|'partial'|'error'|'processing'), completed_at, fields { field_key: value }, cells?, provenance? }, pagination, pending_review_count, links }.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Page size (default 50). | |
| cursor | No | Opaque cursor from pagination.next_cursor. | |
| run_id | No | The RunEnvelope's run_id. Alone: a run_kind 'run' submission (/v1/run). With pipeline_id: scopes the pipeline's rows to that submission. | |
| include | No | Extra per-field detail; heavier payload. | |
| document_id | No | Restrict to one document. | |
| pipeline_id | No | The RunEnvelope's pipeline_id (run_kind 'pipeline'); rows come from /v1/pipelines/{id}/results. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context: partial rows are readable while processing, held/pending-review cells serialize null, and the include parameter makes the payload heavier. It doesn't describe pagination behavior in depth, but the RETURNS section covers the response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured with clear USE WHEN / NOT FOR / ARGS / RETURNS sections. Every sentence earns its place: the first sentence defines the resource, the second gives the trigger condition, the third excludes alternatives, the fourth explains argument selection, and the fifth documents the return shape. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, idempotent tool with 100% schema coverage and a detailed RETURNS section, the description is complete. An agent knows exactly when to call it, which arguments to use, what the response looks like, and how it differs from siblings. The absence of an output schema is compensated by the explicit RETURNS block.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 6 parameters. The description adds meaning beyond the schema by explaining the run_kind distinction (pipeline_id for run_kind 'pipeline' vs run_id alone for run_kind 'run'), the scoping relationship between run_id and pipeline_id, and the semantic effect of include (heavier payload). This goes beyond the schema's per-parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Read a Spec run's structured rows — one row per document with the Spec's fields as clean values... plus the column definitions.' This clearly distinguishes it from siblings like talonic_get_run (progress) and talonic_field_values (per-field provenance).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use it ('USE WHEN: talonic_get_run reports completed...'), what it is not for ('NOT FOR: progress... or per-field provenance...'), and names the alternative tools. It also explains the ARGS selection logic (pipeline_id vs run_id) and optional include values.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_get_specGet a Spec's structureARead-onlyIdempotentInspect
Get one Spec's structure: identity and version state, the schema it materializes onto, nodes[] (the rail as authored, in editing order) and phases[] (the compiled execution plan, in run order — a validation checkpoint expands to one phase per gate, so the two lists differ on purpose), and fields[] (Spec field ↔ schema field).
USE WHEN: you need to explain what a run will do, confirm a Spec is published (version non-null) before talonic_run_spec, or map field names to keys.
NOT FOR: listing Specs (talonic_list_specs) or starting a run (talonic_run_spec).
ARGS: spec_id (UUID from talonic_list_specs); optional include_versions (adds versions[] — published versions newest first, each { version, content_hash, created_at, is_materialized }).
RETURNS: the Spec object { id, name, description, schema_id, version, materialized_version, materialized_at, field_count, node_count, schema, nodes[], phases[], fields[], links } plus optional versions[].
| Name | Required | Description | Default |
|---|---|---|---|
| spec_id | Yes | Spec UUID (from talonic_list_specs). | |
| include_versions | No | Also fetch the published versions list (adds `versions[]`). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds meaningful behavioral context beyond annotations: nodes[] is 'the rail as authored, in editing order' while phases[] is 'the compiled execution plan, in run order', with validation checkpoints expanding to one phase per gate, explaining why the lists differ. It also discloses that include_versions returns versions newest first with specific fields. This adds real value beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but tightly organized with labeled sections: USE WHEN, NOT FOR, ARGS, RETURNS. Each section earns its place and the most important scoping information is front-loaded. No filler or repetition of obvious details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the RETURNS section carries the full burden of describing the response shape, and it does so exhaustively ('{ id, name, description, schema_id, version, materialized_version, materialized_at, field_count, node_count, schema, nodes[], phases[], fields[], links }' plus optional versions[]). Combined with the usage guidance and parameter details, nothing an agent needs to correctly call this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters. The description nonetheless adds value: it specifies spec_id comes from talonic_list_specs and details exactly what include_versions adds ('published versions newest first, each { version, content_hash, created_at, is_materialized }'). This exceeds the schema's bare definitions, though the baseline is already solid.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Get one Spec's structure') and immediately clarifies scope: identity, version state, schema, nodes[], phases[], and fields[]. It actively distinguishes itself from siblings by naming talonic_list_specs and talonic_run_spec in the NOT FOR section. An agent can tell this apart from similarly named tools without inspecting their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit USE WHEN scenarios ('explain what a run will do', 'confirm a Spec is published before talonic_run_spec', 'map field names to keys') and explicit NOT FOR exclusions with named alternatives (talonic_list_specs, talonic_run_spec). This is exactly the when-to-use vs alternative guidance expected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_get_usageGet Talonic UsageARead-onlyInspect
Read the workspace's per-function credit consumption over a trailing window: where the credits actually went.
USE WHEN: the user asks what they have spent credits on, or you want to see which function (extraction, structuring, intelligence ops) dominates spend. NOT FOR: the remaining balance (use talonic_get_balance) or per-unit rates (use talonic_get_pricing). ARGS: days (optional, default 30, clamped 1-365). RETURNS: period_days, total_credits, and by_function[] — each { operation_type, operations, credits }, highest spend first.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Trailing window in days (default 30). |
Output Schema
| Name | Required | Description |
|---|---|---|
| by_function | Yes | Per-function breakdown, highest spend first. |
| period_days | Yes | Length of the reporting window in days. |
| total_credits | Yes | Total credits consumed across all functions in the window. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true and destructiveHint=false. The description adds value by detailing the return structure (period_days, total_credits, by_function) and the days parameter clamping. This extra context goes beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, well-structured with clear sections (lead sentence, USE WHEN, NOT FOR, ARGS, RETURNS). Every sentence is purposeful and front-loaded with key information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one optional parameter, read-only operation with annotations and output schema implied by the return description), the description fully covers when, why, and what to expect. It is complete for effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline 3. The description repeats the parameter's default and clamping, adding slight value over the schema description. However, no additional meaning beyond what the schema already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Read the workspace's per-function credit consumption' using a specific verb and resource. It distinguishes from siblings by explicitly stating it is for seeing where credits went, not for balance or pricing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description includes explicit 'USE WHEN' and 'NOT FOR' sections, directing the agent when to use this tool and providing alternatives (talonic_get_balance, talonic_get_pricing). This makes the usage context crystal clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_heartbeat_agent_taskHeartbeat Agent TaskAInspect
Extend the lease on a claimed Agent-stage task.
USE WHEN: processing may continue past lease_expires_at; heartbeat before expiry using the epoch from claim. NOT FOR: acquiring a task (use talonic_claim_agent_task) or submitting finished outputs (use talonic_submit_agent_task). ARGS: task_id and execution_epoch. Stale or foreign claims return HTTP 409.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | Agent task UUID. | |
| execution_epoch | Yes | Execution epoch returned by the successful claim. Stale epochs are rejected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It discloses the important behavioral outcome that stale or foreign claims return HTTP 409, which is useful beyond the annotations. The lease-extension mutation is consistent with readOnlyHint=false, though the description does not mention success response details or other failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, well-structured, and front-loaded with the core purpose. USE WHEN, NOT FOR, and ARGS/ failure behavior sections each add value without redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity two-parameter tool, the description covers purpose, usage conditions, alternatives, and failure behavior. It stops just short of perfect completeness by not stating what a successful heartbeat returns, but this is not essential for choosing or invoking the operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides a 100% covered description, so the baseline is a solid 3. The tool description raises that by explaining the execution_epoch comes from the claim, that heartbeats must happen before expiry, and that foreign claims are rejected.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Extend the lease') with a specific resource ('a claimed Agent-stage task'), making the operation immediately clear. It also distinguishes this tool from the sibling claim and submit tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit USE WHEN and NOT FOR guidance, including concrete sibling alternatives: talonic_claim_agent_task and talonic_submit_agent_task. This tells the agent exactly when the tool is appropriate and when it is not.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_heartbeat_decision_taskHeartbeat Decision TaskAInspect
Extend the lease on a claimed decision task, never past its sla_deadline_at.
USE WHEN: deciding may run past lease_expires_at; heartbeat before expiry using the epoch from claim. NOT FOR: acquiring a task (talonic_claim_decision_task) or finishing one (talonic_submit_decision_task / talonic_release_decision_task / talonic_fail_decision_task). ARGS: task_id and execution_epoch. A stale or foreign epoch returns HTTP 409: stop, discard the work, and re-list. AUTH: a tlnc_ key needs a per-app 'decide' grant; an OAuth connector session needs the apps:decide scope (consented at connect) and a live workspace role of senior_member or above. A 403 (decide_grant_required, insufficient_scope, insufficient_tier) names what is missing — tell the user, do not retry.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | Decision task UUID. | |
| execution_epoch | Yes | Execution epoch returned by the successful claim. Stale epochs are rejected with 409. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the sparse annotations by disclosing the 409 behavior for stale/foreign epochs, the prescribed reaction (stop, discard, re-list), and the auth failure modes with named error codes. It also specifies the lease extension limit relative to sla_deadline_at, which is useful runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is organized into labeled sections (USE WHEN, NOT FOR, ARGS, AUTH) that are easy to scan and each earns its place. It is dense but free of redundancy, with the most important purpose statement first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schema and meaningful auth/error nuance, the description covers when to use it, what it does, how parameters relate to the claim, what errors mean, and how to handle auth failures. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents both parameters well. The description adds practical meaning by saying 'using the epoch from claim' and clarifying that a foreign epoch also returns 409, plus the recovery action. This is a meaningful but not large increment over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and object ('Extend the lease on a claimed decision task') and adds a hard constraint ('never past its sla_deadline_at'). It also explicitly names sibling tools it is not for, so an agent can distinguish it from claim, submit, release, and fail at a glance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The USE WHEN line gives the exact condition for invoking the tool, and the NOT FOR line lists the specific alternative tools for different lifecycle stages. This is explicit routing guidance with no inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_invoke_agent_toolInvoke Platform Agent ToolARead-onlyInspect
Invoke ONE named tool from the platform's agent tool registry directly, with no model in the loop — you choose the arguments. This is how an external agent uses Talonic's retrieval and provenance while driving control flow itself (e.g. query_data for a read-only SQL SELECT over the extracted data, describe_data for the queryable field list, get_document_markdown to read a document's text).
READ-ONLY BY CONSTRUCTION: API-key credentials are restricted by the platform to the registry's read-only tools (capability data.read); write-capable registry tools are never invocable through this credential, so through an API-key credential this tool reads and never mutates workspace data.
Target API: Talonic agent tool registry — https://talonic.com/docs/api (POST /v1/agent/tools/{name}/invoke; input schemas from talonic_list_agent_tools).
USE WHEN: talonic_list_agent_tools showed a tool with can_invoke: true that does what you need. Pass exactly the args its input_schema declares.
NOT FOR: anything a dedicated talonic_* tool already does (prefer those — they are shaped for you).
ARGS: name (tool name), args (object matching the tool's input_schema), optional document_ids (hard scope for scope-aware tools).
RETURNS: { tool, result (the tool's parsed output), citations?, artifacts?, cards? }. Denied capabilities come back as an error naming the capability required; the platform re-checks every call.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Arguments matching the tool's input_schema. | |
| name | Yes | Tool name from talonic_list_agent_tools, e.g. `query_data`. | |
| document_ids | No | Restrict to these document ids (hard filter, enforced server-side). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It explains the read-only guarantee beyond the annotations: API-key credentials are restricted to registry tools with `data.read` capability, write-capable tools cannot be invoked this way, and the platform re-checks every call. It also discloses denied-capability error behavior and the shape of the response.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but well-structured with labeled ARGS, USE WHEN, NOT FOR, and RETURNS sections. It is front-loaded with the core behavior and stays focused; minor redundancy in the read-only paragraph keeps it from being maximally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a dynamic dispatch tool with no output schema, the description covers the API endpoint, credential constraints, argument contract, return shape, optional return fields, and error behavior. An agent has enough information to decide when to use it and how to interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaningful context: `args` must match the target tool's input_schema, `document_ids` is a hard scope for scope-aware tools, and concrete tool-name examples clarify expected values. This goes slightly beyond the schema without fully duplicating it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: 'Invoke ONE named tool from the platform's agent tool registry directly, with no model in the loop.' It names concrete examples like `query_data` and `get_document_markdown`, and differentiates from the dedicated talonic_* tools by saying those are preferred for their own jobs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit USE WHEN and NOT FOR guidance: use when `talonic_list_agent_tools` shows `can_invoke: true`, and avoid when a dedicated talonic_* tool already exists. It also tells the agent to pass exactly the `args` declared by the selected tool's `input_schema`.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_list_agent_tasksList Agent TasksARead-onlyInspect
List Agent-stage tasks visible to this Talonic workspace credential.
USE WHEN: looking for external-agent work to process; begin with status 'available'. NOT FOR: reading the immutable task payload (use talonic_get_agent_task) or taking a lease (use talonic_claim_agent_task). ARGS: optional status, limit, and cursor. RETURNS: metadata only plus pagination.next_cursor.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Page size (default 50). | |
| cursor | No | Opaque cursor from pagination.next_cursor. | |
| status | No | Optional task-status filter. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the read-only nature is known. The description adds that it 'RETURNS: metadata only plus pagination.next_cursor,' clarifying the scope of results and the absence of payload data, which goes beyond the annotation. This additional context is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: it starts with a clear purpose, then uses labeled sections (USE WHEN, NOT FOR, ARGS, RETURNS) to convey key information without redundancy. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description explicitly states the return value ('metadata only plus pagination.next_cursor'), filling that gap. It also covers scope, filters, pagination, and alternatives, making the tool behavior fully understandable for an agent. The tool is simple, and the description addresses all necessary aspects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with each parameter (limit, cursor, status) already described in the schema. The description's mention of 'optional status, limit, and cursor' adds no new meaning, though it does reference pagination.next_cursor which ties to the cursor parameter. Since the schema does the heavy lifting, a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'List Agent-stage tasks visible to this Talonic workspace credential,' specifying the exact action (list), the resource (agent tasks), and the scope (visible to workspace credential). It also distinguishes itself from siblings by noting NOT for reading payload or taking a lease, making it unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit USE WHEN guidance is provided: 'looking for external-agent work to process; begin with status available.' NOT FOR exclusions mention specific alternatives (talonic_get_agent_task and talonic_claim_agent_task), giving clear when-to-use vs when-not-to-use instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_list_agent_toolsList Platform Agent ToolsARead-onlyIdempotentInspect
List the platform's agent tool registry — every retrieval, provenance and analysis primitive the in-product Talonic agent runs on (find_data, describe_data, query_data for read-only SQL over the extracted data, get_document_markdown, workspace_overview, …) with its input schema and whether THIS credential may invoke it. Target API: https://talonic.com/docs/api (GET /v1/agent/tools).
USE WHEN: you want a capability talonic_* tools do not cover directly (e.g. SQL over the structured data, a workspace overview, cohort discovery) — list here, then call talonic_invoke_agent_tool with the tool name and its args. NOT FOR: discovering fields (talonic_list_fields / talonic_find_data) or documents (talonic_search) — those are shaped for you.
ARGS: only_invocable (default true — hide tools this key cannot run), include_schemas (default true — include each tool's JSON input schema).
RETURNS: { tools[] of { name, description, impact, capability, can_invoke, input_schema? }, invocable_count, totalCount }.
| Name | Required | Description | Default |
|---|---|---|---|
| only_invocable | No | Hide tools this credential cannot invoke. Default true. | |
| include_schemas | No | Include each tool's JSON input schema. Default true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds value beyond that: the credential-based can_invoke semantics, the read-only nature of the SQL primitives, and the explicit return shape. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured with USE WHEN / NOT FOR / ARGS / RETURNS sections and front-loaded purpose. Slightly verbose with the enumerated primitive examples, but each section earns its place and improves scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a registry-listing tool: purpose, usage routing, argument defaults, return format, and credential semantics are all covered. With no output schema present, the RETURNS section is essential and it is provided in detail. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so both parameters (only_invocable, include_schemas) are fully documented in the schema. The description restates them with defaults and meanings, which is redundant but confirms the semantics; baseline 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (the platform's agent tool registry) with concrete examples of what the registry contains. It distinguishes itself from siblings by explicitly naming talonic_list_fields/talonic_find_data and talonic_search as different discovery surfaces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit USE WHEN (capability not covered by talonic_* tools directly, then invoke via talonic_invoke_agent_tool) and NOT FOR (fields/documents discovery) sections, naming exact alternatives. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_list_decision_tasksList Decision TasksARead-onlyInspect
List one External-mode app's decision tasks (runs parked for an outside agent to decide), newest first.
USE WHEN: looking for decisions to make for an app; begin with status 'available'. This is the polling alternative to the app.decision_task.offered webhook. NOT FOR: Agent-stage document tasks (use talonic_list_agent_tasks), reading a task's input package (claim it, then talonic_read_decision_package), or taking a lease (talonic_claim_decision_task). ARGS: app_id, optional status, limit, cursor. RETURNS: task metadata only (id, run_id, status, execution_epoch, lease and sla_deadline_at timing) plus pagination.next_cursor. AUTH: a tlnc_ key needs a per-app 'decide' grant; an OAuth connector session needs the apps:decide scope (consented at connect) and a live workspace role of senior_member or above. A 403 (decide_grant_required, insufficient_scope, insufficient_tier) names what is missing — tell the user, do not retry.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Page size (default 50). | |
| app_id | Yes | App UUID whose decision-task inbox to read. | |
| cursor | No | Opaque cursor from pagination.next_cursor. | |
| status | No | Optional task-status filter. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and non-destructive behavior, so the bar is lower, but the description still adds substantial context: newest-first ordering, return shape limited to task metadata plus pagination cursor, auth requirements, and the 403 error outcome including the instruction not to retry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Though longer than typical, every section earns its place: purpose, when to use, when not to use, parameter/return summary, and auth/error handling. The labeled sections make the content scannable and front-load the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description tells the agent exactly what will be returned and what won't be. It also covers auth failure modes and provides routing to sibling tools, leaving no obvious gap for an agent to call this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds useful usage semantics beyond the schema by recommending starting with status 'available' and explaining that cursor comes from pagination.next_cursor, which helps an agent use the parameters correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: list one app's decision tasks, newest first. It makes the scope explicit (External-mode, runs parked for an outside agent) and differentiates from sibling tools by referencing talonic_list_agent_tasks and related decision-task tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear USE WHEN starting point (status 'available') and an explicit polling-vs-webhook relationship. The NOT FOR section names exact sibling alternatives and sequences them, e.g. claim first, then read the decision package, so an agent knows exactly when to choose other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_list_fieldsList Registry FieldsARead-onlyIdempotentInspect
List the workspace's Field Registry — the canonical concepts Talonic has discovered across every ingested document, each with a stable id, maturity level, data type, synonyms and occurrence count.
USE WHEN: you need to know WHAT data exists before querying it, want to pick the right concept for a question, or need the exact field id for talonic_get_field / talonic_field_values. NOT FOR: locating a specific document (talonic_search) or filtering documents by a value (talonic_filter).
ARGS: search (case-insensitive contains on name), maturity (core | proven | candidate — prefer core/proven for anything you will build on), include_superseded (default false: rows merged into another concept are hidden so you never see two ids for one concept), limit, cursor.
RETURNS: data[] of { id, canonical_name, display_name, data_type, maturity, tier, synonyms, description, occurrence_count, superseded_by, links } plus cursor pagination.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Page size (default 20, max 100). | |
| cursor | No | Opaque cursor from pagination.next_cursor. | |
| search | No | Case-insensitive contains match on canonical_name / display_name. | |
| maturity | No | Filter by maturity: `core` (universal, fully trusted — safe to build on), `proven` (recurring, stable id), `candidate` (newly discovered, may still be merged or renamed). | |
| include_superseded | No | Include rows merged into a survivor (they carry `superseded_by`). Default false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the call as readOnly, idempotent, and non-destructive; the description adds meaningful behavioral detail beyond that: default include_superseded behavior, why merged rows are hidden to avoid duplicate ids, and the returned row shape plus cursor pagination. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description front-loads the purpose and then uses short labeled sections for use conditions, exclusions, args, and returns. Every line earns its place with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the RETURNS line supplies the full result field set and describes cursor pagination. Combined with explicit sibling routing and annotation-covered safety, there is no missing context needed to invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The ARGS section adds non-redundant guidance by recommending core/proven maturity when building and explaining the duplicate-prevention rationale for include_superseded. It still largely mirrors schema descriptions, so it stops short of a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence uses a specific verb 'List' and a precise object, the workspace Field Registry, and enumerates what each entry contains (stable id, maturity, data type, synonyms, occurrence count). It also distinguishes itself by naming siblings it is not (talonic_search, talonic_filter) and providing use cases for related get_field/field_values tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The USE WHEN block gives concrete scenarios: knowing what data exists before querying it, picking a concept, or needing an exact field id for talonic_get_field / talonic_field_values. The NOT FOR section explicitly points to talonic_search and talonic_filter, so an agent can quickly route among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_list_schemasList Talonic SchemasARead-onlyInspect
List the saved schemas in the workspace as compact summaries (id, short_id, name, description, version, field_count).
USE WHEN: 'what schemas do I have', or to find a reusable schema before extracting.
NOT FOR: a one-off extraction with an inline schema (call talonic_extract directly).
ARGS: none.
RETURNS: data[] of schema summaries. Full field definitions are omitted here — read the talonic://schemas resource for those. Pass a schema's id/short_id to talonic_extract as schema_id.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| data | Yes | Saved schemas in the workspace. |
| pagination | No | Cursor-based pagination metadata. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is clear. The description adds value by stating it returns data[] of schema summaries, that full field definitions are omitted, and that there are no arguments. This provides behavioral detail beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with three clear sections: purpose, usage guidance, and returns/args. No unnecessary information. Every sentence serves a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple listing tool with no parameters and minimal complexity, the description covers all necessary information: what it returns, how to use it, and where to find additional details (resource). Output schema exists but the description adequately summarizes the return.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, and the description confirms 'ARGS: none.' The schema coverage is effectively 100%. With 0 parameters, the baseline is 4, and the description meets that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb ('List') and resource ('saved schemas'), and details the output fields (id, short_id, name, description, version, field_count). It distinguishes from sibling tools by noting that full field definitions are omitted and directing users to the talonic://schemas resource for those.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use ('what schemas do I have', find a reusable schema before extracting) and when not to (one-off extraction with inline schema, directing to talonic_extract). This provides clear alternatives and context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_list_specsList SpecsARead-onlyIdempotentInspect
List the workspace's Specs — the configured pipelines (rail + fields) that talonic_run_spec executes. Each row: id, name, description, schema_id (the schema a run materializes onto — a DIFFERENT id from the Spec's), version / materialized_version (null = never published), field_count, node_count, timestamps.
USE WHEN: the user wants to run 'their pipeline' / 'the invoice Spec', or you need a spec_id for talonic_run_spec / talonic_get_spec.
NOT FOR: ad-hoc extraction schemas (talonic_list_schemas) or discovering fields (talonic_list_fields).
ARGS: optional search (name contains, case-insensitive), limit (1–100, default 20), cursor (from pagination.next_cursor), order (asc|desc by updated_at).
RETURNS: { data[] of { id, name, description, schema_id, version, materialized_version, materialized_at, field_count, node_count, created_at, updated_at, links }, pagination { total, limit, has_more, next_cursor } }.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Page size (default 20). | |
| order | No | Sort by updated_at (default desc). | |
| cursor | No | Opaque cursor from pagination.next_cursor. | |
| search | No | Case-insensitive contains match on the Spec name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds meaningful behavioral context beyond that: schema_id is a different id from the Spec's, version/materialized_version null means never published, and it reveals the exact returned fields and pagination shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear USE WHEN, NOT FOR, ARGS, and RETURNS sections. It packs necessary detail into compact labeled blocks without redundancy, and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description fully specifies the return shape including pagination fields. It also covers all parameters, key behavioral nuances, and sibling differentiators, leaving an agent with everything needed to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 100% coverage, so baseline is 3. The description adds value by explaining the purpose of each parameter in the context of the tool (e.g., cursor from pagination.next_cursor), and by specifying the order semantics (by updated_at). It does not merely repeat the schema verbatim.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses a specific verb ('List') and resource ('the workspace's Specs') and immediately clarifies that these are configured pipelines executed by talonic_run_spec. The row-level detail and explicit NOT FOR section distinguish it clearly from sibling tools like talonic_list_schemas and talonic_list_fields.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The USE WHEN section explicitly states the conditions for calling this tool, including needing a spec_id for talonic_run_spec or talonic_get_spec. It also names alternatives it is NOT for, giving the agent clear routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_read_decision_packageRead Decision PackageARead-onlyInspect
Read one page of a claimed decision task's frozen input package: the records the decision must be made from, with their provenance locators. Only the current claimant may read it; each page read is journaled onto the run.
USE WHEN: after a successful claim, walking pages from package.first_cursor while pagination.has_more is true. The first page also carries the source documents list. A mining_round package is one record { system_prompt, first_turn, tools }. NOT FOR: unclaimed tasks (HTTP 409; claim first) or listing tasks (talonic_list_decision_tasks). ARGS: task_id, optional cursor and limit (1 to 2000). Copy evidence locators verbatim from these records for the submit. AUTH: a tlnc_ key needs a per-app 'decide' grant; an OAuth connector session needs the apps:decide scope (consented at connect) and a live workspace role of senior_member or above. A 403 (decide_grant_required, insufficient_scope, insufficient_tier) names what is missing — tell the user, do not retry.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Records per page, 1 to 2000. Defaults to the platform page size (package.page_size from the claim). | |
| cursor | No | Opaque package cursor: package.first_cursor from the claim, then pagination.next_cursor. Omit for the first page. | |
| task_id | Yes | Decision task UUID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotation Contradiction: the annotations declare readOnlyHint=true, but the description says 'each page read is journaled onto the run,' which is a state change. This directly contradicts the read-only hint. The description is otherwise transparent about auth errors, but the contradiction is severe because it makes the tool's side effects falsely appear harmless.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured with clear USE WHEN, NOT FOR, ARGS, and AUTH sections, with the core purpose front-loaded. Every section carries necessary operational detail and there is no filler or redundant repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a paginated read tool with no output schema, the description fully compensates: it describes what the page contains, the first page's source document list, the special mining_round package shape, pagination flow, and auth/error behavior. An agent has enough information to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 even though the description adds little. The ARGS section restates parameter names without adding new semantic meaning beyond the schema; 'Copy evidence locators verbatim' is a usage note rather than parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Read one page of a claimed decision task's frozen input package' and explicitly distinguishes itself from listing tasks with 'NOT FOR: ... listing tasks (talonic_list_decision_tasks).' This makes the tool's unique purpose immediately clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage timing: 'after a successful claim, walking pages from package.first_cursor while pagination.has_more is true.' It also states exclusions (unclaimed tasks, listing tasks) and auth requirements, so the agent knows exactly when and when not to call this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_release_decision_taskRelease Decision TaskAInspect
Release a claimed decision task back to 'available' without deciding it; the next claim bumps the epoch.
USE WHEN: you cannot finish within the lease or SLA but another claimant could decide it. NOT FOR: declaring the task undecidable (talonic_fail_decision_task, which raises a review and applies the app's fallback) or keeping the lease (talonic_heartbeat_decision_task). ARGS: task_id and execution_epoch. Stale epoch is HTTP 409. AUTH: a tlnc_ key needs a per-app 'decide' grant; an OAuth connector session needs the apps:decide scope (consented at connect) and a live workspace role of senior_member or above. A 403 (decide_grant_required, insufficient_scope, insufficient_tier) names what is missing — tell the user, do not retry.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | Decision task UUID. | |
| execution_epoch | Yes | Execution epoch returned by the successful claim. Stale epochs are rejected with 409. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only signal readOnlyHint=false and destructiveHint=false, so the description carries the behavioral burden and does so thoroughly: the non-decision release semantics, epoch-bump side effect, HTTP 409 on stale epoch, and detailed auth failure behavior (403 names what is missing, 'tell the user, do not retry'). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, then organized into scannable labeled sections (USE WHEN, NOT FOR, ARGS, AUTH). Every sentence earns its place given the tool's complex auth and epoch semantics; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with complex auth, epoch semantics, and sibling ambiguity, the description covers what it does, when and when-not to use it, parameter behavior, error codes, and authentication requirements. No output schema exists, but for a release action this is adequate. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented in the input schema, including the stale-epoch-409 behavior on execution_epoch. The description's 'ARGS: task_id and execution_epoch. Stale epoch is HTTP 409' adds marginal reinforcement but little new meaning beyond what the schema states, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Release a claimed decision task back to available without deciding it'. It also names the side effect (next claim bumps the epoch) and explicitly distinguishes itself from talonic_fail_decision_task and talonic_heartbeat_decision_task. An agent can select it correctly without opening sibling schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit USE WHEN condition (cannot finish within lease/SLA but another claimant could decide) and a NOT FOR clause naming the two alternatives (fail, which raises a review; heartbeat, which keeps the lease). This is model guidance for when and when-not, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_request_uploadRequest File UploadAInspect
Get a browser upload link the user opens to add a file to their workspace. Returns the link plus a pre-allocated document_id.
USE WHEN: the user wants to upload a document and you cannot pass it directly — hosted/sandboxed clients (ChatGPT, Claude.ai) or files too large for tool-call arguments. NOT FOR: a document already in the workspace (use its document_id) · a file already on a public URL (use file_url on talonic_extract). ARGS: filename (with extension). RETURNS: upload_url, document_id, expires_at. After the user uploads, poll talonic_get_document on that document_id until status is 'completed', then call talonic_extract. If status becomes ocr_failed, extraction_failed, or error, stop polling and report the failure to the user.
| Name | Required | Description | Default |
|---|---|---|---|
| filename | Yes | The name of the file being uploaded, including extension (e.g. 'invoice.pdf'). Used to pre-allocate the document and infer MIME type. |
Output Schema
| Name | Required | Description |
|---|---|---|
| expires_at | Yes | ISO 8601 timestamp when the upload link expires. |
| upload_url | Yes | URL the user should open in their browser to drop the file. |
| document_id | Yes | The pre-allocated document ID. Use with talonic_get_document to poll status, and with talonic_extract once uploaded. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses full workflow: returns upload_url, document_id, expires_at; after user uploads, poll talonic_get_document until status 'completed', then call talonic_extract; handles failure statuses. Annotations (readOnlyHint=false) are not contradicted and are supplemented by this detailed behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with sections: main description, USE WHEN, NOT FOR, ARGS, RETURNS. It is front-loaded with key information, and every sentence adds value. No unnecessary words or repetitions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (one parameter, no enums, output schema exists), the description completely covers the input, output, and subsequent steps. It provides a workflow for polling and error handling, which is essential for an agent to use the tool correctly. Sibling tools are listed for context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes filename with minLength and a brief description. The description adds that filename should include extension and is used to pre-allocate and infer MIME type, providing extra meaning beyond the schema. With 100% schema coverage, baseline is 3; added value justifies 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get a browser upload link the user opens to add a file to their workspace' and specifies it returns both the link and a pre-allocated document_id. It uses specific verbs and resources, and the context distinguishes this from sibling tools like talonic_extract.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'USE WHEN' and 'NOT FOR' sections provide clear guidance on when to use this tool versus alternatives (e.g., 'hosted/sandboxed clients' or 'files too large'), and when not to (e.g., document already in workspace or file on public URL). The description also directs the agent to talonic_extract for public URLs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_run_specRun a Spec pipelineAInspect
Run a Spec — the customer's configured pipeline — over documents, in one call. Two inputs: document_ids (documents already in the workspace; for a new file, first talonic_request_upload → poll talonic_get_document until completed) OR file_urls (public https files, max 20; Talonic ingests them first). Consumes credits.
USE WHEN: the user wants to 'run the invoice pipeline on these documents', process files through their Spec, or produce the Spec's structured rows.
NOT FOR: one-off extraction with an ad-hoc schema (talonic_extract), or checking progress (talonic_get_run) / reading rows (talonic_get_run_results).
ARGS: spec_id (talonic_list_specs); exactly one of document_ids[] (1–500) or file_urls[] (1–20, https); optional name, pipeline_mode (new default | append to the Spec's existing pipeline); batch_id and flat metadata only with file_urls.
RETURNS: RunEnvelope { run_kind ('pipeline'|'run'), run_id, pipeline_id, spec_id, status ('processing'|'completed'|'failed'), raw_status, input_count, documents?, message?, links }. Then poll talonic_get_run with the pipeline_id when run_kind is 'pipeline', or with the run_id when it is 'run', every 5–10 s until status is completed/failed, then talonic_get_run_results.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Display name for the run. | |
| spec_id | Yes | Spec UUID (talonic_list_specs). | |
| batch_id | No | Caller grouping key (file_urls path only). | |
| metadata | No | Flat caller tags stamped on every ingested document (file_urls path only). | |
| file_urls | No | Public https file URLs (1–20). Mutually exclusive with document_ids. | |
| document_ids | No | Workspace document ids (1–500). Mutually exclusive with file_urls. | |
| pipeline_mode | No | `new` (default) or `append` to the Spec's existing pipeline. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate readOnlyHint=false and destructiveHint=false. The description adds critical side-effect context: it 'Consumes credits,' ingests file_urls first, returns a RunEnvelope with run_kind that dictates the polling target, and instructs to poll every 5–10s until terminal status. This goes well beyond the structured annotations and fully discloses the asynchronous, credit-consuming behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized with clear sections (intro, inputs, USE WHEN, NOT FOR, ARGS, RETURNS). Every sentence earns its place; the core purpose is front-loaded, and the rest is tightly packed without fluff. It's long but justified by the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 7 parameters, two input paths, and a non-trivial return/polling contract, the description is remarkably complete. It explains how to obtain spec_id (talonic_list_specs), how to handle new files via the upload flow, the exact return structure, and the polling logic. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
While schema coverage is 100%, the description adds meaningful semantics beyond the schema: it emphasizes 'exactly one of document_ids or file_urls,' spells out the limits (1–500 and 1–20), and clarifies that batch_id and metadata are only valid with file_urls. It also explains pipeline_mode ('new' vs 'append' to existing pipeline). This adds real value over the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb+resource: 'Run a Spec — the customer's configured pipeline — over documents, in one call.' It explicitly differentiates from siblings by naming what it is NOT for (talonic_extract, talonic_get_run, talonic_get_run_results) and clarifies the two input modes. An agent can immediately grasp the tool's role and scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a dedicated 'USE WHEN' section with concrete user intents (e.g., 'run the invoice pipeline on these documents') and a 'NOT FOR' list that names alternatives. It also gives conditional guidance on choosing document_ids (with the prerequisite upload flow) vs file_urls, and explains the polling flow after the call. This is explicit and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_save_schemaSave Talonic SchemaAInspect
Save a reusable schema to the workspace for use across future extractions.
USE WHEN: the user confirms a schema/template they want to reuse across documents.
NOT FOR: a single one-off extraction (pass the schema inline to talonic_extract instead).
ARGS: name; definition — a JSON Schema ({type:'object',properties:{...}}) or a flat {field:'type'} map.
RETURNS: the saved schema with id and short_id. Pass either to talonic_extract as schema_id.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Human-readable name for the schema, e.g. 'Standard Invoice'. | |
| definition | Yes | Schema definition. Most reliable: full JSON Schema {type:'object', properties:{...}}. Also accepted: a flat key-type map {field_name:'string', amount:'number'} which the API normalises. | |
| description | No | Optional description of what this schema extracts and when to use it. |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | UUID of the newly saved schema. |
| name | Yes | |
| links | No | |
| version | No | Schema version (1 for new schemas; increments on update). |
| short_id | No | Human-readable short id (SCH-XXXXXXXX). |
| created_at | No | |
| definition | No | Final schema definition as stored, normalised by the API. |
| updated_at | No | |
| description | No | Schema description, or null when the schema was saved without one. The API explicitly maps the absent case to null (see SchemaResponse in openapi.yaml). |
| field_count | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate destructiveHint=false, which aligns with saving non-destructive. Description adds that it saves to workspace and returns id/short_id. Could mention if same name overwrites, but overall sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise: one sentence purpose, then usage, then args, then returns. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
All necessary context provided given schema coverage and output description. Parameter count, required fields, and return value all addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 100% coverage. Description adds value by explaining definition accepts JSON Schema or flat map, and that return includes id and short_id for later use.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb 'Save' with specific resource 'a reusable schema to the workspace'. Distinguishes from sibling talonic_extract by stating it's for reuse across documents, not one-off extractions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states 'USE WHEN' and 'NOT FOR' with direct alternative (talonic_extract). Provides clear context for when to choose this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_searchSearch Talonic WorkspaceARead-onlyInspect
Find documents, fields, schemas, or sources in the workspace. One call returns ranked results across all types.
MATCHING IS LITERAL KEYWORD, not semantic. Query with ONE short SINGULAR term or an exact filename: 'invoice', 'bank statement', 'sample-invoice.pdf'. Sentences ('documents related to invoices') and plurals ('invoices') return empty. If a search comes back empty, retry with a shorter singular keyword before concluding the workspace has nothing.
USE WHEN: the user names or describes a document without an id, or you need a document_id or a filterable field name before extract / to_markdown / get_document / filter.
NOT FOR: structured field-value filters like 'amount > 1000' (use talonic_filter).
ARGS: query (short literal keyword); optional limit.
RETURNS: documents[], fields[]/fieldMatches[] (only filterable: true entries work in talonic_filter), schemas[], sources[]. Use the id from documents[] to act on a named file.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum results per entity type. Default: 5. Increase for broader exploration. | |
| query | Yes | ONE short, SINGULAR keyword or an exact filename — 'invoice', 'insurance certificate', 'sample-invoice.pdf'. Matching is literal: sentences and plurals return empty. |
Output Schema
| Name | Required | Description |
|---|---|---|
| hint | No | Present only when nothing matched: actionable guidance to retry with a shorter singular keyword. |
| fields | Yes | Field-registry entries matching the query. filterable: true entries are usable with talonic_filter. |
| schemas | Yes | Saved schemas matching the query. |
| sources | Yes | Source connections matching the query. |
| documents | Yes | Documents matching the query. |
| fieldMatches | Yes | Field-level matches with a filterable flag indicating whether the entry can drive talonic_filter. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false. Description adds critical behavioral details: matching is literal keyword, sentences/plurals return empty, and fields must have filterable:true to work with talonic_filter. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Efficient structure: first sentence states purpose, then details matching behavior, then usage guidelines, then parameter description, then returns listing. Every sentence serves a purpose. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists, the description is complete enough. It covers purpose, behavior (literal matching, empty results), usage constraints, return types with actionable ids, and integrates with sibling tools. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. Description adds value by clarifying that query must be a short singular keyword or exact filename, and gives examples like 'invoice' vs 'invoices'. Also states limit default is 5. This utility boosts the score above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs and resources: 'Find documents, fields, schemas, or sources'. It clearly distinguishes from siblings by noting that this tool is for keyword search across all types, while talonic_filter is for structured filters.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit conditions: 'USE WHEN: the user names or describes a document without an id... NOT FOR: structured field-value filters (use talonic_filter).' Also provides retry advice for empty results, which is actionable context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_submit_agent_taskSubmit Agent TaskAInspect
Submit declared output fields for a claimed Agent-stage task and resume the parked document.
USE WHEN: processing is complete and every required output in output_contract is ready. NOT FOR: undeclared fields or partial lease maintenance (use talonic_heartbeat_agent_task). ARGS: task_id, execution_epoch, outputs keyed exactly by declared field key, and optional summary. The platform validates all fields and types transactionally before writing anything.
| Name | Required | Description | Default |
|---|---|---|---|
| outputs | Yes | Output field key to { value, confidence?, reasoning? }. Use only declared fields. | |
| summary | No | Optional result summary, up to 4,000 characters. | |
| task_id | Yes | Agent task UUID. | |
| execution_epoch | Yes | Execution epoch returned by the successful claim. Stale epochs are rejected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral context beyond the annotations: outputs are validated transactionally before anything is written, the execution epoch must come from a successful claim, stale epochs are rejected, and the document is resumed. This is strong for a write tool even though the annotations do not signal danger or contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core purpose, and organized with 'USE WHEN', 'NOT FOR', and 'ARGS' sections. Every sentence contributes actionable guidance with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description covers the core workflow context: eligibility, arguments, validation behavior, contract enforcement, and resume semantics. It stops short of describing the response or potential failure modes after submission, but the key operational context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents all four parameters, so the baseline is 3. The description adds semantically valuable guidance by clarifying that outputs must be keyed exactly by the declared field key, that the optional summary is allowed, and that execution_epoch comes from a successful claim, which enriches the schema-only understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Submit') and resource ('declared output fields for a claimed Agent-stage task') and explains the resulting effect ('resume the parked document'). It also explicitly distinguishes itself from the sibling tool talonic_heartbeat_agent_task via the 'NOT FOR' clause.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit 'USE WHEN' and 'NOT FOR' guidance, stating this tool is appropriate when processing is complete and every required output is ready, and inappropriate for undeclared fields or partial maintenance. It even names the alternative tool, making the boundaries unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_submit_decision_taskSubmit Decision TaskAInspect
Submit the decision for a claimed decision task; the platform verifies it transactionally and resumes the run.
USE WHEN: the decision is final. outcome must satisfy output_contract from the claim (plain JSON Schema, or the verdict_matrix / record_set envelope), evidence lists the package locators relied on (verbatim; [] only if the app allows unevidenced decisions), rationale is a short summary. Rejections are HTTP 422 with the reason and change nothing; the task stays claimed under your epoch, so fix and resubmit before the lease ends. NOT FOR: giving the task back undecided (talonic_release_decision_task) or declaring it undecidable (talonic_fail_decision_task). Stale epoch is HTTP 409. ARGS: task_id, execution_epoch, outcome, evidence, rationale, optional confidence and service_version. AUTH: a tlnc_ key needs a per-app 'decide' grant; an OAuth connector session needs the apps:decide scope (consented at connect) and a live workspace role of senior_member or above. A 403 (decide_grant_required, insufficient_scope, insufficient_tier) names what is missing — tell the user, do not retry.
| Name | Required | Description | Default |
|---|---|---|---|
| outcome | Yes | The decision, shaped by output_contract from the claim: for a plain JSON Schema contract the object that schema validates; for a verdict_matrix contract { subjects: [{ subject_key, rule_outcomes, auto?, detail? }] }; for a record_set contract its fields-and-rows envelope. | |
| task_id | Yes | Decision task UUID. | |
| evidence | Yes | Provenance locators the decision relied on, each copied verbatim from the input package (or a db:/ref: locator the app holds a read grant on). Required; [] is accepted only when the app allows unevidenced decisions. | |
| rationale | Yes | Short decision rationale (a summary, never private chain-of-thought), up to 4,000 characters. | |
| confidence | No | Optional confidence from 0 to 1. | |
| execution_epoch | Yes | Execution epoch returned by the successful claim. Stale epochs are rejected with 409. | |
| service_version | No | Optional build identifier of the deciding service; recorded as decided_by.label. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false and destructiveHint=false, so the description carries the full burden. It discloses transactional verification, state changes on rejection (nothing changes, task stays claimed), stale epoch 409, and detailed auth requirements (tlnc_ key with 'decide' grant, OAuth with apps:decide scope, senior_member role). No contradiction with annotations; this is thorough behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but well-organized into clear sections (purpose, USE WHEN, NOT FOR, ARGS, AUTH). Every sentence adds necessary context, and the most critical information (purpose and when to use) is front-loaded. It is structured and free of fluff, though slightly verbose for a single tool description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 7 parameters, nested objects, and no output schema, the description covers usage conditions, auth requirements, error codes, and retry behavior. It does not describe the success response format, but that is not essential given the absence of an output schema and the transactional nature described. Overall, it provides sufficient context for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter is already documented in the input schema. The description repeats key constraints (e.g., outcome must satisfy output_contract, evidence verbatim, [] only if allowed) but adds no new information beyond what the schema descriptions already provide. It does not compensate for any gaps, but none exist, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'submit' and the resource 'decision task', explains the platform verifies transactionally and resumes the run, and explicitly differentiates from sibling tools by listing NOT FOR cases with their exact names. This leaves no ambiguity about what the tool does or how it differs from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a dedicated USE WHEN section detailing the conditions (final decision, outcome satisfying output_contract, evidence verbatim, rationale short) and a NOT FOR section naming the two sibling tools to avoid. Also explains error handling (422, 409) and the retry behavior, giving explicit guidance on when to call and when to stop.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talonic_to_markdownDocument to MarkdownAInspect
Get the OCR-converted markdown text of a document.
USE WHEN: the user wants the full text — 'what does it say', summarise, or translate a document.
NOT FOR: specific structured fields (use talonic_extract with a schema).
BY NAME: if the user names a file, call talonic_search first to get its document_id, then call this.
ARGS: prefer document_id (a workspace doc — one cheap call). Otherwise file_url, or file_data+filename for small local files — provide exactly one. A file input ingests the document first and consumes credits; document_id does not.
RETURNS: document_id and markdown (the full text).
| Name | Required | Description | Default |
|---|---|---|---|
| file_url | No | URL to a document file. The Talonic API fetches it server-side. | |
| filename | No | Original filename including extension, e.g. 'invoice.pdf'. Used to infer MIME type when uploading via `file_data`. Required when `file_data` is provided. | |
| file_data | No | Base64-encoded file bytes. Recommended path when the agent already has the file in memory (e.g., the user attached a PDF to the conversation). Pair with `filename` so MIME type can be inferred. | |
| file_path | No | Local path to a document file. Only works if the MCP server has read access to that path. In sandboxed chat clients (Claude Desktop, Cowork) use `file_data` instead. | |
| document_id | No | The Talonic document id whose markdown you want. Get this from a previous talonic_extract or talonic_search response. |
Output Schema
| Name | Required | Description |
|---|---|---|
| cost | No | Per-call cost and post-call balance from the underlying extract step, parsed from the X-Talonic-* response headers. `null` when the document was already ingested (document_id path) and no extract call ran. Not always present on legacy clients. |
| markdown | Yes | OCR-converted markdown text content of the document. |
| document_id | Yes | ID of the document the markdown was extracted from. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnlyHint=false, destructiveHint=false, openWorldHint=true), description adds that file inputs ingest and consume credits while document_id does not, and different input methods have different costs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured with sections, front-loaded with main purpose. Every sentence adds value. No redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given output schema exists, description covers return values (document_id and markdown). Addresses complex input choices and usage scenarios. Complete for tool complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with good descriptions. Description adds hierarchy: prefer document_id, then file_url, then file_data+filename for small local files. Provides selection guidance beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it gets OCR-converted markdown text of a document. Uses verb 'Get' and specifies resource. Distinguishes from sibling talonic_extract for structured fields.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit sections: USE WHEN for full text/summarize/translate, NOT FOR structured fields, BY NAME instructs to first call talonic_search. Provides clear decision rules.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
21 tool updates
- Removed
talonic_ap_bills - Removed
talonic_ap_config - Removed
talonic_ap_issues - Removed
talonic_ap_netsuite_post - Removed
talonic_ap_netsuite_preview - Removed
talonic_ap_netsuite_runs - Removed
talonic_ap_overview - Removed
talonic_ap_save_config - Removed
talonic_contracts_contract - Removed
talonic_contracts_decide - Removed
talonic_contracts_import - Removed
talonic_contracts_import_candidates - Removed
talonic_contracts_import_status - Removed
talonic_contracts_issues - Removed
talonic_contracts_key_date - Removed
talonic_contracts_merge - Removed
talonic_contracts_register - Removed
talonic_contracts_reread - Removed
talonic_contracts_set_term - Removed
talonic_contracts_upcoming - Removed
talonic_contracts_update_document
57 tool updates
- First observed
talonic_ap_bills - First observed
talonic_ap_config - First observed
talonic_ap_issues - First observed
talonic_ap_netsuite_post - First observed
talonic_ap_netsuite_preview - First observed
talonic_ap_netsuite_runs - First observed
talonic_ap_overview - First observed
talonic_ap_save_config - First observed
talonic_ask - First observed
talonic_claim_agent_task - First observed
talonic_claim_decision_task - First observed
talonic_contracts_contract - First observed
talonic_contracts_decide - First observed
talonic_contracts_import - First observed
talonic_contracts_import_candidates - First observed
talonic_contracts_import_status - First observed
talonic_contracts_issues - First observed
talonic_contracts_key_date - First observed
talonic_contracts_merge - First observed
talonic_contracts_register - First observed
talonic_contracts_reread - First observed
talonic_contracts_set_term - First observed
talonic_contracts_upcoming - First observed
talonic_contracts_update_document - First observed
talonic_extract - First observed
talonic_fail_decision_task - First observed
talonic_field_values - First observed
talonic_filter - First observed
talonic_find_data - First observed
talonic_get_agent_task - First observed
talonic_get_answer - First observed
talonic_get_balance - First observed
talonic_get_document - First observed
talonic_get_field - First observed
talonic_get_pricing - First observed
talonic_get_run - First observed
talonic_get_run_results - First observed
talonic_get_spec - First observed
talonic_get_usage - First observed
talonic_heartbeat_agent_task - First observed
talonic_heartbeat_decision_task - First observed
talonic_invoke_agent_tool - First observed
talonic_list_agent_tasks - First observed
talonic_list_agent_tools - First observed
talonic_list_decision_tasks - First observed
talonic_list_fields - First observed
talonic_list_schemas - First observed
talonic_list_specs - First observed
talonic_read_decision_package - First observed
talonic_release_decision_task - First observed
talonic_request_upload - First observed
talonic_run_spec - First observed
talonic_save_schema - First observed
talonic_search - First observed
talonic_submit_agent_task - First observed
talonic_submit_decision_task - First observed
talonic_to_markdown
Publisher details
- Operator
- Talonic · Publisher source
- Operator website
- https://Talonic.com
- Vendor relationship
- Not applicable
- Documentation
- https://github.com/talonicdev/talonic-mcp · Publisher source
- Trust center
- Not applicable
- Restrictions
- Not applicable
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables brand visibility monitoring across major AI platforms like ChatGPT, Claude, Gemini, and Perplexity. It allows users to track visibility scores, analyze competitor data, and receive actionable insights to improve AI-generated brand recommendations.1624 npm1MIT
- AlicenseCqualityBmaintenanceCompetitor Monitor AI - MCP server providing AI-powered tools and automation by MEOK AI Labs117 npmMIT
- AlicenseAqualityCmaintenanceRevnuvo Company Intelligence tells AI agents what changed at a company, with evidence. It observes company websites, technologies, and DNS over time and returns timestamped, confidence-aware changes, signals, and monitoring.9MIT

industrylens-mcpofficial
AlicenseNot gradedqualityBmaintenanceBrowse IndustryLens's published competitive-intelligence reports and head-to-head competitor comparisons from any AI agent — real, source-backed data.MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.