Skip to main content
Glama

knowmind_upload_document

Ingest long-form text as a document to split it into chunks, generate embeddings, and add graph nodes, replacing any existing document with the same title.

Instructions

Ingest a longer text as one document into the tenant corpus: chunk splitting, embeddings, vector chunks and graph nodes. Upsert-by-title: if a document with the same title exists, the new version replaces the old one (server default). Use for long-form or multi-fact content (a report, a page, a doc); for a single short fact use knowmind_store_memory. Returns the document id and the number of chunks written. AFTER INGESTING: extract the facts the text states about people, organisations, projects, products, technologies and hosts, and create the corresponding edges via knowmind_link. Requires write scope.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
titleNoDocument title. Acts as the upsert key - an existing document with the same title is replaced.
sourceNoOrigin - URL or file path. Free text, stored for provenance.
contentYesFull document text (Markdown or plain); chunked and embedded server-side.
relationsNoEdges from this document to the things it talks about - set them HERE, in the same call. Name the counterpart in plain words; the server finds or creates its node and draws the edge. Allowed predicate values: ABOUT (Contract|Document|FileResource → Topic), ASSIGNED_TO (ActionRecord|Task → Agent|Organization|Person|SoftwareAgent), CLIENT_OF (Agent|Organization|Person|SoftwareAgent → Organization), CONTACT_PERSON_FOR (Person → Organization), COVERED_BY (Topic → Contract|Document|FileResource), DELIVERED_AS (Application → Product), DELIVERED_BY (Product → Application), DEPENDS_ON (Application → Application), DEPLOYED_TO (Application → Container), DEVELOPED_BY (Application → Agent|Organization|Person|SoftwareAgent), DEVELOPS (Agent|Organization|Person|SoftwareAgent → Application), ENABLES (Application → Application), FOR_CLIENT (Project → Agent|Organization|Person|SoftwareAgent), HAS_ASSIGNED_TASK (Agent|Organization|Person|SoftwareAgent → ActionRecord|Task), HAS_CHILD (Person → Person), HAS_CLIENT (Organization → Agent|Organization|Person|SoftwareAgent), HAS_CONTACT_PERSON (Organization → Person), HAS_EMPLOYEE (Organization → Person), HAS_OPERATED_APPLICATION (Agent|Organization|Person|SoftwareAgent → Application), HAS_PARENT (Person → Person), HAS_PREDECESSOR (ActionRecord|Project|Task → ActionRecord|Project|Task), HAS_PROJECT (Agent|Organization|Person|SoftwareAgent → Project), HAS_ROLE (Person → Role), HAS_SIBLING (Person → Person), HAS_SKILL (Agent|Organization|Person|SoftwareAgent → Skill), HAS_SUCCESSOR (ActionRecord|Project|Task → ActionRecord|Project|Task), HOSTED_ON (Container → Host), HOSTS (Host → Container), HOSTS_APPLICATION (Container → Application), INFRASTRUCTURE_PROVIDED_BY (Host → Organization), INTEGRATES_WITH (Application → Application), IS_LED_BY (Organization → Person), KNOWS (Person → Person), LEADS (Person → Organization), OPERATED_FOR (Application → Agent|Organization|Person|SoftwareAgent), PAID_BY (Agent|Organization|Person|SoftwareAgent → Agent|Organization|Person|SoftwareAgent), PARTNER_OF (Organization → Organization), PAYS (Agent|Organization|Person|SoftwareAgent → Agent|Organization|Person|SoftwareAgent), PRODUCED_BY (Contract|Document|FileResource → Project), PRODUCES (Project → Contract|Document|FileResource), PROVIDES_INFRASTRUCTURE (Organization → Host), ROLE_OF (Role → Person), SERVED_UNDER (Application → Domain), SERVES (Domain → Application), SKILL_OF (Skill → Agent|Organization|Person|SoftwareAgent), SPOUSE_OF (Person → Person), SUPPLIED_BY (Agent|Organization|Person|SoftwareAgent → Organization), SUPPLIES (Organization → Agent|Organization|Person|SoftwareAgent), SUPPLIES_TECHNOLOGY (Organization → Technology), SUPPORTED_BY (Contract|Document|FileResource → Contract|Document|FileResource), SUPPORTS (Contract|Document|FileResource → Contract|Document|FileResource), TECHNOLOGY_SUPPLIED_BY (Technology → Organization), TECHNOLOGY_USED_BY (Technology → Application), USES_TECHNOLOGY (Application → Technology), WORKED_ON_BY (Project → Agent|Organization|Person|SoftwareAgent), WORKS_FOR (Person → Organization), WORKS_ON (Agent|Organization|Person|SoftwareAgent → Project)

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed4 schema fields changedv0.3.7
    • changedInput schema / properties / content / description
      Previous value: -"Volltext des Dokuments"New value: +"Full document text (Markdown or plain); chunked and embedded server-side."
    • addedInput schema / properties / relations
      Added value: +{
      +  "description": "Edges from this document to the things it talks about - set them HERE, in the same call. Name the counterpart in plain words; the server finds or creates its node and draws the edge. Allowed predicate values: ABOUT (Contract|Document|FileResource → Topic), ASSIGNED_TO (ActionRecord|Task → Agent|Organization|Person|SoftwareAgent), CLIENT_OF (Agent|Organization|Person|SoftwareAgent → Organization), CONTACT_PERSON_FOR (Person → Organization), COVERED_BY (Topic → Contract|Document|FileResource), DELIVERED_AS (Application → Product), DELIVERED_BY (Product → Application), DEPENDS_ON (Application → Application), DEPLOYED_TO (Application → Container), DEVELOPED_BY (Application → Agent|Organization|Person|SoftwareAgent), DEVELOPS (Agent|Organization|Person|SoftwareAgent → Application), ENABLES (Application → Application), FOR_CLIENT (Project → Agent|Organization|Person|SoftwareAgent), HAS_ASSIGNED_TASK (Agent|Organization|Person|SoftwareAgent → ActionRecord|Task), HAS_CHILD (Person → Person), HAS_CLIENT (Organization → Agent|Organization|Person|SoftwareAgent), HAS_CONTACT_PERSON (Organization → Person), HAS_EMPLOYEE (Organization → Person), HAS_OPERATED_APPLICATION (Agent|Organization|Person|SoftwareAgent → Application), HAS_PARENT (Person → Person), HAS_PREDECESSOR (ActionRecord|Project|Task → ActionRecord|Project|Task), HAS_PROJECT (Agent|Organization|Person|SoftwareAgent → Project), HAS_ROLE (Person → Role), HAS_SIBLING (Person → Person), HAS_SKILL (Agent|Organization|Person|SoftwareAgent → Skill), HAS_SUCCESSOR (ActionRecord|Project|Task → ActionRecord|Project|Task), HOSTED_ON (Container → Host), HOSTS (Host → Container), HOSTS_APPLICATION (Container → Application), INFRASTRUCTURE_PROVIDED_BY (Host → Organization), INTEGRATES_WITH (Application → Application), IS_LED_BY (Organization → Person), KNOWS (Person → Person), LEADS (Person → Organization), OPERATED_FOR (Application → Agent|Organization|Person|SoftwareAgent), PAID_BY (Agent|Organization|Person|SoftwareAgent → Agent|Organization|Person|SoftwareAgent), PARTNER_OF (Organization → Organization), PAYS (Agent|Organization|Person|SoftwareAgent → Agent|Organization|Person|SoftwareAgent), PRODUCED_BY (Contract|Document|FileResource → Project), PRODUCES (Project → Contract|Document|FileResource), PROVIDES_INFRASTRUCTURE (Organization → Host), ROLE_OF (Role → Person), SERVED_UNDER (Application → Domain), SERVES (Domain → Application), SKILL_OF (Skill → Agent|Organization|Person|SoftwareAgent), SPOUSE_OF (Person → Person), SUPPLIED_BY (Agent|Organization|Person|SoftwareAgent → Organization), SUPPLIES (Organization → Agent|Organization|Person|SoftwareAgent), SUPPLIES_TECHNOLOGY (Organization → Technology), SUPPORTED_BY (Contract|Document|FileResource → Contract|Document|FileResource), SUPPORTS (Contract|Document|FileResource → Contract|Document|FileResource), TECHNOLOGY_SUPPLIED_BY (Technology → Organization), TECHNOLOGY_USED_BY (Technology → Application), USES_TECHNOLOGY (Application → Technology), WORKED_ON_BY (Project → Agent|Organization|Person|SoftwareAgent), WORKS_FOR (Person → Organization), WORKS_ON (Agent|Organization|Person|SoftwareAgent → Project)",
      +  "items": {
      +    "properties": {
      +      "confidence": {
      +        "description": "Edge confidence in [0..1].",
      +        "type": "number"
      +      },
      +      "object": {
      +        "description": "Proper name of the counterpart.",
      +        "type": "string"
      +      },
      +      "object_class": {
      +        "description": "Class of the counterpart (Person, Organization, Application, Host, Technology, ...).",
      +        "type": "string"
      +      },
      +      "object_description": {
      +        "description": "One sentence about the counterpart, used when its node has to be created. Optional.",
      +        "type": "string"
      +      },
      +      "predicate": {
      +        "description": "Edge type in UPPER_SNAKE_CASE from the allowed list.",
      +        "type": "string"
      +      }
      +    },
      +    "required": [
      +      "predicate",
      +      "object"
      +    ],
      +    "type": "object"
      +  },
      +  "type": "array"
      +}
    • changedInput schema / properties / source / description
      Previous value: -"Quelle, z.B. URL oder Dateipfad"New value: +"Origin - URL or file path. Free text, stored for provenance."
    • changedInput schema / properties / title / description
      Previous value: -"Dokument-Titel"New value: +"Document title. Acts as the upsert key - an existing document with the same title is replaced."
  2. First observedv0.3.1

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond the annotations by disclosing the upsert-by-title replacement behavior, server-side chunking/embedding, and the fact that it returns a document id plus chunk count. It also clarifies the destructive edge case in which an existing document with the same title is replaced, which is exactly the kind of context an agent needs despite destructiveHint being false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and efficient, front-loading the core purpose in the first sentence. Every sentence carries meaningful information: when to use it, the behavior, the output, the alternative tool, and the follow-up workflow. No filler or redundant restating of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is strong overall: it covers return values, write-permission requirements, upsert behavior, and follow-up links. Minor ambiguity remains because it directs the agent to create edges via knowmind_link after ingesting, even though the input schema allows edges via the 'relations' parameter in the same call, leaving two possible workflow interpretations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers all four parameters with full descriptions (100% coverage), so the baseline is 3. The description adds useful high-level context ('title acts as the upsert key', 'content is chunked and embedded server-side') but does not materially improve on the structured parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Ingest a longer text as one document into the tenant corpus') and specifies the processing pipeline (chunking, embeddings, vector chunks, graph nodes). It also explicitly contrasts with knowmind_store_memory for short facts, making the tool's purpose distinguishable from siblings without opening them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use the tool ('long-form or multi-fact content') and when not to ('for a single short fact use knowmind_store_memory'). It also gives a clear post-ingestion workflow via knowmind_link and the required write scope, giving the agent actionable context for choosing and using the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.