Skip to main content
Glama

Upload file

upload_file

Upload one or more files to Clueso. Pick a mode by client + where the file lives:

  1. files — ChatGPT only: files the user attached in the conversation. ChatGPT fills each entry (download_url + file_id) itself; pass the attachments here rather than asking the user to re-upload. Returns one mcp_upload_id per file.

  2. file_name — HOSTED upload, the default for any non-UI / programmatic upload (Claude Code, Cursor, Claude Desktop, scripts). Returns an upload URL on Clueso's OWN base domain + a ready-to-run curl that streams a single local file to it; Clueso relays the bytes to storage server-side. The PUT targets the base domain — NOT cloud storage directly — so it works on desktop/agent clients that can't reach or are blocked from S3. Requirement: the client must be able to PUT bytes to the Clueso base domain (run the returned curl, or any HTTP PUT). The agent (or the user at a shell prompt) runs the curl. Prefer this whenever there's no human at a browser.

  3. file_url: Pass a public https URL. Server fetches and stages the file. Returns mcp_upload_id immediately. Use when the file is already on the open web — no user interaction needed.

  4. request_hosted_upload (UI mode — use ONLY when a human should pick files in a browser: many files at once, or a host with no shell / no PUT capability): Returns a single upload_token + upload_page URL. Share the link with the user; they open it in a new browser tab, drop their files, click Done. Then call check_uploads(upload_token) to retrieve all mcp_upload_ids. Call once for all files.

Hosted uploads cover any number of files per call: one call issues one upload_token, and that token covers every file the user drops on the page. Repeat calls issue additional tokens, each tracking only its own files.

The returned mcp_upload_id (prefixed mup_) can be passed to:

  • add_elements / update_elements (image or video → an element ON a clip: pass it as type_data.mcp_upload_id, on either tool — this is how a local image becomes on-canvas content, and how an existing element's source is swapped). To fill an animation's image slot, pass it inside type_data.parameter_values on update_elements only — parameter_values is an update-path field and is stripped on add.

  • add_audio (audio → project music track that plays under all clips)

  • add_clips(kind='video') (video or audio → sequential clip with auto-transcription)

  • add_clips(kind='pptx') (.ppt/.pptx → slide clips)

  • add_article_media (image/GIF → article asset)

  • analyze_audio (audio → transcript / silences / beats / features)

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
filesNoChatGPT file attachments, one entry per file (filled by ChatGPT).
file_urlNoPublic URL to fetch the file from
file_nameNoFile name with extension. Returns a Clueso upload URL + curl command that streams this single local file to us (single file).
file_namesNoList of file names the user will upload (for hosted mode). Shown on the upload page as guidance.
request_hosted_uploadNoIf true, returns a hosted upload page. Call once for all files — the page accepts multiple uploads under one token.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / properties / files
      Added value: +{
      +  "description": "ChatGPT file attachments, one entry per file (filled by ChatGPT).",
      +  "items": {
      +    "properties": {
      +      "download_url": {
      +        "description": "Temporary download URL supplied by ChatGPT",
      +        "type": "string"
      +      },
      +      "file_id": {
      +        "description": "ChatGPT file id",
      +        "type": "string"
      +      },
      +      "file_name": {
      +        "type": "string"
      +      },
      +      "mime_type": {
      +        "type": "string"
      +      }
      +    },
      +    "required": [
      +      "download_url",
      +      "file_id"
      +    ],
      +    "type": "object"
      +  },
      +  "type": "array"
      +}
  2. Changed4 schema fields changed
    • removedInput schema / properties / context
      Removed value: -{
      -  "description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
      -  "type": "string"
      -}
    • removedInput schema / properties / conversation_id
      Removed value: -{
      -  "description": "Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.",
      -  "type": "string"
      -}
    • removedInput schema / properties / llm_model
      Removed value: -{
      -  "description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
      -  "type": "string"
      -}
    • removedInput schema / required
      Removed value: -[
      -  "context",
      -  "llm_model"
      -]
  3. Changed4 schema fields changed
    • addedInput schema / properties / context
      Added value: +{
      +  "description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
      +  "type": "string"
      +}
    • addedInput schema / properties / conversation_id
      Added value: +{
      +  "description": "Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.",
      +  "type": "string"
      +}
    • addedInput schema / properties / llm_model
      Added value: +{
      +  "description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
      +  "type": "string"
      +}
    • addedInput schema / required
      Added value: +[
      +  "context",
      +  "llm_model"
      +]
  4. Changed2 schema fields changed
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • removedInput schema / additionalProperties
      Removed value: -false
  5. First observed

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false, openWorldHint=true, destructiveHint=false. The description carries the full behavioral burden: it discloses that uploads go to Clueso's own domain, relays bytes server-side, requires PUT capability, returns upload URLs, tokens, and mcp_upload_ids, and explains the curl execution step. It also warns about parameter_values being stripped on add for animation slots. All of this goes beyond the sparse annotations and contradicts nothing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but justifiably so given the tool's complexity with four modes and downstream usage. It is well-structured with numbered modes, bullet points, and clear separation of concerns. Front-loaded with the core purpose and mode selection. A minor deduction because it could be trimmed slightly (e.g., the long paragraph about how mcp_upload_id is used could be condensed), but overall every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of an output schema, the description explicitly covers return values for every mode (mcp_upload_id, upload URL + curl, upload_token + upload_page). It explains the workflow for hosted uploads and how to retrieve IDs via check_uploads. It also maps the ID to consumption tools and notes the parameter_values caveat. Nothing an agent needs to correctly invoke and handle this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema has 100% parameter descriptions, the tool description adds substantial meaning: it explains the semantic difference between file_name (single local file, returns curl) and file_names (list of file names for hosted upload page), clarifies that files is ChatGPT-only and auto-populated, and details how request_hosted_upload triggers a UI flow. This is far more than the one-line schema descriptions, giving agents the context needed to pick the right parameter for the situation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb+resource: 'Upload one or more files to Clueso.' It then breaks down four distinct modes (files, file_name, file_url, request_hosted_upload) with specific use cases, fully distinguishing the tool from its siblings. The description also explains how the returned mcp_upload_id flows into other tools (add_elements, add_audio, add_clips, etc.), which further clarifies its role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance for each mode: files for ChatGPT-only attachments, file_name as the default for non-UI/programmatic uploads, file_url for public web files, and request_hosted_upload for browser-based human selection. It even names the alternative tool check_uploads for polling uploads and states conditions like 'Prefer this whenever there's no human at a browser' and 'use ONLY when a human should pick files in a browser.' This is exemplary routing information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.