Skip to main content
Glama

Upload Asset

upload_asset

Upload a binary asset (image, font, audio, …) to the project's hosted storage. This uploads bytes you actually hold — a file you generated, downloaded, or read yourself. Chat attachments don't qualify: the user's attachments never reach MCP servers (you see attached images through vision only; there is no file, id, or URL behind them you can read), so for those use request_user_upload instead and the user re-picks the file in a card that uploads from their browser. Three modes. ChatGPT conversation files — a generated image, a file ChatGPT itself holds: pass the file as the file parameter and the host attaches a download link itself; this server fetches the bytes directly, at full quality (nothing goes through your sandbox or through base64 in arguments; content_type and size_bytes are optional here). Never downscale or re-encode a generated image to fit the inline cap — pass it as file instead. Files up to 3 MB you hold yourself — pass content_base64 plus size_bytes (the decoded byte count) and the upload completes in this call, returning publicUrl. Larger files — pass size_bytes alone to get an uploadUrl; PUT the raw bytes to it with the same content_type and exact byte count (e.g. curl -X PUT -H 'Content-Type: image/png' --data-binary @file.png '<uploadUrl>'), then reference publicUrl. Some sandboxes (claude.ai Cowork, ChatGPT containers) block egress to S3: if the PUT fails in any way — connection failure, proxy error, or a response without an x-amz-request-id header — that block is permanent for the session, so switch paths instead of retrying or re-encoding smaller: the file parameter in ChatGPT for any file that exists in this conversation, content_base64 for files under 3 MB, request_user_upload for user-provided files, or a PUT from inside the project VM via run_code_in_vm (re-mint the URL first; it is short-lived). For AI imagery generated fresh, use generate_image. A single file can be at most 100 MB via the presigned mode (the inline content_base64 mode is capped at 3 MB).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fileNoChatGPT only: a file from this conversation (e.g. a generated image). The ChatGPT host fills download_url/file_id when you reference the file; the values are host-issued and cannot be constructed by hand — on clients without file-parameter support, leave this unset and use the other modes.
file_nameYes
projectIdYes
size_bytesNoRequired unless `file` is set. Exact byte count of the file, as measured from the file itself (e.g. stat/ls -l). With content_base64 it must equal the decoded length; in presigned mode the PUT must send exactly this many bytes.
content_typeNoRequired unless `file` is set (in that mode the type is derived from the downloaded bytes; pass this only as a hint).
content_base64NoThe file's bytes, base64-encoded (≤3 MB decoded). When set, the upload completes in this call.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false, destructiveHint=false, openWorldHint=true. The description adds substantial behavior the annotations cannot convey: three distinct upload modes with different completion semantics (inline vs presigned), size caps (3 MB inline, 100 MB presigned), short-lived uploadUrl, and the S3 egress-block failure signature (missing x-amz-request-id header) plus the explicit 'do not retry, switch paths' guidance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded on the core verb and the critical disqualifier (chat attachments), then structured by mode. It is long, but given six parameters, three modes, and a hard failure mode, nearly every sentence carries operative information. A tighter edit could fold the failure-path list, but nothing is pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description names the return values (publicUrl, uploadUrl) and the full call lifecycle for each mode. For a 6-parameter tool with nested objects and out-of-band PUT steps, all the information needed to invoke it correctly on the first attempt is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% and the description compensates well, explaining that file_id/download_url are host-issued and cannot be hand-constructed, that size_bytes must equal the decoded length in base64 mode and the exact PUT byte count otherwise, and that content_type is a derived hint in file mode. It does not restate the file_name pattern, but the schema covers that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Upload a binary asset ... to the project's hosted storage') and immediately delimits scope against siblings: request_user_upload for chat attachments, generate_image for fresh AI imagery. An agent can distinguish it from card_upload_asset, copy_file, and write_file without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes between four alternatives with the selecting condition for each: request_user_upload when the bytes are user-provided, generate_image when generating fresh imagery, the `file` param for ChatGPT conversation files, and run_code_in_vm PUT when sandbox egress is blocked. It also states the when-not case (chat attachments never reach MCP servers) with the reason.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources