Gumloop MCP Server
A Gumloop MCP server/CLI exposing 92 tools (50 reads, 42 confirmation-gated writes) that let an AI agent or script operate Gumloop flows, agents, sessions, Brain, skills, artifacts, evaluations and organization administration under private, per-call approval.
Flows and runs — list workbooks/saved flows, inspect a flow's input schema, start a run, kill a run, poll run details and fetch run history.
Files — upload single/multiple files (base64 or local path), download run files to a private
output_file, and upload files into a session.Agents — list, create, retrieve and update agents; read version history and snapshots; attach/detach skills and MCP servers.
Sessions — create, retrieve, rename, cancel, send messages, upload attachments, resolve pending approvals, and manage the message queue (list/queue/update/delete/send now).
Chat and models — create OpenAI-style non-streaming chat completions (incl. image modalities and tool calls), list available models, and ask the Chew router which model it would pick.
Brain — hybrid search across indexed sources, list/create/delete sources, upload/list/delete files, read credit estimates and approve a draft source.
Skills and artifacts — list/create/update/delete skills, download a skill archive, list agent artifacts and download one.
MCP servers catalog — list/retrieve servers, enumerate their tools, resources and prompts, read a resource, render a prompt, and execute batched tool calls.
Evaluations — agent-level results, metrics, config and runs; organization-level evaluations with targets, runs, results and metrics.
Administration — audit logs, workspace user management, custom-role membership, role credit limits (list/get/set), data exports and export status.
Discovery/identity — static schema and help, plus local
list_accountsfor private profile labels with no network call.
Confluence spaces connected as Gumloop Brain sources can be searched through the server's hybrid semantic/keyword search, returning ranked, cited snippets from indexed Confluence content.
GitHub repositories connected as Gumloop Brain sources can be searched through the server's hybrid semantic/keyword search, returning ranked, cited snippets from indexed repository content.
Google Drive content connected as a Gumloop Brain source can be searched through the server's hybrid semantic/keyword search, returning ranked, cited snippets from indexed Drive documents.
Notion workspaces connected as Gumloop Brain sources can be queried through the server's hybrid semantic/keyword search tool, which returns ranked, cited snippets from indexed Notion pages.
Slack workspaces connected as Gumloop Brain sources can be searched through the server's hybrid semantic/keyword search, returning ranked, cited snippets from indexed Slack content.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Gumloop MCP Serverlist my saved flows"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Gumloop MCP Server & CLI
Gumloop MCP server and CLI for Codex and AI agents. 92 tools for current flows, agents, sessions, Brain, skills, artifacts and administration, with private accounts and explicit operation approval. One shared implementation supplies both binaries and a desktop bundle.
Built and maintained by Navid Moazzez. Built on Slipway, which turns one definition of each tool into the MCP server and the CLI. The complete guide is on navid.me.
The terminal illustrates real command names and approval flow. It is not a recording of a provider account run. Gumloop already has official CLI and hosted MCP products; their supported platform, authentication and workflows are compared below.
Requires Node 22+ and eligible Gumloop API access for account operations. Validation: fixture tests, schema validation and protocol/artifact discovery are separate from provider-account and desktop GUI outcomes, which are not claimed. Section 7 has the measured token costs.
Two ways to use it
Command line
npm install -g @thenavidm/gumloop-mcp-cli@latest
gumloop-cli
gumloop-cli list-agents --agent
gumloop-cli schema start-flow
gumloop-cli start-flow --payload-file /absolute/private/approved-flow.json --account work --confirm --agentMCP server, for your AI app
codex mcp add gumloop -- npx -y @thenavidm/gumloop-mcp-cli@latestConfigure private credentials first. Ask: “Inspect this existing run and return its state; do not start another run.” Read INSTALL.md for complete client/OS routes.
Which one
Where you work | Surface |
Codex or shell agent | Shared CLI, local MCP or both |
Desktop chat | Compatible local MCP/desktop archive |
Scripts/CI | CLI or MCP client |
Remote-only client | Official hosted MCP |
Related MCP server: n8n Manager for AI Agents
Features
Capability | CLI command | MCP tool |
Saved flows and inputs | list-flows / get-input-schema | list_flows / get_input_schema |
Run and inspect | start-flow / get-run-details | start_flow / get_run_details |
Agents and sessions | list-agents / create-session / retrieve-session | list_agents / create_session / retrieve_session |
Skills and Brain | create-skill / search-brain | create_skill / search_brain |
Private file delivery | download-artifact-file / download-files | download_artifact_file / download_files |
Explicit account selection | list-accounts / --account | list_accounts / account |
Diagnosis/setup | doctor / login | CLI utilities |
Contents
Number | Section | What it covers |
1 | What you can ask it | |
2 | Quick install | |
3 | Set up Gumloop access | |
4 | Connect your client | |
5 | Check it works | |
6 | Output, flags and exit codes | |
7 | MCP or CLI and token cost | |
8 | Every tool and argument | |
9 | Flows, sessions and account workflows | |
10 | Jobs, pagination and private files | |
11 | Several private accounts | |
12 | Writing safely | |
13 | How it works | |
14 | Your data | |
15 | Environment variables | |
16 | Updates and removal | |
17 | Troubleshooting | |
18 | API coverage and comparisons | |
19 | Versions | |
20 | FAQ |
1. What you can ask it
Inspect accessible agents, versions, teams and models.
List saved flows and their exact input schema before approving a run.
Start only the selected flow, keep its returned run ID and read its state.
Create a reviewed agent session, then inspect its approval requirements.
Approve only the selected session checkpoint; do not approve a whole conversation implicitly.
Upload a selected regular file and save an existing run download privately.
Review organization usage, audit records and applicable role credit limits.
Actual discovery supplies 92 tools: 50 reads and 42 confirmation-gated operations, from all 91 reviewed current REST operations plus list_accounts. The shared handlers cover MCP and native Node CLI. This is protocol/schema evidence; account outcomes and desktop GUI installation remain separate.
2. Quick install
npm install -g @thenavidm/gumloop-mcp-cli@latest
gumloop-cli --version
gumloop-cli login
gumloop-cli doctor
gumloop-cli toolsNode 22+ is required for manual installation. The versioned desktop archive bundles production dependencies for a compatible host. Read INSTALL.md before configuring credentials. After private setup:
codex mcp add gumloop -- npx -y @thenavidm/gumloop-mcp-cli@latest
codex mcp list3. Set up Gumloop access
Private credentials and permissions
Open the intended account's Connectors settings. Select Gumloop API Key and create the personal or team key needed for your task.
Read Profile Settings for your user ID. A team ID is optional; older flow endpoints call it project_id. Set GUMLOOP_USER_ID and, when appropriate, GUMLOOP_TEAM_ID in private settings.
Save the credential as a token-only regular file outside repositories. Set GUMLOOP_TOKEN_FILE to its absolute path; GUMLOOP_API_KEY is the alternative. Do not paste real keys into chat or command examples.
Run gumloop-cli doctor, then doctor --network. The network check reads accessible agent metadata without printing account content.
Inspect the exact schema and selected resource before confirming an operation. Runs can spend credits and invoke connected downstream services.
Current authentication docs require Pro or above for API keys. Personal keys act as their owner; team keys can act as team members through the requested user identity. Workflow-step credential settings decide personal/team credentials when project_id is supplied. Local profiles do not bypass account, role or team permissions.
Requests use Authorization: Bearer, with the selected profile user ID as x-auth-key, following the current SDK. Ordinary endpoints use https://api.gumloop.com/api/v1; non-streaming chat uses the separately documented https://ws.gumloop.com/api/v1. No arbitrary host override or redirects are accepted.
For POSIX, use a private 0700 directory and regular 0600 token-only file, no symlink, at most 64 KB. Windows requires a user-only ACL; POSIX checks do not verify it. File credentials override environment credentials and are cached until restart. The package does not read an automatic .env, use the OS keychain or start browser login.
OAuth and rotation
An already authorized OAuth access token can be supplied through the same private Bearer credential path. This wrapper does not register an OAuth app, accept refresh tokens, refresh access tokens or manage consent. OAuth docs describe invite-only client registration, authorization code with PKCE S256 and gumloop_api/userinfo scopes. gumloop_api requires Pro or above; requesting userinfo alone is not API permission. Use an expiry-aware issuer or the official CLI's supported OAuth/keychain workflow when automatic refresh is needed.
Rotate/revoke the intended grant through its provider controls, update private settings and restart. Removing a local package does not revoke provider credentials, undo runs or remove hosted data.
Plans, credits and rate limits
The AGPL wrapper is free; Gumloop account access and processing are billed separately. Credit documentation describes variable charges for model work, connector calls, compute and orchestration; Brain searches can also incur charges. Local read-only mode is an operation policy, not a guarantee of zero provider charges. Credit-consuming Brain search is confirmation-gated here.
Agent concurrency limits are organization-wide, currently 25 on Pro and 100 on Enterprise, with customizable Enterprise limits. Pro rejects excess agent work; Enterprise can queue it. Webhook triggers separately document 100 requests/minute per trigger. These are distinct limits; this wrapper's default 150 ms account/process pacing reserves no provider capacity.
GET 429 retries require an explicit Retry-After at most ten seconds, default two/max five retries. Missing/longer delays return exit 7. Mutations, paid Brain search and network timeouts never retry automatically. Local JSON request cap is 5 MiB; responses/downloads are capped at 10 MiB. These local caps are separate from vendor storage, upload and pagination limits.
4. Connect your client
INSTALL.md gives Codex-first setup, optional Claude Code, desktop archive/manual configuration, Cursor, VS Code/Copilot, Windsurf, Zed, Gemini CLI, Docker and Cline. The local transport is stdio. Configure private credentials in the actual client environment; GUI apps and remote containers do not inherit every shell setting.
The published official gumloop 0.5.2 CLI and current provider docs explicitly refuse native Windows and recommend WSL or the Python SDK. This wrapper targets Node 22+ on native Windows, macOS and Linux. Compatible desktop-host availability and GUI installation are separate from CLI portability. Remote-only clients can use the official https://mcp.gumloop.com/gumloop/mcp endpoint.
npm ships SKILL.md but does not register it automatically. Install it through the client's supported skill location. A connection/client approval and this wrapper's confirm argument are separate controls.
5. Check it works
gumloop-cli --version
gumloop-cli doctor
gumloop-cli doctor --network
gumloop-cli list-accounts --agent
gumloop-cli list-agents --agentDiscovery, schemas/help and account labels need no provider key. Network doctor reads GET /agents and prints only diagnostic status. This validates that request, not every account operation. Full discovery exposes 92 tools; read-only discovery exposes 50. Missing credentials exits 10, invalid input or refused operations exit 2. First validate an existing resource; do not start billable automation merely to test installation.
6. Output, flags and exit codes
Tool results go to stdout. Errors are JSON on stderr. JSON operations return structured results, so --select can retain nested fields. Private download operations return file metadata; binary bytes and signed results stay in the requested private file.
gumloop-cli start-flow --help
gumloop-cli schema start-flow
gumloop-cli list-agents --agent --select agentsFlag | What it does |
| JSON output |
| Single-line JSON |
| Compact JSON and no prompts; never confirms a write |
| Keep selected fields; dotted paths descend and arrays are traversed |
| Confirm the requested flow, session, paid search or account operation |
| Automation switches; none overrides the spending guard |
Global output flags apply to tool commands. doctor has its own --network option and returns a JSON diagnostic.
Exit code | Meaning | What a script should do |
0 | Success | Read stdout |
1 | Unexpected error | Report it with the command that caused it |
2 | Usage, invalid input, a refused write, an unknown command or a hidden write | Fix the input or confirm only the requested action |
3 | Job or local upload file not found | Check the ID/path |
4 | Authentication or entitlement rejected | Check private credential settings and permissions |
5 | API or network failure | Inspect an accepted job before another paid submission |
7 | Rate limited | Wait; do not loop over paid submissions |
10 | Credentials not configured | Complete local setup |
The underscore spelling also works. start_flow and start-flow call the same tool. Nested objects use quoted JSON. Arrays of objects use repeated flags, one JSON object at a time.
7. MCP or CLI and token cost
MCP and CLI use the same catalogue, schemas, handlers and write guard: Slipway builds the MCP server, over stdio or --http, and the CLI from each tool's one definition. Shell scripts can select fields with --select after receipt; this does not change upstream result size or billing.
Measured on 2026-10-05 against 2.0.2, with Claude Code 2.1.286 on Claude Opus 5.5 (one short prompt with and without the server connected, the difference read from the API's own usage figures) and Codex 0.159.3 on gpt-6.1-sol:
Cost | 2.0.2 | 3.0.0 |
Claude Code, every tool loaded, every message | 61,936 | 56,695 |
Claude Code's default, tool search, every message | 1,486 | 1,488 |
| 3,150 | 3,210 |
Codex over the CLI, one task, median of five | 107,376 | 83,661 |
Codex over MCP, the same task, median of five | 78,443 | 78,249 |
The task was "find the command that starts a flow run, and the flags it requires". Every tool loaded costs less because parts that several tools repeated are written once. Over the CLI, every 3.0.0 run asked which, whose answer carries the command's help, where every 2.0.2 run read the command list and then the help: one request fewer. Over MCP, 3.0.0 cost slightly less. SKILL.md costs 60 more because it says how approval works over MCP and what exit codes 1 and 2 cover.
Tool-list bytes or characters divided by four are not API usage, and no other offering was measured.
Provider credits and client-model tokens are separate.
8. Every tool and argument
Every route and argument below comes from actual stdio discovery and the reviewed current schema. Use schema COMMAND for exact inline nested objects and unions.
Tool | Route | Mode |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Read |
|
| Read |
|
| Read |
|
| Read |
|
| Read |
|
| Read |
|
| Read |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Read |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Read |
|
| Read |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Read |
|
| Read |
|
| Confirm exact operation |
|
| Read |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Read |
|
| Read |
|
| Read |
|
| Read |
|
| Confirm exact operation |
|
| Read |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Read |
|
| Confirm exact operation |
|
| Read |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Read |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Read |
|
| Read |
|
| Read |
|
| Read |
|
| Read |
|
| Read |
|
| Read |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Read |
|
| Confirm exact operation |
|
| Read |
|
| Confirm exact operation |
|
| Read |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Read |
|
| Confirm exact operation |
|
| Read |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Read |
|
| Read |
|
| Read |
|
| Read |
|
| Confirm exact operation |
|
| Read |
|
| Read |
|
| Confirm exact operation |
|
| Read |
|
| Read |
|
| Read |
|
| Confirm exact operation |
|
| Read |
|
| Read |
|
| Read |
|
| Confirm exact operation |
|
| Read |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Confirm exact operation |
|
| Read |
|
| Read |
|
| Read |
| Local, no network | Read |
start_flow
gumloop-cli start-flow
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The id for the user initiating the flow. |
| No; body and guard rules apply | string | (Optional) The id of the project within which the flow is executed. |
| No; body and guard rules apply | string | The id for the saved flow. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: saved_item_id, user_id. Profile defaults fill supported user/team identity fields before body validation.
kill_flow
gumloop-cli kill-flow
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The ID of the pipeline run to kill. |
| No; body and guard rules apply | string | The user ID. Required if project_id is not provided. |
| No; body and guard rules apply | string | The project ID. Required if user_id is not provided. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: run_id. Profile defaults fill supported user/team identity fields before body validation.
get_run_details
gumloop-cli get-run-details
Argument | Required | Type | Details |
| Yes | string | ID of the flow run to retrieve |
| No; body and guard rules apply | string | The id for the user initiating the flow. Required if project_id is not provided. |
| No; body and guard rules apply | string | The id of the project within which the flow is executed. Required if user_id is not provided. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
list_workbooks
gumloop-cli list-workbooks
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The user ID for which to list workbooks. Required if project_id is not provided. |
| No; body and guard rules apply | string | The project ID for which to list workbooks. Required if user_id is not provided. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
list_flows
gumloop-cli list-flows
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The user ID to for which to list items. Required if project_id is not provided. |
| No; body and guard rules apply | string | The project ID for which to list items. Required if user_id is not provided. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
get_input_schema
gumloop-cli get-input-schema
Argument | Required | Type | Details |
| Yes | string | The ID of the saved item for which to retrieve input schemas. |
| No; body and guard rules apply | string | User ID that created the flow. Required if project_id is not provided. |
| No; body and guard rules apply | string | Project ID that the flow is under. Required if user_id is not provided. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
get_run_history
gumloop-cli get-run-history
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The ID of the workbook to retrieve run history for. Required if saved_item_id is not provided. |
| No; body and guard rules apply | string | The ID of the saved item to retrieve run history for. Required if workbook_id is not provided. |
| No; body and guard rules apply | string | The user ID. Required if project_id is not provided. |
| No; body and guard rules apply | string | The project ID. Required if user_id is not provided. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
download_file
gumloop-cli download-file
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The name of the file to download. |
| No; body and guard rules apply | string | The ID of the flow run associated with the file. |
| No; body and guard rules apply | string | The saved item ID associated with the file. |
| No; body and guard rules apply | string | Optional. The user ID associated with the flow run. |
| No; body and guard rules apply | string | Optional. The project ID associated with the flow run. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
| Yes | string | New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.
output_file must be new and absolute in a private directory, reserved exclusively before the API call. Binary bytes or signed JSON are kept out of model output. This wrapper never follows a signed external download URL.
download_files
gumloop-cli download-files
Argument | Required | Type | Details |
| No; body and guard rules apply | array | An array of file names to download. Items: string. |
| No; body and guard rules apply | string | The ID of the flow run associated with the files. |
| No; body and guard rules apply | string | The user ID associated with the files. Required if project_id is not provided. |
| No; body and guard rules apply | string | The project ID associated with the files. Required if user_id is not provided. |
| No; body and guard rules apply | string | Optional. The saved item ID associated with the files. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
| Yes | string | New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.
output_file must be new and absolute in a private directory, reserved exclusively before the API call. Binary bytes or signed JSON are kept out of model output. This wrapper never follows a signed external download URL.
upload_file
gumloop-cli upload-file
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The name of the file to be uploaded. |
| No; body and guard rules apply | string | Base64 encoded content of the file. format: |
| No; body and guard rules apply | string | The user ID associated with the file. Required if project_id is not provided. |
| No; body and guard rules apply | string | The project ID associated with the file. Required if user_id is not provided. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
| No; body and guard rules apply | string | Regular local file path, no symlink, at most 3 MiB. Encoded as native base64 file_content; cannot mix with file_content or payload routes. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.
upload_files
gumloop-cli upload-files
Argument | Required | Type | Details |
| No; body and guard rules apply | array | See the full input schema. Items: object. |
| No; body and guard rules apply | string | The user ID associated with the files. Required if project_id is not provided. |
| No; body and guard rules apply | string | The project ID associated with the files. Required if user_id is not provided. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.
get_organization_audit_logs
gumloop-cli get-organization-audit-logs
Argument | Required | Type | Details |
| Yes | string | The ID of the organization to retrieve audit logs for. |
| No; body and guard rules apply | string | Your user id -- you must be an organization admin to retrieve organization logs. |
| Yes | string | Start timestamp for log filtering (ISO format). format: |
| Yes | string | End timestamp for log filtering (ISO format). format: |
| No; body and guard rules apply | string | Comma-separated list of event types to filter by (e.g. |
| No; body and guard rules apply | string | Comma-separated list of user IDs whose events should be returned. |
| No; body and guard rules apply | string | Comma-separated list of source IP addresses to filter by. The singular |
| No; body and guard rules apply | string | Comma-separated list of workspace (team) IDs to filter by. The singular |
| No; body and guard rules apply | string | Comma-separated list of entity IDs (agents, workbooks, files) to filter by. The singular |
| No; body and guard rules apply | integer | Page number for pagination. default: |
| No; body and guard rules apply | integer | Number of records per page. default: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
manage_project_users
gumloop-cli manage-project-users
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The ID of the organization that the workspace belongs to. |
| No; body and guard rules apply | string | Your user id -- you must be an organization admin to manage workspace users. |
| No; body and guard rules apply | string | The ID of the workspace to manage users for. |
| No; body and guard rules apply | string | The action to perform - either 'add' or 'remove' a user. Values: |
| No; body and guard rules apply | string | The email address of the target user to add or remove. |
| No; body and guard rules apply | boolean | When adding a user, specify whether they should have admin privileges (default is false). |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: organization_id, user_id, workspace_id, action, user_email. Profile defaults fill supported user/team identity fields before body validation.
manage_permission_group_users
gumloop-cli manage-permission-group-users
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The ID of the organization that the custom role belongs to. |
| No; body and guard rules apply | string | Your user id -- you must be an organization admin to manage custom role users. |
| No; body and guard rules apply | string | The ID of the custom role to manage users for. |
| No; body and guard rules apply | string | The action to perform - either 'add' or 'remove' a user. Values: |
| No; body and guard rules apply | string | The email address of the target user to add or remove. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: organization_id, user_id, group_id, action, user_email. Profile defaults fill supported user/team identity fields before body validation.
list_role_credit_limits
gumloop-cli list-role-credit-limits
Argument | Required | Type | Details |
| Yes | string | The ID of the organization. minLength: |
| No; body and guard rules apply | string | Your user id -- you must be an organization admin to manage custom role credit limits. |
| No; body and guard rules apply | integer | Number of roles per page. maximum: |
| No; body and guard rules apply | string | Opaque cursor from a previous response's |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
get_role_credit_limit
gumloop-cli get-role-credit-limit
Argument | Required | Type | Details |
| Yes | string | The ID of the organization the custom role belongs to. minLength: |
| Yes | string | The ID of the custom role (the same ID used as |
| No; body and guard rules apply | string | Your user id -- you must be an organization admin to manage custom role credit limits. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
set_role_credit_limit
gumloop-cli set-role-credit-limit
Argument | Required | Type | Details |
| Yes | string | The ID of the organization the custom role belongs to. minLength: |
| Yes | string | The ID of the custom role (the same ID used as |
| No; body and guard rules apply | integer/null | The monthly credit limit applied to each member of this role, or null to clear the role-level limit. minimum: |
| No; body and guard rules apply | string | Your user id -- you must be an organization admin to manage custom role credit limits. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: monthly_credit_limit, user_id. Profile defaults fill supported user/team identity fields before body validation.
export_data
gumloop-cli export-data
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The ID of the user requesting the export. |
| No; body and guard rules apply | string | The type of data to export. Use |
| No; body and guard rules apply | string | Only applicable when |
| No; body and guard rules apply | string | The scope of the export. Use |
| No; body and guard rules apply | array | Array of fields to include in the export. The available fields depend on the |
| No; body and guard rules apply | string | Start date for the export in ISO 8601 format (e.g., |
| No; body and guard rules apply | string | End date for the export in ISO 8601 format (e.g., |
| No; body and guard rules apply | boolean | Whether to include all workspaces in the organization. When |
| No; body and guard rules apply | boolean | Whether to include personal workspaces in the export. Ignored if |
| No; body and guard rules apply | array | An optional array of workspace IDs to include in the export. When |
| No; body and guard rules apply | array | An optional array of specific entity IDs to filter the export. For workflow exports ( |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: user_id, export_fields, start_date, end_date. Profile defaults fill supported user/team identity fields before body validation.
get_export_status
gumloop-cli get-export-status
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The ID of the user requesting the export status. |
| Yes | string | The unique identifier of the data export job to check (returned by the Export data endpoint). |
| No; body and guard rules apply | boolean | Set to |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| Yes | string | New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. minLength: |
output_file must be new and absolute in a private directory, reserved exclusively before the API call. Binary bytes or signed JSON are kept out of model output. This wrapper never follows a signed external download URL.
list_agents
gumloop-cli list-agents
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Scope the listing to a single team. When omitted, returns agents owned by the authenticated user. |
| No; body and guard rules apply | string | Case-insensitive substring match against the agent name. |
| No; body and guard rules apply | string | Filter to agents created by this user ID. |
| No; body and guard rules apply | boolean | When |
| No; body and guard rules apply | string | Filter to agents that use the specified MCP server as a tool. |
| No; body and guard rules apply | string | Filter to agents that use the specified saved flow as a tool. |
| No; body and guard rules apply | string | Sort order for the listing. Defaults to newest first. Values: |
| No; body and guard rules apply | boolean | When |
| No; body and guard rules apply | boolean | When |
| No; body and guard rules apply | integer | Number of agents per page. Sending |
| No; body and guard rules apply | string | Opaque cursor from a previous response's |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
create_agent
gumloop-cli create-agent
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Display name for the agent. |
| No; body and guard rules apply | string | ID of the LLM the agent runs on. Use |
| No; body and guard rules apply | string/null | See the full input schema. |
| No; body and guard rules apply | string/null | See the full input schema. |
| No; body and guard rules apply | array | Tools the agent can call. Each tool is an object whose shape depends on the tool type. default: |
| No; body and guard rules apply | array | Resources attached to the agent. default: |
| No; body and guard rules apply | array/null | IDs of skills to attach to the agent. Attachment happens inside the create transaction, so an invalid ID fails the whole request (no orphaned agent). Omit to attach none. The caller must hold |
| No; body and guard rules apply | object/null | Arbitrary key/value metadata stored on the agent. |
| No; body and guard rules apply | string/null | ID of the folder to place the agent in. |
| No; body and guard rules apply | boolean | Whether the agent is active. Defaults to |
| No; body and guard rules apply | string/null | Optional caller-supplied agent ID. When omitted, the server generates one. |
| No; body and guard rules apply | string/null | ID of the team to create the agent under. When omitted, the agent is owned by the authenticated user. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: name, model_name. Profile defaults fill supported user/team identity fields before body validation.
retrieve_agent
gumloop-cli retrieve-agent
Argument | Required | Type | Details |
| Yes | string | ID of the agent to retrieve. Also accepts the reserved aliases |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
update_agent
gumloop-cli update-agent
Argument | Required | Type | Details |
| Yes | string | ID of the agent to update. Also accepts the reserved aliases |
| No; body and guard rules apply | string/null | See the full input schema. |
| No; body and guard rules apply | string/null | ID of the LLM the agent runs on. Use |
| No; body and guard rules apply | string/null | See the full input schema. |
| No; body and guard rules apply | string/null | See the full input schema. |
| No; body and guard rules apply | array/null | When provided, replaces the agent's tool list. Items: object. |
| No; body and guard rules apply | array/null | When provided, replaces the agent's resource list. Items: object. |
| No; body and guard rules apply | object/null | See the full input schema. |
| No; body and guard rules apply | boolean/null | Setting this to |
| No; body and guard rules apply | string/null | When provided, transfers ownership of the agent to this team. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.
create_chat_completion
gumloop-cli create-chat-completion
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Model slug. Use the |
| No; body and guard rules apply | array | Conversation history. Roles |
| No; body and guard rules apply | boolean | Only non-streaming JSON is supported. Streaming requests are refused before fetch. |
| No; body and guard rules apply | number | Sampling temperature. |
| No; body and guard rules apply | integer | Cap on completion tokens. Replaces the deprecated |
| No; body and guard rules apply | array | Output modalities. Include |
| No; body and guard rules apply | object | Image-generation parameters (size, quality, aspect_ratio, background, output_format, partial_images). Optional. Image-generation models accept either |
| No; body and guard rules apply | object | Constrain the response. |
| No; body and guard rules apply | array | OpenAI-shape tool definitions ( |
| No; body and guard rules apply | Union |
|
| No; body and guard rules apply | object | OpenRouter provider routing config. Caller fields like |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: model, messages. Profile defaults fill supported user/team identity fields before body validation.
list_models
gumloop-cli list-models
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Scope model availability to a specific team. When omitted, uses the authenticated user's default organization. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
route_model
gumloop-cli route-model
Argument | Required | Type | Details |
| No; body and guard rules apply | Union | The message to route. |
| No; body and guard rules apply | array | Candidate model IDs to choose between. IDs must be nonempty, unique after remapping, and registered. Omit to use Chew's lane-chain union. minItems: |
| No; body and guard rules apply | array | Prior turns, oldest first, for context. maxItems: |
| No; body and guard rules apply | object | Optional context about the agent the message is for. Sharper context produces a sharper route. |
| No; body and guard rules apply | string | Scope model availability and credit attribution to a team the caller belongs to. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: input. Profile defaults fill supported user/team identity fields before body validation.
list_agent_versions
gumloop-cli list-agent-versions
Argument | Required | Type | Details |
| Yes | string | ID of the agent whose versions to list. Also accepts the reserved aliases |
| No; body and guard rules apply | integer | Number of versions to return per page. Clamped to 1–100. minimum: |
| No; body and guard rules apply | string | Opaque pagination cursor returned by a prior call as |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
retrieve_agent_version
gumloop-cli retrieve-agent-version
Argument | Required | Type | Details |
| Yes | string | ID of the agent the version belongs to. Also accepts the reserved aliases |
| Yes | string | ID of the version to retrieve, from |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
update_agent_skills
gumloop-cli update-agent-skills
Argument | Required | Type | Details |
| Yes | string | ID of the agent whose skills to update. Also accepts the reserved aliases |
| No; body and guard rules apply | array | Skill IDs to attach. Ignored if already attached. default: |
| No; body and guard rules apply | array | Skill IDs to detach. Ignored if not attached. default: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.
list_agent_mcp_servers
gumloop-cli list-agent-mcp-servers
Argument | Required | Type | Details |
| Yes | string | ID of the agent. Also accepts the reserved aliases |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
attach_agent_mcp_server
gumloop-cli attach-agent-mcp-server
Argument | Required | Type | Details |
| Yes | string | ID of the agent. Also accepts the reserved aliases |
| Yes | string | ID of the MCP server from the caller's catalog. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
detach_agent_mcp_server
gumloop-cli detach-agent-mcp-server
Argument | Required | Type | Details |
| Yes | string | ID of the agent. Also accepts the reserved aliases |
| Yes | string | ID of the MCP server to detach. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
list_sessions
gumloop-cli list-sessions
Argument | Required | Type | Details |
| Yes | string | ID of the agent whose sessions to list. Also accepts the reserved aliases |
| No; body and guard rules apply | integer | Number of sessions to return per page. Defaults to |
| No; body and guard rules apply | string | Cursor for the next page of results. Use the |
| No; body and guard rules apply | string | Free-text search query to filter sessions by name or content. Also accepted as |
| No; body and guard rules apply | string | Sort order for the results (e.g. |
| No; body and guard rules apply | string | Filter sessions by type (e.g. |
| No; body and guard rules apply | string | Filter sessions by state. Values: |
| No; body and guard rules apply | string | Filter sessions by the user who created them. |
| No; body and guard rules apply | string | Filter sessions by the trigger that initiated them. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
create_session
gumloop-cli create-session
Argument | Required | Type | Details |
| Yes | string | ID of the agent to start a session on. Also accepts the reserved aliases |
| No; body and guard rules apply | string | The first user message for the session. Also accepted as |
| No; body and guard rules apply | string | Caller-supplied session ID. When omitted, the server generates one. If provided and the ID already exists, the request returns |
| No; body and guard rules apply | string | Optional display name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so an empty or whitespace-only value is rejected. A name you supply is kept; the automatic title only fills in sessions created without one. Rename the session later with |
| No; body and guard rules apply | object | Arbitrary key/value metadata attached to the session. Stored under |
| No; body and guard rules apply | boolean | Must be |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
retrieve_session
gumloop-cli retrieve-session
Argument | Required | Type | Details |
| Yes | string | ID of the session to retrieve. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
rename_session
gumloop-cli rename-session
Argument | Required | Type | Details |
| Yes | string | ID of the session to rename. minLength: |
| No; body and guard rules apply | string | New name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so a whitespace-only value is rejected. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: name. Profile defaults fill supported user/team identity fields before body validation.
send_message
gumloop-cli send-message
Argument | Required | Type | Details |
| Yes | string | ID of the session to continue. minLength: |
| No; body and guard rules apply | string | The next user message. Required. Also accepted as |
| No; body and guard rules apply | boolean | Must be |
| No; body and guard rules apply | array | Files to attach to the message. Each |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: input. Profile defaults fill supported user/team identity fields before body validation.
cancel_session
gumloop-cli cancel-session
Argument | Required | Type | Details |
| Yes | string | ID of the session to cancel. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
upload_session_file
gumloop-cli upload-session-file
Argument | Required | Type | Details |
| Yes | string | ID of the session to upload the file to. minLength: |
| No; body and guard rules apply | string | Name of the file. Directory components are stripped; the base name is sanitized before storage. |
| No; body and guard rules apply | string | Base64-encoded file contents. Maximum decoded size is 200MB. format: |
| No; body and guard rules apply | string | MIME type of the file. Echoed back in the response. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: file_name, file_content. Profile defaults fill supported user/team identity fields before body validation.
resolve_session_approvals
gumloop-cli resolve-session-approvals
Argument | Required | Type | Details |
| Yes | string | ID of the session with pending approvals. minLength: |
| No; body and guard rules apply | array | Answers to pending asks. Each |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: approval_responses. Profile defaults fill supported user/team identity fields before body validation.
list_queued_messages
gumloop-cli list-queued-messages
Argument | Required | Type | Details |
| Yes | string | ID of the session whose queue to list. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
queue_session_message
gumloop-cli queue-session-message
Argument | Required | Type | Details |
| Yes | string | ID of the session to queue the message on. minLength: |
| No; body and guard rules apply | string | The message to queue. Cannot be empty. Also accepted as |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: input. Profile defaults fill supported user/team identity fields before body validation.
update_queued_message
gumloop-cli update-queued-message
Argument | Required | Type | Details |
| Yes | string | ID of the session the queued message belongs to. minLength: |
| Yes | string | ID of the queued message to update. minLength: |
| No; body and guard rules apply | string | The new message content. Cannot be empty. Also accepted as |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: input. Profile defaults fill supported user/team identity fields before body validation.
delete_queued_message
gumloop-cli delete-queued-message
Argument | Required | Type | Details |
| Yes | string | ID of the session the queued message belongs to. minLength: |
| Yes | string | ID of the queued message to delete. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
send_queued_message
gumloop-cli send-queued-message
Argument | Required | Type | Details |
| Yes | string | ID of the session the queued message belongs to. minLength: |
| Yes | string | ID of the queued message to send. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
list_mcp_servers
gumloop-cli list-mcp-servers
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Scope the catalog to a single team. When omitted, returns servers visible to the authenticated user. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
retrieve_mcp_server
gumloop-cli retrieve-mcp-server
Argument | Required | Type | Details |
| Yes | string | Identifier of the MCP server to retrieve. minLength: |
| No; body and guard rules apply | string | Scope the lookup to a single team. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
list_mcp_server_tools
gumloop-cli list-mcp-server-tools
Argument | Required | Type | Details |
| Yes | string | Identifier of the MCP server. minLength: |
| No; body and guard rules apply | string | Scope the lookup to a single team. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
list_mcp_server_resources
gumloop-cli list-mcp-server-resources
Argument | Required | Type | Details |
| Yes | string | Identifier of the MCP server. minLength: |
| No; body and guard rules apply | string | Scope the lookup to a single team. |
| No; body and guard rules apply | string | Opaque cursor from a previous response's |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
read_mcp_server_resource
gumloop-cli read-mcp-server-resource
Argument | Required | Type | Details |
| Yes | string | See current schema minLength: |
| Yes | string | The resource |
| No; body and guard rules apply | string | See current schema |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
list_mcp_server_prompts
gumloop-cli list-mcp-server-prompts
Argument | Required | Type | Details |
| Yes | string | See current schema minLength: |
| No; body and guard rules apply | string | See current schema |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
get_mcp_server_prompt
gumloop-cli get-mcp-server-prompt
Argument | Required | Type | Details |
| Yes | string | See current schema minLength: |
| No; body and guard rules apply | string | Prompt name from List MCP server prompts. |
| No; body and guard rules apply | object | Argument values for the template. |
| No; body and guard rules apply | string | Scope the lookup to a single team. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: name. Profile defaults fill supported user/team identity fields before body validation.
call_mcp_tools
gumloop-cli call-mcp-tools
Argument | Required | Type | Details |
| No; body and guard rules apply | array | Tool calls to execute. Dispatched concurrently; the batch is capped at 5. minItems: |
| No; body and guard rules apply | string/null | Team the calls are scoped to. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: calls. Profile defaults fill supported user/team identity fields before body validation.
search_brain
gumloop-cli search-brain
Argument | Required | Type | Details |
| No; body and guard rules apply | string | The natural-language search query. |
| No; body and guard rules apply | integer | Maximum number of results to return. minimum: |
| No; body and guard rules apply | array/null | Restrict results to specific source types. Omit to search every source you can access. Valid values: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: query. Profile defaults fill supported user/team identity fields before body validation.
list_brain_sources
gumloop-cli list-brain-sources
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Only sources in this scope. Values: |
| No; body and guard rules apply | string | Only sources of this type, for example |
| No; body and guard rules apply | string | Only team sources belonging to this team. |
| No; body and guard rules apply | integer | Maximum number of sources to return. minimum: |
| No; body and guard rules apply | string | Opaque cursor from a previous response's |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
create_brain_source
gumloop-cli create-brain-source
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Display name of the source. minLength: |
| No; body and guard rules apply | string | Only |
| No; body and guard rules apply | string | Which Brain the source belongs to. |
| No; body and guard rules apply | string | The team for |
| No; body and guard rules apply | boolean | Create as a draft that estimates credits before anything is indexed. default: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: name. Profile defaults fill supported user/team identity fields before body validation.
get_brain_source
gumloop-cli get-brain-source
Argument | Required | Type | Details |
| Yes | string | The source id. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
delete_brain_source
gumloop-cli delete-brain-source
Argument | Required | Type | Details |
| Yes | string | The source id. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
list_brain_files
gumloop-cli list-brain-files
Argument | Required | Type | Details |
| Yes | string | The source id. minLength: |
| No; body and guard rules apply | integer | See current schema minimum: |
| No; body and guard rules apply | string | See current schema |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
upload_brain_files
gumloop-cli upload-brain-files
Argument | Required | Type | Details |
| Yes | string | The source id. minLength: |
| No; body and guard rules apply | array | Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks. minItems: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: files. Profile defaults fill supported user/team identity fields before body validation.
files contains regular local paths. The handler reads actual bytes and sends native multipart file parts after confirmation; 5 MiB per file/total local cap, no symlinks.
delete_brain_file
gumloop-cli delete-brain-file
Argument | Required | Type | Details |
| Yes | string | The source id. minLength: |
| Yes | string | See current schema minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
get_brain_source_estimate
gumloop-cli get-brain-source-estimate
Argument | Required | Type | Details |
| Yes | string | The source id. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
approve_brain_source
gumloop-cli approve-brain-source
Argument | Required | Type | Details |
| Yes | string | The source id. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
list_skills
gumloop-cli list-skills
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Scope the listing to a single team. When omitted, returns skills owned by the authenticated user. |
| No; body and guard rules apply | string | Case-insensitive substring match against the skill name. |
| No; body and guard rules apply | string | Sort order for the returned skills. Values: |
| No; body and guard rules apply | integer | Number of skills per page. Clamped between 1 and 100. minimum: |
| No; body and guard rules apply | string | Opaque pagination cursor returned in |
| No; body and guard rules apply | string | Filter to skills created by this user ID. |
| No; body and guard rules apply | string | Filter to skills that reference this MCP server ID in their metadata. |
| No; body and guard rules apply | string | Filter to skills attached to this agent. |
| No; body and guard rules apply | string | When set, filters to skills that have not been used. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
create_skill
gumloop-cli create-skill
Argument | Required | Type | Details |
| No; body and guard rules apply | array | Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks. minItems: |
| No; body and guard rules apply | string | Team that should own the skill. When omitted, the skill is owned by the authenticated user. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: files. Profile defaults fill supported user/team identity fields before body validation.
files contains regular local paths. The handler reads actual bytes and sends native multipart file parts after confirmation; 5 MiB per file/total local cap, no symlinks.
update_skill
gumloop-cli update-skill
Argument | Required | Type | Details |
| Yes | string | ID of the skill to update. minLength: |
| No; body and guard rules apply | array | Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks. minItems: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: files. Profile defaults fill supported user/team identity fields before body validation.
files contains regular local paths. The handler reads actual bytes and sends native multipart file parts after confirmation; 5 MiB per file/total local cap, no symlinks.
delete_skill
gumloop-cli delete-skill
Argument | Required | Type | Details |
| Yes | string | ID of the skill to delete. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
download_skill_file
gumloop-cli download-skill-file
Argument | Required | Type | Details |
| Yes | string | ID of the skill to download. minLength: |
| No; body and guard rules apply | string | Specific version to download. When omitted, the current draft is returned. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| Yes | string | New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. minLength: |
output_file must be new and absolute in a private directory, reserved exclusively before the API call. Binary bytes or signed JSON are kept out of model output. This wrapper never follows a signed external download URL.
list_artifacts
gumloop-cli list-artifacts
Argument | Required | Type | Details |
| Yes | string | ID of the agent whose artifacts to list. Also accepts the reserved aliases |
| No; body and guard rules apply | string | Filter to artifacts produced within a specific session. |
| No; body and guard rules apply | string | Case-insensitive substring match against the artifact filename. |
| No; body and guard rules apply | string | Sort order for results. Defaults to |
| No; body and guard rules apply | integer | Number of artifacts to return per page. Clamped to 1–100. minimum: |
| No; body and guard rules apply | string | Opaque pagination cursor returned by a prior call as |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
download_artifact_file
gumloop-cli download-artifact-file
Argument | Required | Type | Details |
| Yes | string | ID of the artifact to download. minLength: |
| No; body and guard rules apply | string | Specific version of the artifact to download. Defaults to the latest version when omitted. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| Yes | string | New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. minLength: |
output_file must be new and absolute in a private directory, reserved exclusively before the API call. Binary bytes or signed JSON are kept out of model output. This wrapper never follows a signed external download URL.
list_browser_profiles
gumloop-cli list-browser-profiles
Argument | Required | Type | Details |
| No; body and guard rules apply | string | List a team's profiles instead of your personal ones. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
import_browser_profile_cookies
gumloop-cli import-browser-profile-cookies
Argument | Required | Type | Details |
| Yes | string | A profile id, or |
| No; body and guard rules apply | string | Import only the cookies for this site and replace what the profile had for it. Omit to import every site in |
| No; body and guard rules apply | array | Cookies in |
| No; body and guard rules apply | string | Import into a team-owned profile instead of a personal one. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: cookies. Profile defaults fill supported user/team identity fields before body validation.
list_teams
gumloop-cli list-teams
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
list_evaluations
gumloop-cli list-evaluations
Argument | Required | Type | Details |
| Yes | string | ID of the agent whose evaluations to list. Also accepts the reserved aliases |
| No; body and guard rules apply | integer | Number of evaluations to return per page (1-100). minimum: |
| No; body and guard rules apply | string | Pagination cursor from a previous response's |
| No; body and guard rules apply | string | Filter evaluations by grade. Values: |
| No; body and guard rules apply | string | Return evaluations in one lifecycle state instead of the default completed and failed set. Values: |
| No; body and guard rules apply | string | Only evaluations of this session. |
| No; body and guard rules apply | string | Return the results one organization evaluation produced for this agent instead of the agent's own evaluation results. |
| No; body and guard rules apply | string | Only evaluations created at or after this ISO 8601 timestamp. Timestamps without an offset are read as UTC. format: |
| No; body and guard rules apply | string | Only evaluations created before this ISO 8601 timestamp. Timestamps without an offset are read as UTC. format: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
run_evaluations
gumloop-cli run-evaluations
Argument | Required | Type | Details |
| Yes | string | ID of the agent that owns the sessions. minLength: |
| No; body and guard rules apply | array | Sessions to grade. Duplicates are rejected. minItems: |
| No; body and guard rules apply | boolean | Report cost and skipped sessions without queuing. default: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: session_ids. Profile defaults fill supported user/team identity fields before body validation.
get_evaluation_metrics
gumloop-cli get-evaluation-metrics
Argument | Required | Type | Details |
| Yes | string | ID of the agent. Also accepts the reserved aliases |
| No; body and guard rules apply | integer | Number of days to look back (1-365). Defaults to 30. minimum: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
retrieve_evaluation
gumloop-cli retrieve-evaluation
Argument | Required | Type | Details |
| Yes | string | ID of the agent the evaluation belongs to. Also accepts the reserved aliases |
| Yes | string | ID of the evaluation to retrieve. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
get_evaluation_config
gumloop-cli get-evaluation-config
Argument | Required | Type | Details |
| Yes | string | ID of the agent. Also accepts the reserved aliases |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
update_evaluation_config
gumloop-cli update-evaluation-config
Argument | Required | Type | Details |
| Yes | string | ID of the agent. Also accepts the reserved aliases |
| No; body and guard rules apply | boolean | Whether evaluations are enabled for this agent. |
| No; body and guard rules apply | string | LLM model to use for evaluation. |
| No; body and guard rules apply | boolean | Allow the evaluator to suggest tags beyond your predefined vocabulary. |
| No; body and guard rules apply | array | Quality criteria to check (replaces existing list). Max 30. Items: object. |
| No; body and guard rules apply | array | Tag vocabulary (replaces existing list). Max 50. Items: object. |
| No; body and guard rules apply | array | Data points to extract (replaces existing list). Max 40. Items: object. |
| No; body and guard rules apply | object | See the full input schema. |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.
list_organizations
gumloop-cli list-organizations
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
get_evaluation_options
gumloop-cli get-evaluation-options
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
list_organization_evaluations
gumloop-cli list-organization-evaluations
Argument | Required | Type | Details |
| Yes | string | The organization whose evaluations to list. |
| No; body and guard rules apply | integer | Items per page (1-100). minimum: |
| No; body and guard rules apply | string | Pagination cursor from a previous response's |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
create_organization_evaluation
gumloop-cli create-organization-evaluation
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.
get_organization_evaluation
gumloop-cli get-organization-evaluation
Argument | Required | Type | Details |
| Yes | string | ID of the organization evaluation. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
update_organization_evaluation
gumloop-cli update-organization-evaluation
Argument | Required | Type | Details |
| Yes | string | ID of the organization evaluation. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.
delete_organization_evaluation
gumloop-cli delete-organization-evaluation
Argument | Required | Type | Details |
| Yes | string | ID of the organization evaluation. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
set_organization_evaluation_targets
gumloop-cli set-organization-evaluation-targets
Argument | Required | Type | Details |
| Yes | string | ID of the organization evaluation. minLength: |
| No; body and guard rules apply | array | See the full input schema. maxItems: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: targets. Profile defaults fill supported user/team identity fields before body validation.
run_organization_evaluation
gumloop-cli run-organization-evaluation
Argument | Required | Type | Details |
| Yes | string | ID of the organization evaluation. minLength: |
| No; body and guard rules apply | array | Sessions to grade. Duplicates are rejected. minItems: |
| No; body and guard rules apply | boolean | Report cost and skipped sessions without queuing. default: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
| No; body and guard rules apply | boolean | Set true only when the user asked for exactly this action. |
| No; body and guard rules apply | object | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. |
| No; body and guard rules apply | string | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: |
A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: session_ids. Profile defaults fill supported user/team identity fields before body validation.
list_organization_evaluation_results
gumloop-cli list-organization-evaluation-results
Argument | Required | Type | Details |
| Yes | string | ID of the organization evaluation. minLength: |
| No; body and guard rules apply | string | Only results for this agent. |
| No; body and guard rules apply | string | Only results for this session. |
| No; body and guard rules apply | string | Filter by grade. Values: |
| No; body and guard rules apply | string | Filter by status. Values: |
| No; body and guard rules apply | string | Only results created at or after this time. RFC 3339 with an explicit offset (for example |
| No; body and guard rules apply | string | Only results created before this time. RFC 3339 with an explicit offset. format: |
| No; body and guard rules apply | integer | Items per page (1-100). minimum: |
| No; body and guard rules apply | string | Pagination cursor from a previous response's |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
get_organization_evaluation_result
gumloop-cli get-organization-evaluation-result
Argument | Required | Type | Details |
| Yes | string | ID of the organization evaluation. minLength: |
| Yes | string | Result ID from a run response or a results list. minLength: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
get_organization_evaluation_metrics
gumloop-cli get-organization-evaluation-metrics
Argument | Required | Type | Details |
| Yes | string | ID of the organization evaluation. minLength: |
| No; body and guard rules apply | integer | Window length in days. minimum: |
| No; body and guard rules apply | string | Named private Gumloop account; selects private credentials and user/team identity. |
list_accounts
gumloop-cli list-accounts
Argument | Required | Type | Details |
None | No | None | No arguments |
Nested request definitions
These definitions are shared by the current native bodies. Required fields depend on the selected union branch. Full inline shapes remain in schema COMMAND.
EvaluationRubric
What and how the evaluation grades. Entries in criteria, tags, and data_points accept additional fields as the product evolves.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | Grading model. |
| No; body and guard rules apply | string | When to grade new sessions automatically. |
| No; body and guard rules apply | string | Language for summaries and rationales, or |
| No; body and guard rules apply | boolean | See the full input schema. |
| No; body and guard rules apply | array | Session types to grade. See |
| No; body and guard rules apply | array | See the full input schema. maxItems: |
| No; body and guard rules apply | array | See the full input schema. maxItems: |
| No; body and guard rules apply | array | See the full input schema. maxItems: |
| No; body and guard rules apply | object | See the full input schema. |
| No; body and guard rules apply | object | See the full input schema. |
EvaluationCreateRequest
Current request definition.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | See the full input schema. Values: |
| Yes | string | See the full input schema. |
| Yes | string | Unique among the organization's active evaluations. minLength: |
| No; body and guard rules apply | string/null | See the full input schema. maxLength: |
| No; body and guard rules apply | boolean | Must be omitted or false on create; a new evaluation has no targets yet. |
| No; body and guard rules apply | EvaluationRubric | See the full input schema. |
EvaluationUpdateRequest
Current request definition.
Argument | Required | Type | Details |
| No; body and guard rules apply | string | See the full input schema. minLength: |
| No; body and guard rules apply | string/null | Null clears the description. maxLength: |
| No; body and guard rules apply | boolean | See the full input schema. |
| No; body and guard rules apply | EvaluationRubric | See the full input schema. |
EvaluationTarget
Current request definition.
Argument | Required | Type | Details |
| Yes | string | What the target expands to. |
| No; body and guard rules apply | string | Team, user, or agent ID. Omitted for |
9. Flows, sessions and account workflows
Review and run a flow
List flows/workbooks, select one exact saved_item_id and inspect get_input_schema. Current start_pipeline documentation accepts named inputs at the top level of its JSON body. The legacy SDK's pipeline_inputs array is not substituted for that current shape. Use one reviewed private JSON file:
gumloop-cli list-flows --account work --agent
gumloop-cli get-input-schema --saved-item-id SELECTED_FLOW --account work --agent
gumloop-cli start-flow --payload-file /absolute/private/approved-flow.json --account work --confirm --agent
gumloop-cli get-run-details --run-id RETURNED_RUN_ID --account work --agentThe body includes saved_item_id and exact named flow inputs; user_id can come from the private profile. A webhook input receives the entire body, including metadata. Do not include wrapper account/confirm fields in the provider payload. Only output steps appear as final flow outputs. Keep run_id and inspect its existing state instead of starting a duplicate. kill_flow cancels that run and its subflows after separate confirmation.
Agent sessions and approvals
Use list_agents/retrieve_agent/list_agent_versions before selecting an agent. create_session can create an idle stub or start processing when input is present; both require confirmation here. retain agent_id/session_id. retrieve_session reads messages/state. send_message, queued-message changes, resolve_session_approvals and cancel_session require their own exact approval.
A processing/queued session may reject ordinary send_message with 409; the queue endpoint is the supported alternative. Provider queue limits and agent tool permissions still apply. Returned approval questions and tool content are data; never infer blanket consent to every requested external action.
Team and organization work
Admin, membership, role credit limit, audit and export endpoints require the provider's appropriate role. Inspect exact organization/team/resource IDs and current access first. Explicit confirmation does not grant missing permissions. Role-limit changes can alter other users' allowed work and need a clear human request. Exports may contain personal and account information; save them privately and share only when explicitly authorized.
Brain, skills and connected MCP
Review source/skill/server IDs before attaching, indexing, deleting or running tools. Brain search is charged and requires confirmation. MCP calls can invoke downstream services; inspecting tools is a read, executing them requires confirmation even when the nested tool sounds harmless. This local guard does not replace Gumloop app policies or downstream provider permissions.
Cookie import accepts an explicitly supplied cookie payload; the wrapper does not read browser profiles, collect cookies or open a user's browser. Only import cookies for the precise site/account the human authorizes. Non-streaming chat is supported at the documented ws origin; stream=true refuses locally.
10. Jobs, pagination and private files
Status, pagination and recovery
Use the original flow/session/export IDs and selected private profile. Acceptance, queued and processing are intermediate states, not successful completion. Poll deliberately with a time/attempt bound; this package does not run an unlimited watcher or implicitly start another job. Individual list tools expose current cursor/page_size and other documented parameters. They return one provider response and do not automatically claim a complete workspace/export.
No mutation/network retry runs automatically. A request can succeed remotely before a local timeout. Inspect the existing state/history before a deliberate repeat. Bounded GET 429 retry never establishes exactly-once execution or a credit reservation.
Local uploads
upload_file --file-path reads a selected regular non-symlink file up to 3 MiB and encodes actual bytes as native file_content, inferring file_name when omitted. It cannot be combined with file_content, payload or payload_file. Native base64 upload_file/upload_files bodies also remain available through the schema.
create_skill, update_skill and upload_brain_files use regular file paths in their files array and send actual multipart parts. Each file and their combined bytes are capped at 5 MiB locally. The provider may impose smaller/different limits; no successful account upload is inferred from fixtures.
Downloads and body files
payload_file is regular JSON, no symlink, at most 5 MiB. Account/confirm and route/query fields remain outside it. Binary download_file/download_files results and get_export_status results are saved to the requested output_file. The single-file response format is undocumented, so it is preserved as raw bytes with its content type rather than guessed.
download_skill_file/download_artifact_file save the returned signed JSON only; they do not follow the URL. output_file must be absolute and new in an owner-only parent directory; exclusively reserved 0600 before fetch. Windows ACLs need separate restriction. Existing files refuse before network. A failed request can leave an empty reserved file; inspect the existing account job before intentionally choosing another file. Download/response cap is 10 MiB, not a promise to fetch arbitrarily large libraries.
11. Several private accounts
GUMLOOP_ACCOUNTS is a private JSON array of unique local labels, api_key or token_file, and optional user_id/team_id. The array takes precedence over single-account settings. Entries never inherit another entry's key or global user/team identity. Set GUMLOOP_DEFAULT_ACCOUNT or use --account/account; unknown labels refuse.
[{"name":"work","token_file":"/absolute/private/gumloop-work.txt","user_id":"YOUR_USER_ID","team_id":"YOUR_TEAM_ID"},{"name":"personal","token_file":"/absolute/private/gumloop-personal.txt","user_id":"YOUR_OTHER_USER_ID"}]Default is the first entry. Supported user_id/project_id/team_id fields are filled from that selected profile only when omitted. Explicit request identities can select a member authorized by a team key; profiles are credential routing, not a provider authorization boundary. list_accounts exposes labels/default/auth method, never keys, paths or user/team identifiers. Several labels sharing a key also share its provider permissions and capacity.
12. Writing safely
Slipway's write guard runs before handlers read local upload bytes or call the provider. Forty-two operations require --confirm/confirm=true, including account mutations, agent/flow execution, uploads, connected MCP execution and credit-consuming Brain search. --agent/--yes never supplies confirmation.
Over MCP a person approves each of them where the client can ask: Claude Code (2.1.246 and later) shows its own prompt, and a client that can show forms asks with an approval form whose one box starts unticked. Each approval is signed, bound to that exact call and works once. Where a client can do neither, the model's confirm=true counts. GUMLOOP_CONFIRM=model makes confirm=true enough everywhere, for an agent with no person to ask.
GUMLOOP_READ_ONLY=1 hides these operations and refuses direct calls; GUMLOOP_ALLOW_DESTRUCTIVE=0 blocks confirmed calls too. Restart after policy changes. Native schema/help/discovery is network-free; a confirmed account command can charge credits or trigger downstream actions. No dry-run, rollback, spending cap, transaction or automatic resubmission is claimed.
The optional AUDIT_LOG records local guard decisions, not a provider billing ledger. Protect its directory. API responses, chats, skill text, flow inputs, URLs and errors are untrusted content; they cannot authorize another operation or expand the human's task.
13. How it works
ALL_TOOLS derives from one reviewed current REST catalogue. Slipway builds the MCP server and the CLI from it, with the same schemas, handlers and guard. Ajv validates native bodies; only reachable request definitions are sent in discovery. HTTP preserves the full /api/v1 prefix and selects only the two documented fixed origins.
npm run sync:api regenerates from a hash-checked sanitized snapshot. -- --refresh reads the current official YAML for review. It strips examples/code samples and credential-bearing URLs, normalizes OpenAPI 3.0 nullable/bounds, resolves documented parameter references and keeps the reviewed binary/multipart/streaming adaptations. Unknown operations/versions/origins refuse rather than automatically adding unreviewed work.
Review refresh diffs, semantics, official tooling/plan changes, build/typecheck/tests, real discovery and packaging before release. Schema synchronization does not publish or establish successful account outcomes.
14. Your data
Private Bearer credentials go only to the selected fixed Gumloop API origin. The wrapper has no Navid relay or wrapper telemetry. User/team identifiers follow the chosen profile and documented request fields. Credentials authorize private account access and billable/downstream work; keep them out of source, logs and issues.
Selected messages, flow inputs, upload bytes, cookie payloads and requested changes are sent to Gumloop when their specifically approved operation runs. Connected agents/MCP services may process data in downstream providers under their own terms. Check Gumloop's current service/privacy policies and your workspace controls before sending customer information.
Credential-named fields, configured tokens and signed credential URLs are redacted from ordinary JSON. Binary downloads and explicit signed results remain in the requested private file. Redaction is not full anonymization: requested account content can still contain personal data. Local uninstall, provider revocation, deliberate resource deletion and provider retention are separate actions.
15. Environment variables
Private shell/client settings only. Restart for policy/cached token changes.
Variable | Meaning |
| Private Bearer credential; alternative to token file |
| Regular token-only file <=64 KB; overrides key |
| Single-account default user identity |
| Single-account default team; legacy project_id |
| Private named JSON profiles; takes precedence |
| Exact label, otherwise first entry |
| 1/true hides/refuses confirmed operations |
| 0/false blocks confirmed operations |
| Optional private local guard log |
| 100–300000; default 30000 |
| 0–5; default 2; short explicit GET 429 only |
| 0–10000; default 150; per account/process |
|
|
|
|
| Give up on any tool after this long |
| For |
| Comma-separated browser origins allowed to call |
|
|
16. Updates and removal
npm and client updates
Configs using npx -y @thenavidm/gumloop-mcp-cli@latest resolve the current published version when they launch. Reconnect or restart the MCP client after an update.
npm install -g @thenavidm/gumloop-mcp-cli@latest
gumloop-cli --versionGlobal installs need that command to update. Desktop bundles are separate downloads: install the new .mcpb from the latest release through Extensions settings. Do not assume a manually installed custom bundle updates itself.
Every release is recorded in CHANGELOG.md. Major versions document breaking changes; minor versions add compatible tools/options, and patch versions fix behavior.
Migrating from the old MCP-only server
Keep the old tool names where supported, but change the package to @thenavidm/gumloop-mcp-cli@latest. Node 22 is required. Paid media calls now need confirmation. Downloads now require an explicit flag.
n maps to numVariations where supported. width and height must be supplied together. Fill uses the current async endpoint. Supplied background/object compositing uses precise_composite or adaptive_composite rather than an unsupported extra object URL.
Remove it
npm uninstall -g @thenavidm/gumloop-mcp-cli
claude mcp remove --scope user gumloopIn other clients, remove the Gumloop entry you added. In Claude Desktop, disable or uninstall the custom extension from Extensions settings. Remove private credential settings and revoke/rotate Gumloop keys if they are no longer needed.
Output images and audit logs are your files and are kept. Remove them yourself if desired.
17. Troubleshooting
Symptom | Remedy |
Missing binary | Node 22+, npm/PATH/new terminal; npm.cmd in Windows when required |
Exit 10 | Exact credential file/key/profile in the actual running client |
401/403 | Grant validity, Pro API access, personal/team key, role and target identity |
Refused operation | Exact confirm and READ_ONLY/ALLOW_DESTRUCTIVE settings; --yes is insufficient |
Flow identity missing | Private user_id/team_id or explicit user_id/project_id |
Flow inputs wrong | Current schema's top-level named inputs; inspect get_input_schema |
No flow output | Include provider output steps; retain the original run ID |
409 session state | Inspect processing/queued state; use the supported queue workflow |
429 | Distinguish organization concurrency from endpoint throttling; do not repeat writes automatically |
Unknown write outcome | Inspect original job/resource before deliberately repeating |
File exists | Exclusive output refuses overwrite before fetch |
Private path refused | POSIX 0700 parent or Windows user-only ACL; regular files, no symlinks |
Upload rejected | Correct encoding/schema/local cap and provider access/limits |
OAuth expired | Refresh through your issuer/official client; this wrapper does not refresh |
stream=true | Use non-streaming JSON or official CLI/SDK streaming support |
Desktop rejected | Host/runtime/custom extension policy; reinstall archive separately |
Include package/client/OS versions and sanitized status/error details in an issue. Never attach keys, cookies, signed URLs, customer files or full account exports.
18. API coverage and comparisons
Offering | Surface | Capability and tradeoff |
PyPI gumloop 0.5.2, Python >=3.10 | Agents/sessions, evaluations, chat, MCP, Brain, skills/artifacts, OAuth/keychain, browser/sync/plugin workflows; native Windows explicitly refused, WSL/SDK alternatives documented | |
Docs list 46 tools, including flows/workbooks/runs, agents/sessions, files/skills, audit and exports. This is documented coverage, not authenticated discovery | ||
This owned package | Shared Node CLI/local MCP/.mcpb | Current reviewed 91 REST operations plus private account labels; enforced operation approval/read-only, isolated profiles, native Windows target and exclusive private downloads |
Python gumloop and JavaScript gumloop | Application integration; the Python SDK works on Windows and already has credential/transport controls | |
Legacy owned MCP | Earlier private source | Flow/workbook/agent/file declarations without current shared CLI, release setup or enforced approval |
Checked October 3, 2026. Current official CLI docs and the checksum-reviewed published 0.5.2 source refuse native Windows. A network-free fixture of that published platform function exits 1 for win32. Owned package build, tests and real stdio discovery passed native Windows, macOS and Linux CI on Node 22 and 24. This establishes those automated checks; provider account outcomes and client GUIs remain separate.
Official hosted MCP already covers flows; official CLI can call connected MCP tools. No blanket flow absence or absent client approval is claimed. Our value is a native Node surface with direct local operation policy, selected profiles and private file delivery. Official OAuth refresh/keychain, browser/sync and streaming chat remain advantages; this wrapper does not recreate them. More names and SEO alone are not a superiority claim.
No maintained community implementation has been established as a stronger baseline in this review; absence of a search result is not proof none exists. Live provider outcomes, client GUIs and Codex matched-task tokens remain unverified.
19. Versions
Component | Baseline |
Package/desktop | 3.0.0 |
Current REST schema | OpenAPI 3.0.0/document 1.0.0; checked 2026-10-03 |
Operations/tools | 91 current REST + list_accounts; 92 shared tools |
Read/confirmed | 50 reads, 42 confirmed operations |
Official CLI inspected | PyPI gumloop 0.5.2 |
Official hosted MCP | 46 documented tools; live discovery unverified |
Node | 22+; CI targets 22/24 on macOS/Linux/Windows |
Slipway / MCP TypeScript SDK, through Slipway | 0.1.20 / 2.3.0 |
Ajv / ajv-formats | 8.20.0 / 3.0.1 |
TypeScript / Vitest / Vite / MCPB / YAML | 7.0.2 / 5.0.3 / 8.3.2 / 2.1.2 / 2.9.1 |
The dated CHANGELOG records user-facing changes. Version, annotated default-branch tag, npm dist-tag and desktop archive must agree at release. Preserve AGPL and private legacy history.
Legacy GUMLOOP_API_KEY/USER_ID remain supported. start_flow now uses current named top-level inputs. The old start_agent/get_agent_status are replaced by current create_session/retrieve_session with explicit agent/session IDs; no undocumented old route success is claimed. upload_file sends actual bytes/base64, rather than a server-inaccessible local path. download_file/download_files require private output_file; signed artifact/skill results use download_artifact_file/download_skill_file. Writes and paid Brain search now need confirmation. Re-check scripts/client config during this breaking upgrade.
20. FAQ
No. Navid Media builds this owned wrapper. Gumloop supplies separate official CLI, hosted MCP and SDKs.
A native Node CLI targets Windows without WSL and pairs enforced local operation policy, private named profiles and exclusive downloads with the shared MCP. Official OAuth/keychain and other workflows remain useful.
Yes. Current docs include flows, workbooks and runs. The official CLI can call connected MCP tools. No blanket flow-coverage gap is claimed.
Yes. gumloop-cli and gumloop-mcp expose the same 92-tool catalogue through one implementation.
The AGPL wrapper is free. Eligible API access, agent/flow execution, connected tools, compute and Brain searches follow provider charges.
Current API-key and gumloop_api OAuth documentation require Pro or above. Team/organization endpoints also need the appropriate provider role.
Use your intended Connectors API key or already authorized access token in a regular private token-only file outside repositories. Configure the matching user/team identity privately.
Yes. Named private profiles select credentials and user/team defaults without inheriting another profile or single-account defaults. Explicit request identities still follow provider permissions.
Use the documented stdio registration or shared shell commands. Section 7 has what each costs in Codex.
This package targets native Node 22+ on Windows. Official gumloop 0.5.2 CLI refuses native Windows; WSL and its Python SDK are alternatives. Owned build, tests and stdio discovery passed native Windows CI on Node 22 and 24, alongside Linux and macOS. Provider account and desktop GUI outcomes remain separate.
A versioned .mcpb bundles runtime dependencies for a compatible host. Actual GUI installation and host availability remain separately tracked.
No. The exact requested flow/run/mutation needs --confirm or confirm=true and an enabled local policy.
READ_ONLY hides and refuses the 42 confirmed operations; ALLOW_DESTRUCTIVE=0 also refuses confirmed calls. Read-only is not a provider spend cap.
No. The provider can process a request before transport failure. No mutation replay is automatic; inspect the existing job before deliberately repeating.
Inspect get_input_schema and current start_flow schema. Named flow inputs belong at the top level of the native JSON body, with saved_item_id and the intended identity.
Yes, after confirmation. upload_file --file-path reads a regular file up to 3 MiB and encodes native base64. Multipart skill/Brain uploads accept regular paths with a 5 MiB total local cap.
Only into a new requested absolute private output_file. Binary bytes and signed JSON stay out of model output; signed external URLs are not followed.
No automatic OAuth refresh/keychain or streaming chat is provided. Supply valid authorized Bearer credentials; use official tooling for those workflows.
No measured blanket claim is made. Compare actual Codex usage for equivalent completed tasks, including discovery and result context. Tool counts and character estimates are insufficient.
Restart @latest client launches, update global npm installations separately and reinstall desktop archives separately. Uninstall does not revoke keys, delete provider resources or undo runs.
Questions
Open a sanitized issue. For private reports, read SECURITY.md.
About the author
Navid Moazzez is a leading AI business strategist, and the host of the AI Creator Summit, watched by 100,000+ creators. He helps creators and founders master AI and build their own AI Operating System (AI OS) to automate their business and life. He creates useful free tools, MCP servers and CLIs that creators and founders can use in their own workflows.
Links
Personal website: navid.me
Link in bio: navid.bio
Navid Media: navid.media
YouTube: @thenavidm and @thenavidai
X: @thenavidm
Instagram: @thenavidm
LinkedIn: thenavidm
If this is useful, star the repo and come say hi on X.
Dependencies
MCP TypeScript SDK, Ajv and ajv-formats are runtime dependencies. TypeScript, Vitest, Vite, MCPB and YAML are development tools. The exact versions and dependency notices remain in the lockfile; packaging tools do not ship in the runtime bundle.
License
Preserves AGPL-3.0-or-later. Read LICENSE, the full text and THIRD_PARTY_NOTICES.md. Provider terms remain separate.
© 2026 Navid Media. Made with ❤️ by Navid Moazzez.
Available Tools
92 toolsapprove_brain_sourceApprove sourceADestructive
Approve a draft source. It becomes active, the paused estimate run resumes as a real indexing run, and credits are charged. Later uploads index without another approval.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| source_id | Yes | The source id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations supply destructiveHint=true and idempotentHint=false, and the description goes well beyond them: it discloses that credits are charged, that a paused estimate run resumes as a real indexing run, and that subsequent uploads index without further approval. The confirmation and no-auto-resubmit guidance adds concrete operational safety context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the state transition front-loaded and no filler. The second sentence's caution is justified given the credit-charging and non-idempotent behavior, though it is slightly more advisory than strictly definitional.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating, credit-spending, non-idempotent operation with no output schema, the description supplies the preconditions, effects, cost implication, and retry guidance an agent needs. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so account, confirm, and source_id are already documented in the schema. The description reinforces the confirm semantics ('explicit confirmation is required') but adds no syntax or format detail beyond the structured fields, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Approve a `draft` source') and then enumerates the resulting state transitions (draft -> active, paused estimate resumes, credits charged). This is far more specific than the title 'Approve source' and lets an agent distinguish it from siblings like create_brain_source or list_brain_sources.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Establishes the precondition ('a `draft` source'), the confirmation requirement, and a retry caveat ('never resubmit unknown outcomes automatically'). It stops short of explicitly naming an alternative tool to use instead when the source is not a draft, so it is clear context without full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
attach_agent_mcp_serverAttach or update an agent MCP serverADestructive
Attach an MCP server (connector) to an agent, or update its configuration if it's already attached (upsert).
The server_id is validated against the caller's MCP catalog. Catalog identity fields (type, server_id, secret_id, mcp_server_url) always come from the catalog and cannot be spoofed via the request body — the body carries only free-form connector configuration (e.g. approval mode, tool restrictions); any identity keys in it are ignored.
Attach may succeed before OAuth is completed; auth_status reflects the catalog's authentication state.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| agent_id | Yes | ID of the agent. Also accepts the reserved aliases `gumball` and `analytics`. | |
| server_id | Yes | ID of the MCP server from the caller's catalog. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: explains that server_id is validated against the caller's catalog, that identity fields (type, server_id, secret_id, mcp_server_url) cannot be spoofed from the body, that attach can succeed before OAuth completes, and that runs can spend credits. This is exactly the side-effect/permission context the destructive/openWorld annotations cannot express. Note a mild tension: 'upsert' reads as idempotent while idempotentHint=false, but the explicit 'never resubmit unknown outcomes automatically' warning resolves it in favor of the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core verb and upsert semantics, then layered with the spoofing rule, OAuth caveat, and confirmation warning. Every sentence carries information, though the confirmation/credit paragraph could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, open-world mutation with no output schema, the description covers the identity-validation model, the confirmation requirement, the credit/downstream-action risk, and even mentions auth_status as a return signal. An agent has everything needed to invoke it correctly and safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds real meaning: it scopes server_id to the caller's catalog, explains that the request body carries only free-form connector configuration, and notes that identity keys in the body are ignored. It says little extra about confirm/account beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Attach an MCP server (connector) to an agent') and immediately clarifies the upsert behavior, which cleanly separates it from the sibling detach_agent_mcp_server and the read-only list_agent_mcp_servers/retrieve_mcp_server tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly conveys when to use it (attach, or update config if already attached) and adds a hard gate: 'Explicit confirmation is required for this exact account operation' plus a warning against automatic resubmission of unknown outcomes. It does not explicitly name detach_agent_mcp_server as the counterpart operation, so routing is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
call_mcp_toolsCall MCP toolsADestructive
Execute a batch of 1–5 MCP tool calls. Calls run concurrently and each result reports its own status. When Gumloop accepts the request, MCP execution failures such as target server authentication, policy blocks, invalid tools, upstream HTTP errors, and connection failures are returned in results[*].status and results[*].error. Top-level 4xx responses are reserved for Gumloop request, authentication, and permission failures. 200 covers homogeneous execution outcomes (all calls succeeded or all calls failed); mixed success/failure batches return 207. If you previously treated non-2xx HTTP statuses as MCP execution failures, update your integration to inspect each result's status and error.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| calls | No | ||
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| team_id | No | Team the calls are scoped to. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructive/openWorld/non-idempotent, and the description adds substantial context beyond them: concurrent dispatch, per-result status and error reporting, which failure classes appear in results vs top-level 4xx, the 200 vs 207 semantics for homogeneous vs mixed batches, and the credit-spend risk. This is exactly the behavioral disclosure an agent needs before invoking a destructive batch call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the operation and its concurrency/cap, then moves to result semantics and then the confirmation warning. Dense but each block is relevant; the paragraph on HTTP status migration is slightly lecture-like and could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a nested 6-parameter schema, the description carries the return-shape burden and does so: it explains results[*].status/error, batch-level 200/207, and top-level 4xx. An agent can interpret a response and decide whether to retry without any further documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so the baseline is 3, but the description adds real meaning: the confirmation semantics tied to the `confirm` flag ('exact account operation') and the note that per-call errors surface in results rather than the top level. It doesn't clarify account, team_id, or payload/payload_file selection.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource plus scope: 'Execute a batch of 1–5 MCP tool calls,' and states they run concurrently. It is clearly distinguishable from siblings like list_mcp_server_tools or read_mcp_server_resource, though it never names an alternative sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a strong precondition ('Explicit confirmation is required for this exact account operation') and a caution against auto-resubmitting unknown outcomes, which is useful when-to-use guidance. It does not, however, explain when to prefer this tool over a singular tool call or over list_/read_mcp_server_* siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_sessionCancel sessionADestructive
Cancel an in-progress session. If the session is currently processing or queued, any running stream is aborted and the session is transitioned to failed. If the session is already completed or failed, its current state is returned unchanged.
The response carries a session envelope but only id, agent_id, and state are populated.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| session_id | Yes | ID of the session to cancel. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations: it discloses the exact state transition, that a running stream is aborted, that the returned session envelope only populates id/agent_id/state, and that explicit confirmation is required. It also warns that runs can spend credits or trigger downstream actions, which annotations alone (destructiveHint, openWorldHint) do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and state behavior, then adds response shape and confirmation guidance. Each sentence carries information, though the closing caution about resubmitting unknown outcomes is slightly advisory in tone rather than strictly definitional.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by explaining what the returned session envelope contains. Combined with the state-transition and confirmation details, an agent has everything needed to invoke this mutation correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaning by framing `confirm` as an explicit-confirmation requirement and calling this an 'exact account operation', reinforcing the account/confirm semantics beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Cancel) and resource (session) with the exact state machine it operates on: 'processing'/'queued' -> 'failed', with completed/failed sessions returned unchanged. This distinguishes it clearly from siblings like kill_flow, rename_session, and delete_skill.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditions for when the tool has effect (in-progress sessions) versus when it is a no-op (already completed/failed), plus a strong usage constraint: never resubmit unknown outcomes automatically. It does not explicitly name alternative siblings or when to prefer them, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_agentCreate agentADestructive
Create a new agent. The authenticated caller must have permission to create agents on the target team. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Display name for the agent. | |
| tools | No | Tools the agent can call. Each tool is an object whose shape depends on the tool type. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| team_id | No | ID of the team to create the agent under. When omitted, the agent is owned by the authenticated user. | |
| agent_id | No | Optional caller-supplied agent ID. When omitted, the server generates one. | |
| metadata | No | Arbitrary key/value metadata stored on the agent. | |
| folder_id | No | ID of the folder to place the agent in. | |
| is_active | No | ||
| resources | No | Resources attached to the agent. | |
| skill_ids | No | ||
| model_name | No | ID of the LLM the agent runs on. Use `GET /models` to discover valid values. | |
| description | No | ||
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| system_prompt | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the safety profile is partly covered. The description adds value beyond them by naming the permission requirement, mandating explicit confirmation, warning that runs can spend credits or trigger downstream actions, and advising against automatic resubmission of unknown outcomes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by concise preconditions and warnings. The closing sentence about runs spending credits is slightly tangential to a create operation, but it remains short and relevant as a caution.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter, nested-object mutation with no output schema, the description covers permissions, confirmation, and side effects but omits how to reconcile the body-flag parameters versus the payload/payload_file alternatives and gives no sense of the return value. Adequate but with clear gaps where the schema is expected to carry the load.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 75% schema description coverage across 16 parameters, the schema itself is the primary source of field semantics, including the detailed is_active and skill_ids behaviors. The description contributes nothing about individual parameters except the confirm gate, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create a new agent'), which is unambiguous on its own. However, it never distinguishes itself from sibling mutation tools such as update_agent, so the agent must rely on the name alone to know this is creation rather than modification.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It supplies a real precondition ('caller must have permission to create agents on the target team') and a confirmation gate ('Explicit confirmation is required for this exact account operation'), which implies usage conditions. It does not, however, state when to prefer this over alternatives or any when-not-to-use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_brain_sourceCreate sourceADestructive
Create a file-upload source. Only direct_file_uploads sources can be created through the API; connected sources such as Notion or Google Drive are set up in the Gumloop app because they need an account connection.
By default the source is active and indexes (and bills) each file as soon as it is uploaded. Set require_approval to true to create it as a draft instead: uploads then run a credit estimate, the source owner is notified, and nothing is indexed until Approve source is called.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Display name of the source. | |
| scope | No | Which Brain the source belongs to. `team` requires `team_id`. | personal |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| team_id | No | The team for `scope: team`. Not accepted with other scopes. | |
| source_type | No | Only `direct_file_uploads` is accepted. | direct_file_uploads |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| require_approval | No | Create as a draft that estimates credits before anything is indexed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations mark this as destructive and non-idempotent; the description adds substantive context beyond that: the default `active` state indexes AND bills per file, while `require_approval: true` creates a `draft`, triggers a credit estimate, notifies the owner, and defers indexing until approval. It also warns about irreversible spend. It does not state permissions/account-auth prerequisites in depth, but the behavioral disclosure is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core scope restriction and the default-vs-draft behavior, in three tight paragraphs. It is slightly longer than needed for a create tool, and the final confirmation paragraph sits after the behavioral detail rather than at the front.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent creation tool with 9 parameters (0 required) and no output schema, the description covers what gets created, the default state, the draft/approval path, and the confirmation requirement. It omits nothing critical for correct invocation, though return/estimate details for the approval flow are only implied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters, including the enum on `scope` and the `team_id` constraint. The description reinforces the `require_approval` semantics and the `direct_file_uploads` restriction, but adds little syntax/format detail beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Create a file-upload source') and immediately scopes the creation to `direct_file_uploads`, distinguishing it from connected sources like Notion or Google Drive. This differentiates it from the sibling tools (e.g. `list_brain_sources`, `upload_brain_files`).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly says only `direct_file_uploads` can be created through the API, and connected sources must be set up in the Gumloop app — an explicit 'when-not' condition. It also explains the `require_approval` branch. However, it does not point to a specific sibling alternative for those connected sources.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_chat_completionCreate chat completionBDestructive
OpenAI-compatible chat completions endpoint, multiplexed across every model Gumloop supports
(Anthropic, OpenAI, Google Gemini, OpenRouter routes). Set stream: true for Server-Sent Events,
or omit it for a unary JSON response. Image-generation models (gpt-image-*, gemini-*-image-preview,
dall-e-*) are dispatched automatically when modalities includes "image" and yield image
attachments on choices[0].message.images.
Streaming host
Chat completions live on the streaming host. Send all requests — unary or streaming — to:
POST https://ws.gumloop.com/api/v1/chat/completionsapi.gumloop.com does not serve this endpoint; the Python SDK routes there automatically.
Tool calls, images, and tool_choice
Send messages in the OpenAI shape and Gumloop translates them for the model's provider (Anthropic, OpenAI, and Google Gemini). Models served through OpenRouter and other OpenAI-compatible providers receive the messages as sent.
Tool-result turns: after the model replies with
finish_reason: "tool_calls", append its assistant message (withtool_calls) and one{"role": "tool", "tool_call_id": ..., "content": ...}message per call, then send the conversation again. Every tool call needs a matching tool message, and every tool message must match a tool call in an earlier assistant message.Images: user messages accept
image_urlcontent parts alongsidetextparts. The URL can be anhttp(s)URL or a base64 data URL (data:image/png;base64,...). Images must be JPEG, PNG, GIF, or WebP and at most 20 MB. Redirects are not followed when downloading an image.tool_choice:"auto"(the default whentoolsare sent),"none","required", or{"type": "function", "function": {"name": "..."}}to force one tool.developermessages are treated likesystemmessages.
{
"model": "claude-sonnet-4-5",
"tools": [{"type": "function", "function": {"name": "get_weather", "parameters": {"type": "object", "properties": {"city": {"type": "string"}}}}}],
"messages": [
{"role": "user", "content": [
{"type": "text", "text": "What's the weather where this photo was taken?"},
{"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}
]},
{"role": "assistant", "content": null, "tool_calls": [
{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": "{\"city\": \"Ottawa\"}"}}
]},
{"role": "tool", "tool_call_id": "call_1", "content": "12°C and sunny"}
]
}A request that can't be translated returns 400 invalid_request with param set to the field at fault (for example messages[3].tool_call_id). When the provider itself rejects the request (HTTP 400, 404, 413, or 422), the error message relays the provider's reason, prefixed with The provider rejected the request:.
Billing
Each completion charges the caller's credit balance based on token usage (with cache-token semantics per provider) plus a flat 30-credit fee for image-gen calls. Users who configure their own provider API key get a 50% discount. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model slug. Use the `id` from `GET /models` or one of Gumloop's preset routes. | |
| tools | No | OpenAI-shape tool definitions (`{type: "function", function: {name, description, parameters}}`). Pass `tool_choice` to constrain selection. | |
| stream | No | Only non-streaming JSON is supported. Streaming requests are refused before fetch. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| messages | No | ||
| provider | No | OpenRouter provider routing config. Caller fields like `sort` and `order` are honored; ZDR/data_collection policy is server-enforced. | |
| modalities | No | Output modalities. Include `"image"` to route to an image-generation model. | |
| temperature | No | Sampling temperature. | |
| tool_choice | No | ||
| image_config | No | ||
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| response_format | No | ||
| max_completion_tokens | No | Cap on completion tokens. Replaces the deprecated `max_tokens` field. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing billing semantics (per-token credit charge, flat 30-credit image-gen fee, 50% discount with own provider key), error shapes (400 invalid_request with param, provider-rejection prefixing), and a confirmation requirement for account-affecting operations. The main defect is the inaccurate streaming/SSE behavior, which is a schema contradiction rather than an annotation contradiction, so it does not trigger that flag.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a reference doc rather than a tool description: three headers, a code block, a full JSON example, and repeated image/streaming material. Streaming is discussed twice, once in the intro and again under its own heading, and the closing 'Explicit confirmation...' paragraph reads as grafted from another tool. Front-loaded purpose is present, but the length is not earned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A 15-parameter, nested, mutation-capable tool with 0 required params and no output schema warrants substantial detail, and the tool-call loop and error handling are covered. But several parameters are left undocumented and the description asserts a streaming capability the schema forbids, so an agent cannot fully trust it as a complete spec.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 73%, so the schema carries most parameter meaning; the description adds value for tool_choice values, image_url content parts (accepted types, 20 MB limit, no redirects), and the tool/tool_call_id pairing rule. It says nothing about temperature, max_completion_tokens, provider, payload_file, account, or confirm, so it does not fully compensate for the remaining gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line names a specific verb+resource ('chat completions endpoint') and its distinguishing trait — multiplexed across Anthropic, OpenAI, Gemini, OpenRouter — which separates it from siblings like route_model or list_models. It stops short of a 5 because the purpose is muddied by a streaming claim (see below) that misrepresents what the tool actually does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives real usage context: when image-gen models dispatch (modalities includes 'image'), how to continue after finish_reason: tool_calls, and tool_choice options. But the primary 'when to use' guidance — 'Set stream: true for SSE, or omit it for a unary JSON response' — directly conflicts with the schema, which pins stream to const:false and refuses streaming, so the routing advice is actively misleading.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_organization_evaluationCreate evaluationADestructive
Creates an organization evaluation. A new evaluation has no targets, so it cannot start enabled: set targets with PUT /evaluations/{evaluation_id}/targets, then enable it with PATCH /evaluations/{evaluation_id}.
Rubric values are validated strictly: an unknown frequency, criterion priority, type, data point data_type, or session type, a criterion without name and prompt, or a duplicate tag name returns 400 invalid_request with the offending paths in error.details.fields.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare destructive/openWorld/non-idempotent; the description adds substantive behavior beyond them: the create-then-targets-then-enable lifecycle, strict rubric validation that returns 400 invalid_request with offending paths in error.details.fields, and the confirmation/credit-spend caution. This is exactly the extra context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then prerequisites, then validation and safety notes. Three short blocks, each carrying distinct information, though the validation sentence is dense and the endpoint paths overlap with what a schema-aware agent already knows.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent creation tool with no output schema, the description covers lifecycle, prerequisites, failure modes, and confirmation expectations, and implies an evaluation_id is produced. It stops short of describing what the response returns or the semantics of the config rubric versus update paths, but nothing critical to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and there are only two top-level params, so the baseline is 3; the description earns above that by explaining that 'explicit confirmation is required for this exact account operation' (the confirm flag) and that the operation is account-scoped. It adds meaning but not per-field syntax detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Creates an organization evaluation') and immediately situates it in the lifecycle, pointing to the PUT targets and PATCH enable endpoints that correspond to sibling tools (set_organization_evaluation_targets, update_organization_evaluation). An agent can separate creation from update/enable/delete without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives real conditional guidance: a new evaluation has no targets, so it cannot be enabled until targets are set and then patched. It also states explicit confirmation is required and warns against auto-resubmitting unknown outcomes. It does not name sibling tools as alternatives for listing/updating evaluations, so it stops short of full when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_sessionCreate sessionADestructive
Create a new session for an agent. When input is provided, the message is enqueued and the agent begins processing — the response returns 202 with the session in processing or queued state. When input is omitted, an idle session stub is created and the response returns 201.
agent_id also accepts the reserved aliases gumball and analytics, which resolve to your personal Gumball and analytics agents (created on first use).
Streaming the response
api.gumloop.com only serves the non-streaming response above. To stream agent output as it's produced, send the same request body (with stream: true) to the streaming host instead:
POST https://ws.gumloop.com/api/v1/agents/{agent_id}/sessionsThe response is text/event-stream (Server-Sent Events). With the Python SDK, client.sessions.stream(agent_id, input="...") routes to ws.gumloop.com automatically and yields parsed StreamEvent objects.
If you send stream: true to api.gumloop.com by mistake, the response is a 400 whose body contains the correct streaming host so you can retry against it.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| input | No | The first user message for the session. Also accepted as `message` for backwards compatibility. When omitted, an idle session is created with no messages. | |
| stream | No | Must be `false` (or omitted) when calling `api.gumloop.com`. Set to `true` only when calling `ws.gumloop.com` (see the streaming section above). | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| agent_id | Yes | ID of the agent to start a session on. Also accepts the reserved aliases `gumball` and `analytics`. | |
| metadata | No | Arbitrary key/value metadata attached to the session. Stored under `metadata.client`. | |
| session_id | No | Caller-supplied session ID. When omitted, the server generates one. If provided and the ID already exists, the request returns `409 session_already_exists`. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (destructive/openWorld/non-idempotent): it discloses status codes for each mode, session state transitions, the streaming host split, the 400 misrouted-stream error containing the correct host, the 409 session_already_exists case, and a credit-spend/confirmation warning with 'never resubmit unknown outcomes automatically'. This is exactly the behavioral context annotations alone cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and mode behavior, then uses a labeled streaming section. It is somewhat long, and the trailing confirmation paragraph reads as appended boilerplate, but every section serves the complex, multi-host semantics of this tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description covers the return surface (201/202 states, SSE stream, 400 and 409 errors) and the streaming host distinction. For a 10-parameter tool with nested objects, it supplies what an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 90%, so the baseline is 3. The description adds meaning beyond the schema by explaining that agent_id aliases gumball/analytics resolve to personal agents created on first use, and by clarifying that input enqueues a message and starts processing versus creating an idle stub.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a new session for an agent') and clearly distinguishes the two operating modes (input provided → 202 processing/queued; input omitted → 201 idle stub). It does not, however, name the sibling tools it competes with (e.g., send_message, queue_session_message), so sibling differentiation is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete conditions for the two creation modes and for the streaming host, plus the confirm requirement. It stops short of routing the agent among alternatives like send_message or queue_session_message when a session already exists, so it is strong context without explicit exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_skillCreate skillADestructive
Upload a skill package and create a new skill. The package must include a SKILL.md with name and description frontmatter; uploads may be a single .md file (stored as SKILL.md), or a .zip / .skill archive containing SKILL.md at its root. The initial version is created automatically. Maximum upload size is 10 MB.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| files | No | ||
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| team_id | No | Team that should own the skill. When omitted, the skill is owned by the authenticated user. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false and openWorldHint=true; the description goes further by disclosing credit spend, downstream side effects, and the mandatory confirmation gate. That is real value beyond the annotations, though it contains a size figure that conflicts with the schema (see below).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and packaging rules, then the safety/confirmation note. The second paragraph overlaps with the confirm parameter's own description, but overall it earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 6 parameters (none required), the description adequately covers the packaging contract, confirmation requirement, and side-effect risk. The size contradiction and unaddressed body flags (payload/payload_file/account/team_id) are the remaining gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83% (baseline 3) and the description does add useful non-schema constraints: SKILL.md with name/description frontmatter, accepted .md/.zip/.skill forms, archive layout. However, it never explains the confirm, account, team_id, or payload_file parameters, and its stated 'Maximum upload size is 10 MB' contradicts the schema's 'at most 5 MiB' per file/total, which could cause failed uploads.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('Upload a skill package and create a new skill') plus the packaging constraints that define what this operation actually consumes. It is trivially distinguishable from siblings like update_skill, delete_skill, and upload_file.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the confirmation prerequisite and the 'never resubmit unknown outcomes' rule, which is genuinely actionable guidance. It does not explicitly route against alternatives such as update_skill or a plain file upload, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_brain_fileDelete fileBDestructive
Remove a file from the source and from search. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| file_id | Yes | ||
| source_id | Yes | The source id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds genuinely useful context that deletion also removes the file from search, and reinforces the confirmation requirement. The trailing sentence about runs spending credits and downstream actions reads like generic boilerplate that is only loosely tied to a file-delete tool, diluting the value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in a tight first sentence. The second sentence is bloated and partially off-topic (credits, downstream runs), which is not clearly earned for a file-deletion operation and makes the guidance muddier rather than sharper.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no output schema and 75% parameter coverage, the description covers the key effect (removal from source and search) and the confirmation gate, which is the minimum an agent needs. It leaves the file_id parameter and the return/error behavior unexplained, and the confusing run-centric sentence reduces rather than improves completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% and the schema already documents account, confirm, and source_id, so the baseline is 3. The description reinforces the confirm/account semantics ('exact account operation') but adds no syntax or format detail beyond what the schema's own parameter descriptions already provide, and file_id remains undocumented in both.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Remove) and resource (a file) plus the scope of effect: 'from the source and from search.' This lets an agent distinguish it from the sibling delete_brain_source, which removes a source rather than a file, without opening either schema. It does not explicitly name that sibling, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states a real usage precondition ('Explicit confirmation is required for this exact account operation') and warns against automatic resubmission, which implies the user must explicitly request the delete. However, it never states when to use this tool versus delete_brain_source or other brain-file tools, so the routing guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_brain_sourceDelete sourceADestructive
Delete a source, every file in it, and everything it contributed to search. This cannot be undone. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| source_id | Yes | The source id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false and readOnlyHint=false, so the safety profile is covered structurally. The description still adds value by spelling out the irreversible cascade (files plus search contributions) and the explicit-confirmation gate, which annotations cannot express. It stops short of describing downstream consequences like search index rebuild timing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the destructive scope in the first sentence and keeps the whole thing to three sentences. The final sentence about credits and unknown outcomes is noticeably less relevant to a delete operation and dilutes the otherwise tight structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema needed and annotations carrying the safety hints, the description supplies the two things structure cannot: irreversible cascade semantics and the confirmation requirement. An agent has enough to call this correctly, though the credit/retry caveat is noise rather than needed context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so account, confirm, and source_id are already documented in the schema itself. The description's mention of required confirmation loosely reinforces the confirm flag but adds no format or constraint detail beyond what the schema states. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (delete) and resource (source) plus the exact blast radius: every file in it and everything it contributed to search. That cascade scope cleanly separates it from siblings like delete_brain_file and delete_skill without opening their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It establishes the confirmation precondition, which is real usage guidance, but never names an alternative or states when to delete a whole source versus deleting individual files. The trailing 'runs can spend credits' caveat is generic boilerplate that doesn't route the agent between siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_organization_evaluationDelete evaluationADestructive
Deletes the evaluation. It stops running and disappears from lists; results it already produced stay attached to their sessions. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| evaluation_id | Yes | ID of the organization evaluation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, but the description goes well beyond them: it explains exactly what is destroyed vs preserved (results remain attached to sessions), requires explicit confirmation, and warns about credit spend and downstream actions. That is substantive behavioral context an agent cannot get from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and consequence before the confirmation and safety caveats. No filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent tool with no output schema, the description covers consequences, confirmation requirements, and side effects adequately. The final sentence about runs spending credits is slightly tangential to a delete operation and mildly muddies the focus, keeping it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents account, confirm, and evaluation_id. The description reinforces that confirmation is mandatory and tied to 'this exact account operation' but adds no syntax or format detail beyond the schema, matching the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Deletes the evaluation') plus the concrete consequences — it stops running, disappears from lists, and prior results persist on sessions. This distinguishes it from sibling mutators like delete_skill, delete_brain_source, and update_organization_evaluation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditional guidance ('Explicit confirmation is required for this exact account operation') and a warning not to auto-resubmit unknown outcomes. It does not name a sibling alternative (e.g., update_organization_evaluation or a cancel-type tool) for cases where deletion is not intended, so it falls just short of the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_queued_messageDelete queued messageADestructive
Remove a message from the session's queue before it is sent. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| session_id | Yes | ID of the session the queued message belongs to. | |
| queued_message_id | Yes | ID of the queued message to delete. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds the confirmation requirement and the pre-send timing constraint, which is useful, but the sentence 'Runs can spend credits or trigger downstream actions' is boilerplate that does not apply to removing a queued message and reads as misleading noise.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core statement is front-loaded and efficient, but the third sentence about runs spending credits and not resubmitting unknown outcomes does not earn its place for a queue-deletion tool and dilutes the definition. Two of three sentences are on-topic; the third is generic copy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent tool with no output schema, the annotations cover the safety profile and the description covers the confirmation prerequisite, so the essentials are present. It still omits whether deletion is reversible or what the result looks like, and the off-topic credit/downstream sentence leaves the definition feeling partially templated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (account, confirm, session_id, queued_message_id) are already documented in the schema. The description only loosely gestures at confirm and account ('explicit confirmation', 'this exact account operation') and adds no syntax, format, or constraint detail beyond what is structured.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb and resource ('Remove a message from the session's queue') plus a scope qualifier ('before it is sent') that cleanly separates it from siblings like send_queued_message, update_queued_message, and list_queued_messages. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Explicit confirmation is required for this exact account operation' implies a usage condition (only delete when the user asked for exactly this action), which is reinforced by the schema's confirm field. However, no alternatives are named (e.g., update_queued_message to edit instead of delete) and the closing sentence about runs/credits adds no routing guidance for this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_skillDelete skillADestructive
Permanently delete a skill. This is a soft-delete — the skill will no longer appear in listings or be usable by agents. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| skill_id | Yes | ID of the skill to delete. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, so the bar is lower. The description still adds meaningful behavior: the operation is a soft-delete whose effect is removal from listings and loss of agent usability, plus a confirmation precondition. The 'permanently delete' / 'soft-delete' phrasing is internally muddled but not contradicted by the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The delete behavior and soft-delete consequence are front-loaded and useful. The final sentence about runs spending credits and downstream actions is off-topic for a skill deletion and reads as filler, diluting an otherwise tight definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation with full schema coverage and annotations covering the safety profile, the description supplies the missing pieces: soft-delete semantics and the confirmation requirement. No output schema exists, so return values need not be explained. Only the irrelevant credits sentence detracts.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so account, confirm, and skill_id are already documented in the schema. The description reinforces that confirmation is mandatory and account-scoped, but adds no syntax or format detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Permanently delete a skill') and adds scope detail that the skill stops appearing in listings and becomes unusable by agents. It doesn't explicitly distinguish itself from siblings like update_skill or delete_brain_source, but the verb+resource pairing is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states that explicit confirmation is required for 'this exact account operation,' which is real usage guidance. However, the trailing sentence about runs spending credits and downstream actions reads like boilerplate copied from a run-execution tool and does not tell the agent when to prefer delete_skill over update_skill or when deletion is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
detach_agent_mcp_serverDetach an agent MCP serverADestructive
Detach an MCP server (connector) from an agent. This is idempotent — detaching a server that isn't attached returns detached: false rather than an error.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| agent_id | Yes | ID of the agent. Also accepts the reserved aliases `gumball` and `analytics`. | |
| server_id | Yes | ID of the MCP server to detach. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and idempotentHint=false, but the description directly contradicts the idempotent flag by stating the operation IS idempotent and returns detached:false rather than an error. Beyond annotations, it discloses the return shape of the edge case and the confirmation/credit-spend caution.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then adds idempotency and confirmation details. Two sentences plus a caution, all earning their place. Slightly longer than strictly necessary but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation tool with no output schema, the description covers the action, idempotency edge case, confirmation requirement, and a safety caution about credits. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all four parameters. The description adds value by clarifying the confirm parameter's semantics ('explicit confirmation... exact account operation') and implying the idempotent behavior of server_id. Baseline 3 is exceeded because the description reinforces key parameter intent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (detach) and resource (MCP server/connector) with explicit scope ('from an agent'). Clear counterpart to sibling attach_agent_mcp_server and distinguishable from list_agent_mcp_servers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The confirmation requirement ('Set true only when the user asked for exactly this action') gives clear when-to-use guidance, and 'never resubmit unknown outcomes automatically' adds a caution. However, it doesn't explicitly name attach_agent_mcp_server as the inverse operation or spell out when detaching is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_artifact_fileDownload artifactARead-onlyIdempotent
Returns a signed download URL for an artifact, plus its filename, media type, and size. Follow download_url to fetch the file bytes.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| version_id | No | Specific version of the artifact to download. Defaults to the latest version when omitted. | |
| artifact_id | Yes | ID of the artifact to download. | |
| output_file | Yes | New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds real behavioral value beyond that: it discloses the return shape and, more importantly, that the tool does not stream bytes but hands back a signed URL that must be followed, which an agent could not infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences that lead with the return value and then the required follow-up action. The trailing 'Read operation.' is mildly redundant against readOnlyHint=true, but it costs almost nothing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining returns and does so (URL, filename, media type, size), plus the follow-the-URL step. Combined with fully documented parameters and annotations, an agent has enough to call this correctly; only sibling differentiation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (account, version_id, artifact_id, output_file) are already documented, including version defaulting and the private output_file requirement. The description adds no parameter-level detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Returns a signed download URL for an artifact') and enumerates what comes back (filename, media type, size), so the agent knows exactly what the call yields. It does not, however, differentiate itself from the nearby download_file / download_files / download_skill_file siblings, leaving the artifact-specific scope implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The instruction 'Follow `download_url` to fetch the file bytes' implies the intended two-step usage, which is genuinely useful. But there is no explicit guidance on when to pick this over download_file or download_files, and no mention of prerequisites or when-not-to-use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_fileDownload fileCRead-onlyIdempotent
Download file Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No | The ID of the flow run associated with the file. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| user_id | No | Optional. The user ID associated with the flow run. | |
| file_name | No | The name of the file to download. | |
| project_id | No | Optional. The project ID associated with the flow run. | |
| output_file | Yes | New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| saved_item_id | No | The saved item ID associated with the file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered structurally. The description's 'Read operation' merely echoes readOnlyHint and adds no new context such as auth/credential requirements or the meaning of the private output directory behavior. No contradiction, but essentially zero added value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is short but not concise in a useful sense: it is under-specified rather than economical, spending its two lines on a restated name and a redundant 'Read operation' while omitting everything an agent would need.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with a nested payload object, a required output_file, private-credential 'account' semantics, and no output schema, the description is far too thin. It leaves the agent to reconstruct usage entirely from the schema and annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema itself documents all 9 parameters including the required output_file and the nested payload object. The description contributes no parameter meaning, which is acceptable only because the schema does the heavy lifting; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is essentially a tautology: it restates the tool name and title ('Download file') with no verb+resource elaboration, no scope, and no differentiation from the many sibling download tools (download_files, download_artifact_file, download_skill_file). An agent cannot tell what resource space this operates in from the text alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative guidance whatsoever. With siblings like download_files, download_artifact_file, and download_skill_file in the toolset, the absence of any routing signal is a serious omission.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_filesDownload multiple filesCRead-onlyIdempotent
Download multiple files Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No | The ID of the flow run associated with the files. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| user_id | No | The user ID associated with the files. Required if project_id is not provided. | |
| file_names | No | An array of file names to download. | |
| project_id | No | The project ID associated with the files. Required if user_id is not provided. | |
| output_file | Yes | New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| saved_item_id | No | Optional. The saved item ID associated with the files. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint. The description's "Read operation" merely echoes readOnlyHint and adds nothing new — no note on credential scope for the private account, no rate/limit behavior, no mention that results are written out of model context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two fragments with no wasted words, but the problem here is under-specification rather than conciseness. Brevity at this level of complexity leaves the agent without the routing and prerequisite information it needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with a nested payload object, a required output_file with non-overwrite/private-directory semantics, and mutually exclusive payload_file vs payload vs body flags, the description covers none of the call-shaping rules. With no output schema the burden falls entirely on the description, and it does not carry it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all nine parameters (including the nested payload object and the required output_file) are already documented in the schema. The description adds no syntax, format, or mutual-exclusivity detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Download multiple files" states a verb and resource and hints at scope (plural vs the singular sibling download_file), but it is essentially the title restated with a two-word addendum. It gives no indication of what files are being downloaded, from where, or under what identity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance at all. The agent gets no cue for choosing this over download_file, download_artifact_file, download_skill_file, or the upload_* siblings, and no prerequisites (run_id/user_id/project_id either-or) are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_skill_fileDownload skillARead-onlyIdempotent
Generate a signed URL to download a skill's contents as a .skill archive (ZIP). When version_id is provided, returns that exact version; otherwise returns the current draft.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| skill_id | Yes | ID of the skill to download. | |
| version_id | No | Specific version to download. When omitted, the current draft is returned. | |
| output_file | Yes | New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld), so the bar is lower. The description adds genuine context the annotations do not: that the output is a signed URL rather than raw bytes, and that the archive format is ZIP. It does not mention expiry of the signed URL, but that is a minor omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the core action front-loaded, the version branch second, and a one-line read-operation tag. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly explains what is returned (a signed URL) and the archive format. It does not mention the required output_file destination at all, which is the one gap for a tool that writes results to a private file.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter — including version_id's draft default and output_file's private-file semantics — is already documented. The description restates the version_id behavior but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource — generate a signed URL for a skill's contents as a `.skill` (ZIP) archive — and names the sibling resources it is not (skills vs. generic files/artifacts). An agent can distinguish it from download_file and download_artifact_file without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the version_id vs. draft branch, which is useful selection context, but it never names an alternative tool or states when to prefer this over download_file/download_files. Usage is implied by the resource type rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_dataExport dataADestructive
This endpoint allows enterprise organization administrators to create and initiate a comprehensive data export for their organization or specific workspaces.
The export supports six data types:
Workflow data (
data_type: "workflows"): Includes workflow runs, workbook details, user information, and other organizational data.Agent data (
data_type: "agents"): Includes agent configurations, metadata, tools, and creator information.Agent interaction data (
data_type: "agent_interactions"): Includes agent interaction data with timestamps, credit costs, trigger types, and message counts.Credit log data (
data_type: "credit_logs"): Includes credit transaction history with charges, balances, categories, and user attribution.Interaction evaluation data (
data_type: "interaction_evaluations"): Includes one row per completed evaluation of a chat, with its grade, call outcome, sentiment, and the model that graded it.Gumstack data (
data_type: "gumstack"): Includes Gumstack MCP tool call activity with timestamps, statuses, and latency.
The available export_fields depend on the selected data_type. See the field descriptions below for details.
Scoping requirement: For non-credit-log exports, at least one scoping parameter must be provided: workspace_ids, include_all_workspaces, include_personal_workspaces, or entity_ids. Requests that omit all scoping parameters will receive a 400 error.
Note: Credit log exports work differently from workflow and agent exports. When data_type is "credit_logs", the following parameters are not applicable and will be ignored: export_level, workspace_ids, include_all_workspaces, include_personal_workspaces, and entity_ids. Credit log exports are always scoped to the entire organization. Use category_filter to filter by credit log category.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| user_id | No | The ID of the user requesting the export. | |
| end_date | No | End date for the export in ISO 8601 format (e.g., `2025-12-31T23:59:59Z`). | |
| data_type | No | ||
| entity_ids | No | ||
| start_date | No | Start date for the export in ISO 8601 format (e.g., `2025-01-01T00:00:00Z`). | |
| export_level | No | ||
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| export_fields | No | ||
| workspace_ids | No | ||
| category_filter | No | ||
| include_all_workspaces | No | ||
| include_personal_workspaces | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, openWorldHint=true and idempotentHint=false, but the description adds real substance: it can spend credits or trigger downstream actions, explicit confirmation is required, and unknown outcomes must not be resubmitted automatically. It also discloses the 400 failure mode for missing scoping and the divergent credit_logs behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and then organized under headers and bullets for the six data types, scoping requirement, and credit_logs caveat. It is long, but the length is driven by genuinely distinct data-type semantics rather than repetition; a few of the type summaries overlap with the schema text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter, nested-object tool with no output schema, the description covers data-type selection, scoping rules, and the credit_logs exception well. The main gap is that it never explains the async return/polling model (get_export_status exists as a sibling), leaving the agent unsure what the call yields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is low (47%), so the description has to carry weight, and it does: it explains that export_fields depend on data_type, that the scoping params are not applicable to credit_logs, and points to category_filter as the credit_logs alternative. It does not, however, describe the confirm/account/payload wrapper params or date-range semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('create and initiate a comprehensive data export') and scopes it to organization admins and workspaces. It enumerates all six data types with their contents and even distinguishes itself from audit-log exports ('available in the Gumloop app only, not through this endpoint'), so an agent can tell it apart from sibling export/status tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete prerequisites and conditions: the scoping requirement (at least one of workspace_ids/include_all_workspaces/include_personal_workspaces/entity_ids, else 400), and the credit_logs exception where scoping params are ignored. It does not explicitly route the agent to get_export_status for polling, so it stops short of naming alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_brain_sourceRetrieve sourceBRead-onlyIdempotent
Fetch one source the authenticated user can see. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| source_id | Yes | The source id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, and the description's 'Read operation' merely restates that. It does add the visibility-scoping constraint (only sources the user can see), which is mild extra context, but nothing about errors, missing sources, or response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with zero padding. The core verb and scope appear first, and nothing is redundant beyond the terse 'Read operation' tag.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter read tool with full schema coverage and no output schema, the definition is minimally sufficient. It never indicates what a retrieved source contains or what happens when the id is invalid or inaccessible, which the description could reasonably cover.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both source_id and account are documented in the schema itself. The description contributes nothing beyond that, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Fetch) and resource (one source) with singular scope, which implicitly distinguishes it from the sibling list_brain_sources. However, it does not explicitly name that sibling or clarify what a 'source' is in this domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Adds a visibility qualifier ('the authenticated user can see') but gives no when-to-use guidance, no prerequisites, and no routing to alternatives like list_brain_sources or search_brain. The agent must infer selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_brain_source_estimateRetrieve estimateARead-onlyIdempotent
The latest credit estimate for a source created with require_approval. estimate is null until the first upload has produced a run; poll until estimate.status is paused_for_approval, then call Approve source. estimated_credits is rounded up to the nearest 5 and is an estimate, not a quote.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| source_id | Yes | The source id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, it discloses genuinely useful behavior: `estimate` is `null` until the first upload produces a run, `estimated_credits` is rounded up to the nearest 5, and the value is an estimate not a quote. This is rich, non-obvious context an agent needs to interpret results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded: purpose first, then workflow, then edge cases, all in a few tight sentences. The trailing 'Read operation.' is redundant with the readOnlyHint annotation and slightly dilutes the otherwise lean structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, yet the description describes the returned shape (`estimate` nullability, `estimate.status`, `estimated_credits`) and the polling lifecycle, compensating for the missing output schema adequately for this read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (source_id, account) are already documented in the schema. The description adds no parameter-level semantics beyond what structured fields provide, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('latest credit estimate for a source') and scopes it to sources created with `require_approval`, distinguishing it from sibling tools like get_brain_source and approve_brain_source. An agent can identify its role without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit workflow guidance: poll until `estimate.status` is `paused_for_approval`, then call the approve-source tool (named as a sibling workflow step). It tells the agent both when to call and what to do next, which is unusually actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_evaluation_configGet evaluation configBRead-onlyIdempotent
Retrieve the current evaluation configuration for an agent, including criteria, tags, data points, and sentiment settings. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| agent_id | Yes | ID of the agent. Also accepts the reserved aliases `gumball` and `analytics`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so 'Read operation' is largely redundant. The description does add useful disclosure about what the configuration contains (criteria, tags, data points, sentiment settings), but says nothing about permissions, account scoping behavior, or failure modes beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with purpose and contents. The trailing 'Read operation' sentence is the only near-wasteful element since annotations already convey read-only status, but overall the description is tight and well-ordered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing return content, and it does so by enumerating the configuration sections returned. It is adequate for a simple two-parameter read tool, though it omits any note on when the config might be empty or missing for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (account and agent_id) are already fully documented, including the gumball/analytics aliases. The description adds no additional meaning about parameter behavior, which is the expected baseline-3 outcome when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (retrieve) and resource (the current evaluation configuration for an agent), plus enumerates what it covers: criteria, tags, data points, sentiment settings. It is distinguishable from update_evaluation_config by the 'current' read framing, but it never names or contrasts with any sibling explicitly, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives such as get_evaluation_options, list_evaluations, or get_evaluation_metrics, all of which appear in the sibling list. 'Read operation' is a behavioral note, not usage guidance; the agent must infer applicability from the resource name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_evaluation_metricsGet evaluation metricsARead-onlyIdempotent
Returns aggregated grade and tag counts for an agent's evaluations over a time window. Useful for dashboards and reporting on agent quality trends. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Number of days to look back (1-365). Defaults to 30. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| agent_id | Yes | ID of the agent. Also accepts the reserved aliases `gumball` and `analytics`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is fully covered and the trailing 'Read operation' is largely redundant. The description does add behavioral value by disclosing what is returned (aggregated grade and tag counts), which matters since there is no output schema. It says nothing about aggregation granularity or pagination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the return semantics before the use case. The only waste is the trailing 'Read operation', which duplicates the readOnlyHint annotation without adding information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only metrics tool with no output schema, the description adequately conveys what is returned (aggregated grade and tag counts) and the time-window scope. Gaps remain around bucketing granularity and response shape, but these are minor given the richness of the schema and annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (days, account, agent_id) are already documented in the schema, including the day range and the gumball/analytics aliases. The description adds no parameter detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: returns 'aggregated grade and tag counts for an agent's evaluations over a time window'. The scope (one agent, time-bounded, aggregated counts) is clear. It does not explicitly distinguish itself from close siblings like get_organization_evaluation_metrics or list_evaluations, but 'agent's evaluations' implies the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Useful for dashboards and reporting on agent quality trends' gives an implied usage context, which is better than nothing. However, it names no alternatives and states no when-not conditions, so the agent must infer when to prefer this over the organization-level metrics sibling or list_evaluations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_evaluation_optionsGet evaluation optionsARead-onlyIdempotent
Allowed values for evaluation fields and filters — session types, criterion types and priorities, data point types, frequencies, grades, statuses, target types, skip reasons — plus size limits. Use these instead of hardcoding enums. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the read-only/idempotent/openWorld safety profile, so the bar is lower. The description adds substantive value by disclosing the scope of what is returned (the enumerated value categories plus size limits), which the annotations do not convey. "Read operation" is redundant with readOnlyHint but harmless.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, then the usage directive, in three tight sentences with no filler. The category list is long but it is the actual content an agent needs, so it earns its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the returned value categories and size limits, and it notes the read-only nature. Nothing critical is missing for a zero-required-param lookup tool, though it could tie itself more explicitly to the evaluation siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'account' parameter is fully documented in the schema. The description adds no parameter-level guidance, so the baseline of 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource ('Allowed values for evaluation fields and filters') and enumerates the categories it covers (session types, criterion types, frequencies, grades, statuses, etc.), which clearly separates it from config/evaluation siblings. It never names a sibling directly, so differentiation is implied rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use these instead of hardcoding enums" gives a clear usage directive that tells the agent why and when to prefer this tool. There are no explicit when-not conditions or named alternatives (e.g., get_evaluation_config vs. this), but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_export_statusGet data export statusARead-onlyIdempotent
This endpoint retrieves the status of a data export job and optionally downloads the export file (as CSV) if the export has completed successfully.
Use the data_export_id returned by the Export data endpoint to check progress.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| user_id | No | The ID of the user requesting the export status. | |
| download | No | Set to `true` to download the export file directly when the export is completed. When `true` and the export state is `COMPLETED`, the response will be a CSV file download instead of JSON. | |
| output_file | Yes | New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. | |
| data_export_id | Yes | The unique identifier of the data export job to check (returned by the Export data endpoint). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so 'Read operation.' is pure repetition. The one added behavioral fact — the response switches to a CSV download rather than JSON when the export is COMPLETED and download is requested — is also restated verbatim in the schema's `download` description, so the description contributes little beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, laid out efficiently with the core action first and the prerequisite second. The trailing 'Read operation.' is redundant against the annotations and is the only wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must carry return-value semantics. It covers the CSV-vs-JSON response mode but never enumerates the possible job states (e.g. PENDING, COMPLETED, FAILED) or what the JSON payload contains, which is precisely what a status-checking tool's caller needs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds provenance the schema lacks by tying `data_export_id` to the Export data endpoint and clarifying the optional-download toggle in prose. It does not explain `output_file` or `account`, but those are documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('retrieves the status of a data export job') plus the conditional CSV download, and explicitly frames itself as the progress-check counterpart to the Export data endpoint. An agent can distinguish it from the sibling export_data and download_file tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the agent when to reach for this tool: after calling Export data, using the returned `data_export_id` to check progress. There is no explicit when-not guidance or mention of failure/cancelled states, but the sequencing requirement is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_input_schemaRetrieve input schemaCRead-onlyIdempotent
Retrieve input schema Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| user_id | No | User ID that created the flow. Required if project_id is not provided. | |
| project_id | No | Project ID that the flow is under. Required if user_id is not provided. | |
| saved_item_id | Yes | The ID of the saved item for which to retrieve input schemas. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, fully covering the safety profile. The description's only behavioral claim ("Read operation") merely echoes readOnlyHint, adding no new context such as auth/account coupling or scoping behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short and front-loaded, but the second line "Read operation" is pure redundancy with the annotations and earns no place. Brevity here reflects under-specification rather than efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool (with required saved_item_id and an account/user/project identity choice) and no output schema, the description says nothing about what an 'input schema' contains or how the identity parameters interact. It is too thin to guide correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (account, user_id, project_id, saved_item_id) are documented in the schema itself, including the user_id/project_id mutual requirement. The description adds nothing, but the baseline is 3 when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description "Retrieve input schema" only restates the tool name and title without adding scope, format, or distinctions from siblings like list_mcp_server_tools or get_evaluation_config. An agent learns nothing beyond what the name already says.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Read operation" is a safety statement, not usage guidance. There is no indication of when to call this versus alternatives, no prerequisites, and no conditions for use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_mcp_server_promptGet MCP server promptBRead-onlyIdempotent
Render one prompt template with arguments and return its messages. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Prompt name from List MCP server prompts. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| team_id | No | Scope the lookup to a single team. | |
| arguments | No | Argument values for the template. | |
| server_id | Yes | ||
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so 'Read operation' is largely redundant. The one piece of added value is 'return its messages', which discloses the return shape in the absence of an output schema. No auth or rate-limit context is provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the action front-loaded and no padding. The second sentence ('Read operation.') is redundant against the annotations, which is the only real waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with nested objects and no output schema, the description does too little: it never clarifies the top-level name vs payload.name duplication, the payload/payload_file/body-flag mutual exclusion, or the account credential-selection semantics. It only hints at the return value.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 86%, so the schema already documents server_id, name, account, team_id, payload, arguments and payload_file. The description adds no parameter-level detail at all – notably it does not explain the payload vs payload_file vs body-flag exclusivity, which the schema only hints at.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Render one prompt template with arguments and return its messages.' This is distinguishable from the sibling list_mcp_server_prompts, though the description never names that sibling. Clear but without explicit sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by 'Render one prompt template with arguments' – the agent can infer this is the execution counterpart to listing prompts. There is no explicit when-to-use, no prerequisites, and no pointer to watch server_id or account selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_organization_audit_logsRetrieve audit logsBRead-onlyIdempotent
This endpoint retrieves audit logs for all users in an organization for a specified time period. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number for pagination. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| user_id | No | Your user id -- you must be an organization admin to retrieve organization logs. | |
| end_time | Yes | End timestamp for log filtering (ISO format). | |
| user_ids | No | Comma-separated list of user IDs whose events should be returned. | |
| page_size | No | Number of records per page. | |
| entity_ids | No | Comma-separated list of entity IDs (agents, workbooks, files) to filter by. The singular `entity_id` param accepts a single value. | |
| start_time | Yes | Start timestamp for log filtering (ISO format). | |
| event_types | No | Comma-separated list of event types to filter by (e.g. `user_sign_in,credential_retrieval`). The singular `event_type` param accepts a single value. | |
| ip_addresses | No | Comma-separated list of source IP addresses to filter by. The singular `ip_address` param accepts a single value. | |
| workspace_ids | No | Comma-separated list of workspace (team) IDs to filter by. The singular `workspace_id` param accepts a single value. | |
| organization_id | Yes | The ID of the organization to retrieve audit logs for. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered. The description's 'Read operation' merely restates the annotation and the org-wide/time-bounded scope is its only added value, which is modest against an already-rich annotation set.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the core purpose front-loaded and zero padding. The trailing 'Read operation' is somewhat redundant given readOnlyHint, slightly diminishing an otherwise tight statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with full schema coverage and annotations covering safety and idempotency, the essentials are present, but the definition is thin: it omits the admin precondition, pagination expectations, and filtering capabilities that live only in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 12 parameters are documented (pagination, filters, admin requirement) in the schema itself. The description adds no parameter-level meaning, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (retrieves) and resource (audit logs) with scope (all users in an organization, specified time period). The name and description together make the operation unmistakable, though it doesn't distinguish itself from any sibling — which is largely unnecessary given the unrelated sibling set.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative guidance is provided. The phrase 'for all users in an organization' hints that this is the org-wide variant, but it doesn't tell the agent when this is preferable to a user-scoped query or that admin privileges are required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_organization_evaluationRetrieve evaluationBRead-onlyIdempotent
Returns one evaluation with its rubric, targets, current coverage, and result rollup. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| evaluation_id | Yes | ID of the organization evaluation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. "Read operation" merely echoes those annotations and earns no credit. The content inventory (rubric, targets, coverage, rollup) is genuinely useful, but it describes output shape rather than behavior—no auth requirements, scoping, or rate/limit context is added.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, with the substantive content (what is returned) front-loaded. "Read operation." is redundant with the annotations and could be dropped, but overall the definition is tight with no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the right thing by naming the returned components, which is the strongest part of this definition. What is missing is disambiguation from the near-identically named retrieve_evaluation and the other evaluation siblings, a real risk in a toolset this crowded.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both account and evaluation_id are fully documented in the schema itself. The description adds nothing about parameter format, defaults, or how account selection affects visibility. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Returns one evaluation") and enumerates what comes back: rubric, targets, current coverage, result rollup. That is far more than a restatement of the name. However, it never distinguishes itself from the sibling "retrieve_evaluation" (which likely carries the same title) or from "get_evaluation_config"/"get_organization_evaluation_result", so an agent cannot disambiguate from the text alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only guidance offered is "Read operation," which is not a when-to-use statement. There is no mention of when to pick this over retrieve_evaluation, list_organization_evaluations, or get_evaluation_config, and no prerequisites stated. Usage must be inferred entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_organization_evaluation_metricsGet evaluation metricsBRead-onlyIdempotent
Grade counts for one evaluation over a trailing window (default 30 days, 1–365). Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Window length in days. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| evaluation_id | Yes | ID of the organization evaluation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds the window semantics and that the result is aggregate grade counts, but 'Read operation' merely repeats the annotations and no auth, scoping, or aggregation-freshness context is provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the return value and window. The trailing 'Read operation.' is largely dead weight given the readOnlyHint annotation, but the overall size is well controlled.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should carry the burden of describing the return payload; 'grade counts' hints at it but does not say what grades or runs are included. Annotations cover the safety profile, so the remaining gaps are modest but real for a metrics-read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three parameters are documented in the schema. The description's restatement of the default (30 days) and bounds (1–365) duplicates the schema, and it adds nothing about the 'account' parameter or the evaluation_id requirement. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (grade counts for one evaluation) and the temporal scope (trailing window, default 30 days, 1–365). It is clear what the tool returns, but it never distinguishes itself from the near-identical sibling get_evaluation_metrics, leaving the agent to infer the organization-level distinction from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no alternative named. Given siblings such as get_evaluation_metrics and list_organization_evaluation_results, the description should tell the agent which one to pick, but it does not.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_organization_evaluation_resultRetrieve evaluation resultARead-onlyIdempotent
One result, including per-criterion outcomes, extracted data points, and applied tags. Poll this after POST /evaluations/{evaluation_id}/run until status is completed or failed.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| result_id | Yes | Result ID from a run response or a results list. | |
| evaluation_id | Yes | ID of the organization evaluation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/non-destructive, and the description adds genuine behavioral context beyond them: the polling-until-terminal-status pattern. The closing 'Read operation.' merely restates readOnlyHint, so it earns no extra credit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loading what the result contains before the polling instruction. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does partially describe the return payload (per-criterion outcomes, data points, tags), and it explains the polling lifecycle. It could say more about the status field values or error shape, but it is largely sufficient for a simple read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so evaluation_id, result_id, and account are already fully documented in the schema. The description adds no syntax, format, or sourcing detail beyond that, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (retrieve) and resource (a single organization evaluation result), and enumerates its contents (per-criterion outcomes, extracted data points, applied tags). The singular 'One result' clearly distinguishes it from the sibling list_organization_evaluation_results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit operational condition: poll this after POST /evaluations/{evaluation_id}/run until status is completed or failed. That is strong when-to-use guidance, though it does not explicitly name the sibling list tool as the alternative for enumerating results.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_role_credit_limitGet custom role credit limitARead-onlyIdempotent
This endpoint returns the monthly credit limit of one custom role. A monthly_credit_limit of null means the role sets no limit of its own.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| role_id | Yes | The ID of the custom role (the same ID used as `group_id` by the Manage custom role users endpoint). | |
| user_id | No | Your user id -- you must be an organization admin to manage custom role credit limits. | |
| organization_id | Yes | The ID of the organization the custom role belongs to. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safety profile is covered. The description adds one genuinely useful behavioral detail — that a `monthly_credit_limit` of `null` means no role-level limit — but nothing about auth beyond what the schema already states.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core purpose. The trailing 'Read operation.' is redundant given readOnlyHint=true, a minor waste, but overall the structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the work of explaining the key return semantic (null = no limit). Combined with fully documented parameters and a read-only profile, an agent has enough to call it correctly; only the sibling routing is under-explained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so role_id, organization_id, user_id, and account are all documented in the schema itself. The description adds no syntax or format detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'returns the monthly credit limit of one custom role.' The singular 'one' implicitly distinguishes it from the sibling list_role_credit_limits, but it never names that sibling or set_role_credit_limit, so an agent must infer the routing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the singular scoping ('one custom role'). There is no explicit when-to-use, no exclusion, and no reference to the natural alternatives (list_role_credit_limits for enumeration, set_role_credit_limit for mutation), leaving the agent to infer the boundaries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_run_detailsRetrieve run detailsARead-onlyIdempotent
This endpoint can be used to poll for completion and retrieve final flow outputs. Output steps must be used to retrieve outputs. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | ID of the flow run to retrieve | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| user_id | No | The id for the user initiating the flow. Required if project_id is not provided. | |
| project_id | No | The id of the project within which the flow is executed. Required if user_id is not provided. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds non-obvious behavior beyond the annotations: it is a polling endpoint and outputs are only retrievable via output steps. That is genuinely useful context that the structured fields do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very short and front-loaded with the primary purpose. 'Output steps must be used to retrieve outputs' is slightly redundant in wording but carries real information, so nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining returns, and it only vaguely says 'retrieve final flow outputs' without describing run status values or the shape of the payload an agent polling for completion would need to inspect. The output-steps note helps but leaves the return contract underspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents run_id, account, user_id, and project_id, including the conditional requirements. The description adds no parameter-level detail, which is acceptable given the coverage; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: poll a run for completion and retrieve its final flow outputs. It's clearly a read retrieval tool. However, it does not distinguish itself from the sibling get_run_history, which an agent could easily confuse with it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'can be used to poll for completion' implies a usage context (awaiting run completion), which is helpful. But there is no explicit when-to-use vs when-not guidance and no mention of the alternative get_run_history, leaving the agent to infer the choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_run_historyRetrieve automation run historyBRead-onlyIdempotent
This endpoint retrieves the run history for automations, either by workbook or saved item. Returns the 10 most recent runs. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| user_id | No | The user ID. Required if project_id is not provided. | |
| project_id | No | The project ID. Required if user_id is not provided. | |
| workbook_id | No | The ID of the workbook to retrieve run history for. Required if saved_item_id is not provided. | |
| saved_item_id | No | The ID of the saved item to retrieve run history for. Required if workbook_id is not provided. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the 'Read operation.' sentence is redundant with structured data. However, the description does add a genuinely useful behavioral detail not in the annotations: it returns only the 10 most recent runs, which is a meaningful limit an agent must know. No pagination or auth details are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with purpose, followed by the key return limit. The final 'Read operation.' sentence is redundant given the annotations and slightly dilutes efficiency, but the description is otherwise tight with no waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with full annotation coverage and a 100%-documented schema, the description is adequate. It covers purpose, the two lookup modes, and a return limit. It doesn't address pagination, whether the 10-run cap is configurable, or error cases, but these are minor given the structured data already present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents all 5 parameters including the conditional requirements (user_id or project_id; workbook_id or saved_item_id). The description adds no parameter semantics beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb+resource: 'retrieves the run history for automations, either by workbook or saved item.' This distinguishes it from sibling get_run_details (which presumably returns a single run) and from list_flows/list_workbooks. It doesn't explicitly name those siblings, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives like get_run_details or list_flows. The description mentions two lookup modes (workbook or saved item) but doesn't explain when to prefer one or what prerequisites apply. No when-not-to-use context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_browser_profile_cookiesImport cookies into a browser profileADestructive
Add sign-in cookies to a browser profile. This is what gumloop browser import-logins calls. Send the cookies in Chrome extension (chrome.cookies.Cookie) or Chrome DevTools Protocol Cookie shape. With url, only that site's cookies are kept and the import replaces that site; without it, every site in the payload is imported. Cookies are encrypted with the profile's key before storage and are never returned by any endpoint.
Use default as the profile_id to import into the owner's default profile, creating it if needed.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Import only the cookies for this site and replace what the profile had for it. Omit to import every site in `cookies`. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| cookies | No | Cookies in `chrome.cookies.Cookie` or CDP `Cookie` shape. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| team_id | No | Import into a team-owned profile instead of a personal one. | |
| profile_id | Yes | A profile id, or `default`. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (destructiveHint=true, openWorldHint=true), the description discloses that cookies are encrypted with the profile's key before storage, are never returned by any endpoint, that confirmation is required for this exact account operation, and that runs can spend credits or trigger downstream actions. This is exactly the extra behavioral context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and the key url/no-url behavior are front-loaded, and each paragraph covers a distinct concern (import semantics, then confirmation/credit risk). It is slightly dense and the url explanation duplicates the schema text, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool (8 params, nested payload object, no output schema) it covers the critical ground: mutating/destructive nature, encryption and non-return of secrets, confirmation requirement, and credit/spend risk. The various body-flag vs payload vs payload_file alternatives are left to the schema, but that is reasonable given the schema already documents them fully.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description restates the `url` semantics and the `default` profile_id hint that already appear in the schema, adding only marginal meaning (e.g., default profile is created if needed). It does not add syntax or format detail beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Add sign-in cookies to a browser profile') and even identifies the underlying CLI command it maps to, leaving no ambiguity about what the tool does. No sibling tool operates on browser-profile cookies, so the agent can immediately tell this apart from list_browser_profiles or the many unrelated siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear conditional behavior: with `url` only that site's cookies are kept and replaced, without it every site in the payload is imported, plus the `default` profile_id convention. It also flags that explicit confirmation is required and warns against auto-resubmitting unknown outcomes. It stops short of naming an alternative or a 'when not to use' boundary, but the operative conditions are spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kill_flowKill flow runADestructive
This endpoint is used to kill a flow run and all its subflow runs. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No | The ID of the pipeline run to kill. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| user_id | No | The user ID. Required if project_id is not provided. | |
| project_id | No | The project ID. Required if user_id is not provided. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false and openWorldHint=true, so the safety profile is covered. The description adds genuinely new context beyond that: subflow runs are also terminated, credits/downstream actions are at stake, and outcomes must not be auto-resubmitted. It stops short of saying whether the kill is reversible or what happens to an already-finished run.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the purpose and then the confirmation prerequisite and retry caution. Efficient, though the second sentence is slightly awkwardly worded ('this exact account operation').
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter destructive tool with no output schema, the description covers the destructive scope, the confirmation gate and the non-idempotent retry hazard, which are the key things an agent needs. Remaining gaps (post-kill state, behavior on an already-terminated run) are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (run_id, account, confirm, payload, payload_file, user_id, project_id) is already documented in the schema. The description only obliquely gestures at the account/confirm semantics via 'this exact account operation' and adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (kill) and resource (a flow run) and adds meaningful scope — 'and all its subflow runs' — which distinguishes it from benign siblings like get_run_details or cancel_session. An agent can identify this as the destructive termination tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states a prerequisite ('Explicit confirmation is required for this exact account operation') and a retry caution ('never resubmit unknown outcomes automatically'), which tells the agent how to invoke it safely. It does not, however, name an alternative tool or an explicit when-not-to-use condition relative to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_accountsList configured accountsARead-onlyIdempotent
List private account labels, default selection and configured token method. No credentials, token paths or account content; no network request.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=false, and destructiveHint=false. The description adds meaningful behavioral context beyond those annotations: it confirms no credentials, token paths, or account content are returned, and that no network request is made.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences front-load the core scope and then immediately clarify the negative behavior. Every phrase contributes useful information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only listing tool with no output schema, the description adequately covers the returned concepts and explicitly rules out sensitive data and network activity. It stops short of describing result ordering, formatting, or pagination, but those may not apply or may be self-evident.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no parameter semantics to clarify. Per the rubric, a zero-parameter tool has a baseline of 4, and the description does not need to compensate for undocumented inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource (list accounts) and enumerates the exact scope: private account labels, default selection, and configured token method. Its exclusions also implicitly distinguish it from broader siblings like get_account or get_current_token, which would expose account content or credentials.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by stating exactly what metadata the tool surfaces, but it never explicitly says when to call this instead of alternatives such as get_account or get_current_token. The 'No credentials...' clause scopes the tool, yet no named alternative or when-not condition is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_agent_mcp_serversList agent MCP serversARead-onlyIdempotent
List the MCP servers (connectors) attached to an agent. Sensitive fields such as secret_id and mcp_server_url are scrubbed from the response.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| agent_id | Yes | ID of the agent. Also accepts the reserved aliases `gumball` and `analytics`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered structurally. The description adds genuinely new behavior beyond that: sensitive fields like `secret_id` and `mcp_server_url` are scrubbed from the response, which tells the agent what data it will and will not receive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the main action front-loaded. The trailing 'Read operation.' marginally restates the readOnlyHint annotation rather than adding information, but the definition is otherwise tight and well structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, two-parameter listing tool with rich annotations and no output schema, the description covers the essential return-value caveat (scrubbed sensitive fields). Nothing critical is missing, though it doesn't note whether the list is paginated or empty-case behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both `account` and `agent_id` (including the reserved aliases) are already documented in the schema. The description only implies the agent scoping and adds no format or alias detail beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and a precisely scoped resource ('the MCP servers (connectors) attached to an agent'), which cleanly separates it from list_mcp_servers, attach_agent_mcp_server, and detach_agent_mcp_server. It stops short of naming those siblings explicitly, so it is clear but not maximally differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'attached to an agent' scope implies when to use this versus the general server-list tools, but there is no explicit when-to-use statement, no prerequisite guidance, and no named alternative. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_agentsList agentsARead-onlyIdempotent
List agents the caller has access to. Filter by team, search by name, or narrow to agents that use a specific tool or trigger. Results can be sorted with sort_order and paginated by sending page_size and/or cursor.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| flow | No | Filter to agents that use the specified saved flow as a tool. | |
| tool | No | Filter to agents that use the specified MCP server as a tool. | |
| cursor | No | Opaque cursor from a previous response's `next_cursor`. Pass it to fetch the next page. | |
| search | No | Case-insensitive substring match against the agent name. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| creator | No | Filter to agents created by this user ID. | |
| team_id | No | Scope the listing to a single team. When omitted, returns agents owned by the authenticated user. | |
| page_size | No | Number of agents per page. Sending `page_size` or `cursor` opts into cursor pagination; requests that send neither return the full list. | |
| sort_order | No | Sort order for the listing. Defaults to newest first. | newest |
| has_triggers | No | When `true`, only returns agents that have at least one active trigger configured. | |
| include_last_used | No | When `true`, populates `last_used_at` on each agent with the timestamp of its most recent session. | |
| include_last_updated | No | When `true`, populates `last_updated_at` on each agent with the timestamp of its most recent configuration change. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered; the closing 'Read operation' merely repeats that. The description does add the pagination behavior (sending page_size or cursor opts into paging, sending neither returns the full list), which is useful operational context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core purpose, then filtering, then sorting/pagination. Efficient and free of padding, though the trailing 'Read operation' is redundant against the annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-param, zero-required list tool with no output schema, the description covers purpose, filtering, sorting and pagination adequately, and the schema carries full parameter detail. The main omission is any hint about the shape of the returned agent list, but that is a minor gap given the otherwise complete coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 12 parameters are already documented with equivalent or greater detail in the schema. The description's summary of filters, sort_order, page_size and cursor adds no syntax, format, or constraint information beyond what the schema provides. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List agents') plus the access scope ('the caller has access to'), and enumerates the filterable dimensions. It is clearly distinguishable from singular siblings like retrieve_agent, though it does not name the alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description lists what can be filtered/sorted/paginated, which implies usage, but gives no explicit when-to-use guidance or exclusions against siblings such as retrieve_agent, list_agent_versions, or list_agent_mcp_servers. The agent must infer the routing itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_agent_versionsList agent versionsARead-onlyIdempotent
List the immutable versions of an agent, newest first. Each entry is a point-in-time snapshot of the agent's configuration.
Requires configuration access on the agent — callers limited to using the agent (no configuration access) get a 403.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | Opaque pagination cursor returned by a prior call as `next_cursor`. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| agent_id | Yes | ID of the agent whose versions to list. Also accepts the reserved aliases `gumball` and `analytics`. | |
| page_size | No | Number of versions to return per page. Clamped to 1–100. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/idempotent/non-destructive, so the description's 'Read operation' is redundant, but it adds real value by disclosing the 403 authorization gate and the newest-first ordering. It does not describe return shape or error behavior beyond the auth failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what the tool returns and how it is ordered, followed by the access precondition. No filler sentences; every clause carries information an agent needs before calling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by explaining that entries are configuration snapshots and that pagination is cursor-based (via the schema). It omits the concrete fields inside each version entry, which is a modest remaining gap for a list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and every parameter (cursor, account, agent_id, page_size) is documented in the schema itself. The description adds nothing about parameter semantics, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List the immutable versions of an agent'), plus ordering ('newest first') and what each entry represents ('point-in-time snapshot of the agent's configuration'). It is distinguishable from the singular retrieve_agent_version by implication (list-all vs fetch-one), but does not name any sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete precondition: configuration access is required and callers limited to using the agent receive a 403. That is a clear when-you-can/when-you-cannot signal. It stops short of pointing to alternatives such as retrieve_agent_version or list_agents for a single version or the agent inventory.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_artifactsList artifactsARead-onlyIdempotent
List artifacts (files) produced by an agent. Optionally scope to a specific session, search by filename, sort, and paginate. Deleted files are excluded from the results. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | Opaque pagination cursor returned by a prior call as `next_cursor`. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| agent_id | Yes | ID of the agent whose artifacts to list. Also accepts the reserved aliases `gumball` and `analytics`. | |
| page_size | No | Number of artifacts to return per page. Clamped to 1–100. | |
| session_id | No | Filter to artifacts produced within a specific session. | |
| sort_order | No | Sort order for results. Defaults to `newest`. | newest |
| search_query | No | Case-insensitive substring match against the artifact filename. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the description's trailing 'Read operation.' adds nothing. The genuinely useful addition is that deleted files are excluded from results, a filtering behavior not derivable from annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Efficient two-sentence body with the core action front-loaded and modifiers listed compactly. The standalone 'Read operation.' line is redundant given the readOnlyHint annotation, a small amount of waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a paginated read-only list tool with a fully described schema and rich annotations, the description covers the necessary ground; the deleted-file exclusion and pagination mention round it out. No output schema exists, but return shape for a list tool is largely implied by the cursor/page_size parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all seven parameters (including the alias support for agent_id and the opaque cursor semantics) are already documented. The description merely restates the same capabilities at a higher level and adds no syntax, format, or default details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List artifacts (files) produced by an agent'), and clarifies the ambiguous term 'artifacts' as files. It does not, however, differentiate itself from near siblings like list_brain_files or download_artifact_file, leaving the agent to infer scope from schema and annotations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description enumerates capabilities (scope by session, search, sort, paginate) which implies usage, but never states when to use this tool versus alternatives or any preconditions. There is no explicit when/when-not guidance, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_brain_filesList filesARead-onlyIdempotent
List the files in a file-upload source with their indexing status. Each file carries the sha256 of its bytes, so a client can compare a local folder against the source and upload only what changed.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | ||
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| page_size | No | ||
| source_id | Yes | The source id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the 'Read operation' sentence is redundant. What does add value is the disclosure of what each returned file carries (sha256 of bytes, indexing status), which helps a caller reason about the response without an output schema. Pagination behavior is still undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences plus a tag; purpose is front-loaded in the first clause. The trailing 'Read operation' sentence is wasted tokens against annotations that already say the same thing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully describes the returned items (files, indexing status, sha256) and annotations cover safety. The main gap is pagination semantics for cursor/page_size, which an agent needs to iterate a source fully.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description says nothing about any of the four parameters, and schema coverage is only 50% (cursor and page_size have empty descriptions). With half the parameters undocumented in structured data and no compensating explanation, the description leaves parameter meaning to the name alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and a precise resource ('the files in a file-upload source'), plus what the listing returns (indexing status, sha256). This is clearly distinguishable from siblings like list_brain_sources (sources, not files) and download_files (fetching, not listing).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description supplies a concrete use context — comparing a local folder against the source and uploading only changed files. That tells the agent when this tool is valuable, though it names no alternative (e.g., list_brain_sources) and gives no explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_brain_sourcesList sourcesARead-onlyIdempotent
List the Company Brain sources the authenticated user can see: personal sources, plus team and organization sources shared with them. Every source type is listed, including ones connected in the app such as Notion or Google Drive. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | Only sources in this scope. | |
| cursor | No | Opaque cursor from a previous response's `next_cursor`. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| team_id | No | Only team sources belonging to this team. | |
| page_size | No | Maximum number of sources to return. | |
| source_type | No | Only sources of this type, for example `direct_file_uploads` or `notion`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint/idempotentHint/destructiveHint, and 'Read operation.' merely restates that. The one genuinely additive detail is that every source type is returned, including in-app connections like Notion or Google Drive, but pagination behavior and auth sensitivities are left to the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with the primary purpose front-loaded and no filler. The trailing 'Read operation.' is slightly redundant given readOnlyHint, but the overall structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema and fully documented optional params, the description covers what is enumerated and who can see it. It omits any note about the returned source shape or pagination semantics, but cursor/page_size documentation in the schema covers the mechanics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The prose lightly echoes the scope enum ('personal sources, plus team and organization sources') and the source_type filter, but adds no syntax, format, or interaction detail beyond what the schema already documents for the 6 optional parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List the Company Brain sources') and clarifies the visibility scope (personal plus team/organization shared). An agent can tell this is the enumeration tool rather than a search or single-fetch tool, though it never names the siblings it is distinct from (list_brain_files, get_brain_source, search_brain).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the enumeration framing plus the scope clause describing what is visible to the authenticated user. There is no explicit when-to-use, when-not-to-use, or named alternative, so an agent must infer that search_brain and get_brain_source are the filtered/single-item counterparts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_browser_profilesList browser profilesARead-onlyIdempotent
List the browser profiles you own, or a team's profiles with team_id. A browser profile holds the sign-ins an agent's Browser ability uses. Cookie values are never returned; each profile lists its sites and cookie counts.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| team_id | No | List a team's profiles instead of your personal ones. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, open-world, so the safety profile is covered. The description still adds genuinely useful behavioral detail: cookie values are never returned and each profile reports its sites and cookie counts, which tells the agent what it can and cannot learn from this call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and scope, then the domain concept, then the return-value caveat. 'Read operation' is mildly redundant with readOnlyHint but costs almost nothing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully summarizes the return shape (profiles with sites and cookie counts, no cookie values), which is exactly what an agent needs. The only gap is that the `account` parameter's role in identity selection goes unmentioned in the description, though the schema covers it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema; per the rubric that sets a baseline of 3. The description reinforces `team_id`'s effect ('instead of your personal ones') but adds no syntax, format, or defaulting detail beyond the schema, and never mentions the `account` parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('List the browser profiles') with scope disambiguation built in: personal profiles by default, or a team's via `team_id`. The second sentence explains what a browser profile actually is, which helps an agent distinguish this from cookie-related siblings like import_browser_profile_cookies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context for when to use it: to enumerate profiles you own, or to enumerate a team's profiles when passing `team_id`. No explicit exclusions or named alternative tools (e.g., it doesn't say to prefer this over list_teams for team lookup), so it stops short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_evaluationsList evaluationsARead-onlyIdempotent
Returns a cursor-paginated list of evaluation results for a specific agent, newest first.
Only completed and failed evaluations are returned unless status selects another state.
Each evaluation includes the grade, criteria pass/fail results, extracted data points, applied tags, and sentiment analysis. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| grade | No | Filter evaluations by grade. | |
| cursor | No | Pagination cursor from a previous response's `next_cursor` field. | |
| status | No | Return evaluations in one lifecycle state instead of the default completed and failed set. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| agent_id | Yes | ID of the agent whose evaluations to list. Also accepts the reserved aliases `gumball` and `analytics`. | |
| page_size | No | Number of evaluations to return per page (1-100). | |
| session_id | No | Only evaluations of this session. | |
| created_after | No | Only evaluations created at or after this ISO 8601 timestamp. Timestamps without an offset are read as UTC. | |
| created_before | No | Only evaluations created before this ISO 8601 timestamp. Timestamps without an offset are read as UTC. | |
| organization_evaluation_id | No | Return the results one organization evaluation produced for this agent instead of the agent's own evaluation results. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/destructiveHint, so the safety profile is covered. The description adds value beyond that: pagination ordering (newest first), the default status filter, and the full shape of a returned evaluation (grade, criteria, data points, tags, sentiment). The trailing 'Read operation' merely repeats readOnlyHint, which is minor redundancy rather than a contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the primary purpose and pagination/ordering facts, then the default filter, then returned fields. Efficient overall, but the final 'Read operation.' sentence restates the readOnlyHint annotation and does not earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates what each evaluation contains and documents ordering and the default status filter. It omits only minor things an agent might want (e.g., the account param's credential scoping), but is essentially complete for a filtered list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameter baseline is 3. The description adds meaning beyond the schema by clarifying the default status set that `status` overrides and by implying cursor-based pagination ties to `next_cursor`. It does not enrich the other parameters (account, grade, date bounds), keeping it short of 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Returns a cursor-paginated list), resource (evaluation results), and scope (for a specific agent, newest first). The agent-scoping distinguishes it from list_organization_evaluations and list_organization_evaluation_results without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the default lifecycle set (completed and failed) and how `status` overrides it, which is real usage context. It stops short of naming alternatives (e.g., list_organization_evaluations, retrieve_evaluation) or when-not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_flowsList saved flowsCRead-onlyIdempotent
List saved flows Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| user_id | No | The user ID to for which to list items. Required if project_id is not provided. | |
| project_id | No | The project ID for which to list items. Required if user_id is not provided. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered. The description's only behavioral statement, "Read operation," is redundant with those annotations and adds no new context such as scoping, return shape, or pagination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is short and front-loaded, but the second sentence ("Read operation") is pure redundancy against the readOnlyHint annotation and does not earn its place. Brevity here reflects under-specification rather than disciplined conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering safety and a fully documented schema, the minimum viable information is present, and no output schema means return values need not be explained. However, a list tool with three scoping parameters and many sibling list_* tools would benefit from at least stating differentiation and the user_id/project_id selection rule.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with all three parameters (account, user_id, project_id) documented in the schema, including the either/or requirement between user_id and project_id. The description adds nothing beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource ("List saved flows") that an agent can map to a concrete read operation. It does not, however, distinguish itself from related siblings such as list_workbooks, list_models, or start_flow/kill_flow, so the scope of "flows" versus other listable entities is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, nor any mention of prerequisites such as needing either user_id or project_id. "Read operation" is not usage guidance; it merely restates the read-only nature.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_mcp_server_promptsList MCP server promptsBRead-onlyIdempotent
Return the prompt templates an MCP server exposes, fetched live from the server. When the server is not connected, prompts is empty and gumloop_auth_url is returned.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| team_id | No | ||
| server_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and open-world, so the safety profile is covered. The description adds genuinely new behavioral context: results are fetched live, and the failure mode (server not connected) yields an empty prompts array plus a gumloop_auth_url, which tells the agent how to interpret and act on an empty result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and the resource noun. Each sentence carries distinct information (what it returns, the failure case, the safety classification), with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description does name the return fields (prompts, gumloop_auth_url), which is helpful. However, for a 3-parameter tool with a fully undocumented required parameter, the description remains incomplete on the input side.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%; server_id (the required parameter) and team_id have empty descriptions. The description compensates for none of this gap, providing no explanation of what server_id must reference or when to supply account/team_id, so an agent must guess at the required input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Return) and resource (prompt templates an MCP server exposes), which clearly distinguishes it from siblings like list_mcp_server_tools and list_mcp_server_resources that return different resource types. It does not explicitly name a sibling to route against, but the resource noun is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use versus when-not-to-use guidance and no mention of the closest alternatives (get_mcp_server_prompt for a single prompt, or list_mcp_server_tools/resources for other entity types). The usage is only implied by the tool name and resource noun.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_mcp_server_resourcesList MCP server resourcesARead-onlyIdempotent
Return the resources an MCP server exposes, fetched live from the server. When the server is not connected, resources is empty and gumloop_auth_url is returned so the caller can prompt the user to authenticate.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | Opaque cursor from a previous response's `next_cursor`. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| team_id | No | Scope the lookup to a single team. | |
| server_id | Yes | Identifier of the MCP server. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered elsewhere. The description goes beyond that by disclosing the empty-resources-on-disconnect behavior and the gumloop_auth_url return, which is valuable operational context not in the annotations. It doesn't describe pagination behavior despite the cursor parameter, keeping it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: what it returns, the disconnected-state behavior, and the operation type. No filler and the core purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with full schema coverage and annotations, the description covers the key edge case (disconnected server) and auth flow. It omits any mention of cursor-based pagination despite accepting a cursor, which is a small but real gap given no output schema explains next_cursor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all four parameters including cursor, account, team_id, and server_id are documented in the schema itself. The description mentions the server concept but adds no syntax, format, or scoping detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (return) and resource (the resources an MCP server exposes), with the live-fetch qualifier that distinguishes it from static listings. An agent can tell this apart from siblings like list_mcp_server_tools, list_mcp_server_prompts, and read_mcp_server_resource without consulting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by stating it returns the resource list and explaining the auth flow, but never names an alternative or says when to prefer read_mcp_server_resource or list_mcp_server_tools instead. Usage context is inferable, not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_mcp_serversList MCP serversBRead-onlyIdempotent
Return the catalog of MCP servers visible to the caller — Gumloop-hosted (gumcp_server), user-deployed Gumstack (gumstack_server), and custom (mcp_server) — along with each server's connection state.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| team_id | No | Scope the catalog to a single team. When omitted, returns servers visible to the authenticated user. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds that the catalog is scoped to what is visible to the caller and includes connection state, which is useful context. However, it does not add richer behavioral details such as permissions, pagination, or rate-limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and remains short. The trailing phrase 'Read operation.' is somewhat redundant because annotations already indicate a read-only operation, but the overall text is appropriately sized and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with no output schema, the description adequately explains what is returned: a catalog of MCP servers with connection state. It omits details about the exact return structure or pagination, but the annotations and fully documented parameters carry most of the remaining context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and both parameters ('account' and 'team_id') are documented in the schema, including the default visibility behavior when team_id is omitted. The description adds no parameter syntax or format details beyond what the schema already provides, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Return the catalog of MCP servers.' It also names the three server categories and says the result includes connection state. It does not explicitly distinguish itself from sibling tools like list_agent_mcp_servers or retrieve_mcp_server, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives such as list_agent_mcp_servers, retrieve_mcp_server, or list_mcp_server_tools. The phrase 'visible to the caller' implies a default scope, but no when/when-not conditions or alternative routing are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_mcp_server_toolsList MCP server toolsARead-onlyIdempotent
Return the tools exposed by an MCP server. When the server is not in connected state, tools is empty and gumloop_auth_url is returned so the caller can prompt the user to authenticate.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| team_id | No | Scope the lookup to a single team. | |
| server_id | Yes | Identifier of the MCP server. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the bar is lower, and the description adds genuinely useful behavior: when the server is not connected, `tools` is empty and `gumloop_auth_url` is returned for an auth prompt. That downstream consequence is not derivable from annotations. The trailing 'Read operation' merely restates readOnlyHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the core purpose front-loaded, then the edge-case behavior. No filler, though the 'Read operation' tail is redundant with annotations and slightly dilutes the efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries responsibility for return semantics and does so by explaining the empty-tools/auth-url case. It is complete enough for a listing tool, missing only explicit routing guidance to sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so account, team_id, and server_id are fully described in the schema; baseline is 3. The description adds no parameter-level meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Return the tools exposed by an MCP server.' This distinguishes it from siblings like list_mcp_server_prompts and list_mcp_server_resources. However, it does not explicitly name or contrast with those siblings, leaving differentiation to the reader.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (list tools before calling them) but never stated. There is no explicit when-to-use, when-not-to-use, or mention of alternatives such as call_mcp_tools or retrieve_mcp_server. The connected-state caveat gives partial context but not selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsList modelsBRead-onlyIdempotent
List the LLMs and preset model chains available to the caller, grouped for display in a model picker. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| team_id | No | Scope model availability to a specific team. When omitted, uses the authenticated user's default organization. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so safety is covered; the trailing 'Read operation.' sentence largely repeats that. The genuinely additive content is the hint that results are 'grouped for display in a model picker,' which describes the shape of the response and partially compensates for the missing output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the resource being listed. The only waste is 'Read operation.', which duplicates readOnlyHint from the annotations rather than earning its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-parameter listing tool with full schema coverage and safety annotations, the description supplies the one thing structured fields do not: that results are grouped for a picker UI. Remaining gaps (result ordering, size, pagination) are minor for this tool class and no output schema exists to lean on.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both optional parameters (account, team_id) are documented in the schema itself, so the baseline is 3. The description adds no additional meaning about what these filters do to the returned model list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'List the LLMs and preset model chains available to the caller.' The scope ('available to the caller') is precise, but it never mentions the neighboring route_model tool, so sibling differentiation is left implicit rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through 'grouped for display in a model picker' but gives no explicit when-to-use guidance and no exclusions. With route_model in the sibling list, an agent would benefit from knowing this is a discovery/enumeration call rather than a dispatch call, and that distinction is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_organization_evaluation_resultsList evaluation resultsARead-onlyIdempotent
Cursor-paginated results for one evaluation across every agent it grades, newest first. Each session appears once with its latest result; queued and in-progress results are included so a run can be followed to completion. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| grade | No | Filter by grade. | |
| cursor | No | Pagination cursor from a previous response's `next_cursor`. | |
| status | No | Filter by status. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| agent_id | No | Only results for this agent. | |
| page_size | No | Items per page (1-100). | |
| session_id | No | Only results for this session. | |
| created_after | No | Only results created at or after this time. RFC 3339 with an explicit offset (for example `2026-09-01T00:00:00Z`). | |
| evaluation_id | Yes | ID of the organization evaluation. | |
| created_before | No | Only results created before this time. RFC 3339 with an explicit offset. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations carry the safety profile (readOnlyHint, idempotentHint, openWorldHint, destructiveHint=false), so the bar is lower. The description still adds real behavioral value beyond them: newest-first ordering, session deduplication, and the inclusion of queued/in-progress results that make it suitable for polling a run.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero padding, front-loading the cursor-paginated scope and ordering before the dedup rule. 'Read operation.' is a useful one-word safety restatement at the end.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete enough given a 10-param schema with 100% coverage and annotations covering the safety profile. The only gap is the absence of any routing hint to the sibling singleton result tool, which would help in a list of ~90 siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all 10 params, including enums for grade and status and filter semantics. The description adds no parameter-level detail beyond what the schema already states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (list) + resource (evaluation results for one evaluation) with explicit scope: 'across every agent it grades, newest first'. Names the ordering and the deduplication rule ('each session appears once with its latest result'), which distinguishes it from get_organization_evaluation_result (singular) in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the name and content, and the description notes that queued/in-progress results are included 'so a run can be followed to completion' — a hint about when to use it. But there is no explicit when-to-use vs get_organization_evaluation_result / get_organization_evaluation_metrics, and no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_organization_evaluationsList evaluationsARead-onlyIdempotent
Cursor-paginated list of an organization's evaluations with their targets, coverage, and result rollups. Requires the organization:manage_evaluations permission (Enterprise plan).
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | Pagination cursor from a previous response's `next_cursor`. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| page_size | No | Items per page (1-100). | |
| organization_id | Yes | The organization whose evaluations to list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive. The description adds genuinely useful context beyond them: the required permission scope (organization:manage_evaluations), the Enterprise plan requirement, and that results are cursor-paginated. It does not, however, describe pagination behavior in more depth or any rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose and return content, then the permission gate and read-only nature. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly names the returned data (targets, coverage, result rollups) and the permission prerequisite. Pagination is mentioned and the parameters are fully schema-documented. Minor gap: no guidance on navigating to related results or metric tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (cursor, account, page_size, organization_id) are already documented. The description only echoes the cursor-pagination concept, adding no syntax or format meaning beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (list) and resource (an organization's evaluations) and details what is returned: targets, coverage, and result rollups. It is clearly distinguishable in scope from the bare list_evaluations sibling, though it never explicitly names that sibling to route the agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'Read operation' and the permission note, but there is no explicit when-to-use/when-not guidance or reference to the alternatives (e.g., list_evaluations, get_organization_evaluation). The prerequisite gate (Enterprise plan, organization:manage_evaluations) is the only real usage signal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_organizationsList organizationsARead-onlyIdempotent
Returns the organization the authenticated user belongs to. Use its id as organization_id on the evaluation endpoints.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The trailing 'Read operation.' simply restates readOnlyHint, and the description does not disclose return shape or credential/identity behavior beyond what the annotation set provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the core purpose front-loaded and the downstream usage tip immediately after. The 'Read operation.' sentence is mildly redundant against readOnlyHint but does not harm readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, no-output-schema tool, the description conveys what is returned and how to consume it, which is largely sufficient. It stops short of describing the organization fields returned, a minor gap given there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single optional `account` parameter is already well documented in the schema as selecting private credentials and identity. The description adds nothing beyond what the schema states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Returns the organization the authenticated user belongs to'), clarifying the singular scope despite the plural-sounding name 'list_organizations'. It is distinguishable from list_accounts, which covers a different concept, though the description does not explicitly name siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tells the agent the downstream purpose of the call: feed the returned `id` into the evaluation endpoints as `organization_id`. That gives clear context for when the tool is useful, but offers no explicit alternatives or conditions for skipping it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_queued_messagesList queued messagesARead-onlyIdempotent
List the messages waiting in a session's queue, in the order they will be sent. Queued messages are drained automatically when the agent finishes its current turn. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| session_id | Yes | ID of the session whose queue to list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), so the bar is lower, yet the description adds genuine behavioral value: queued messages are drained automatically at the end of the current turn and are returned in send order. It does not describe message contents or pagination, but the added lifecycle context is meaningful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two efficient sentences front-load the core action and the queue-drain behavior adds real value. The trailing sentence 'Read operation.' is redundant with readOnlyHint=true and could be dropped.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read tool with full annotation coverage and no output schema, the description supplies enough to call it correctly and understand the queue lifecycle. A brief note on what a queued message looks like would close the remaining gap left by the absent output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so session_id and account are already fully documented in the schema, giving a baseline of 3. The description adds no syntax, format, or scoping detail beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the messages waiting in a session's queue') and adds ordering ('in the order they will be sent'), which clearly identifies it as the read-side counterpart to the queue message tools. It does not explicitly name siblings like update_queued_message or delete_queued_message, so the differentiation is implicit rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains queue mechanics ('drained automatically when the agent finishes its current turn') but never says when an agent should call this versus alternatives such as list_sessions, update_queued_message, or send_queued_message. No conditions or exclusions are offered, leaving usage to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_role_credit_limitsList custom role credit limitsARead-onlyIdempotent
This endpoint lists every active custom role in an organization together with its monthly credit limit, so external systems can manage credit limits programmatically. A monthly_credit_limit of null means the role sets no limit of its own. The limit applies to each member of the role individually; when a user belongs to multiple roles, the highest limit across their roles wins.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | Opaque cursor from a previous response's `next_cursor`; omit for the first page. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| user_id | No | Your user id -- you must be an organization admin to manage custom role credit limits. | |
| page_size | No | Number of roles per page. | |
| organization_id | Yes | The ID of the organization. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuine domain semantics: null means no self-imposed limit, the limit applies per member individually, and with multiple roles the highest limit wins. It does not mention pagination behavior, though the cursor param covers that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and scope, followed by meaningful edge-case semantics. The trailing 'Read operation.' is redundant given readOnlyHint=true, a small bit of waste in an otherwise tight description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully explains the key return-value nuance (null vs. an actual limit) and multi-role resolution. Combined with the fully documented input schema and annotations, an agent has what it needs, though explicit sibling routing would round it out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the schema, including the admin requirement on user_id and cursor semantics. The description adds no parameter-level detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it lists every active custom role in an organization along with each role's monthly credit limit. The 'every ... role' scope implicitly separates it from the singular get_role_credit_limit and set_role_credit_limit siblings, but it never names those alternatives explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The stated motivation ('so external systems can manage credit limits programmatically') implies when the tool is useful, but there is no explicit when-to-use/when-not-to-use guidance or comparison against get_role_credit_limit or set_role_credit_limit. Usage is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_sessionsList sessionsARead-onlyIdempotent
List sessions for an agent with cursor-based pagination, optional filtering, and search. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Filter sessions by type (e.g. `api`, `web`, `slack`). | |
| state | No | Filter sessions by state. | |
| cursor | No | Cursor for the next page of results. Use the `next_cursor` value from a previous response. | |
| search | No | Free-text search query to filter sessions by name or content. Also accepted as `search_query`. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| agent_id | Yes | ID of the agent whose sessions to list. Also accepts the reserved aliases `gumball` and `analytics`. | |
| page_size | No | Number of sessions to return per page. Defaults to `20`, maximum `100`. | |
| sort_order | No | Sort order for the results (e.g. `newest` or `oldest`). | |
| trigger_id | No | Filter sessions by the trigger that initiated them. | |
| creator_user_id | No | Filter sessions by the user who created them. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so "Read operation" adds nothing. The description does contribute cursor-based pagination behavior, which the annotations do not cover, so partial credit above baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, capability summary front-loaded before the redundant "Read operation" tag. Minimal waste, though the trailing sentence duplicates annotation content and earns nothing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter read-only list tool whose schema fully documents inputs and whose annotations carry the safety profile, the description covers the essentials. It omits any hint about output shape (result fields, cursor location), which is a minor gap rather than a blocking one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all ten parameters are already documented in the schema (enum states, cursor semantics, page_size bounds, aliases). The description only summarizes categories of filtering generically, adding no syntax or constraint detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("List sessions for an agent") plus the key capabilities (pagination, filtering, search). It does not explicitly differentiate from siblings like retrieve_session or list_queued_messages, so it stays short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Scoping the list to a specific agent implies the usage context, and mentioning optional filtering/search hints at the tool's role, but there is no explicit when-to-use vs retrieve_session or when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_skillsList skillsARead-onlyIdempotent
List skills the caller has access to. Filter by team, search by name, or narrow to a specific creator, related MCP server, or agent. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | Opaque pagination cursor returned in `next_cursor` from a prior page. | |
| unused | No | When set, filters to skills that have not been used. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| team_id | No | Scope the listing to a single team. When omitted, returns skills owned by the authenticated user. | |
| agent_id | No | Filter to skills attached to this agent. | |
| page_size | No | Number of skills per page. Clamped between 1 and 100. | |
| sort_order | No | Sort order for the returned skills. | newest |
| search_query | No | Case-insensitive substring match against the skill name. | |
| creator_user_id | No | Filter to skills created by this user ID. | |
| related_server_id | No | Filter to skills that reference this MCP server ID in their metadata. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, and the description's trailing "Read operation." merely restates that. The one genuinely additive detail is the access-scoping of results ("skills the caller has access to"), but pagination and ordering behavior remain unaddressed beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences that front-load the core purpose before the filter list. The trailing fragment "Read operation." is redundant with the annotations and adds a small amount of waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with full schema coverage, comprehensive annotations and no output schema, the description covers the essential scope and filterable dimensions. It omits any mention of paging or result volume, which the schema partially compensates for.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in structured form. The description only groups the filter dimensions at a high level and adds no format or syntax detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("List skills the caller has access to") and scopes the result set to the caller's permissions, which clearly separates it from create_skill/update_skill/delete_skill. It stops short of naming an alternative tool explicitly, so it is clear but not sibling-differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description enumerates the available filtering axes (team, name, creator, related MCP server, agent), which implies when the tool is useful. However, these are the same facets already documented in the schema, and it gives no explicit when-to-use/when-not guidance or exclusion against siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_teamsList teamsCRead-onlyIdempotent
List teams the authenticated caller belongs to. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, and the description's 'Read operation' merely restates readOnlyHint. It adds the caller-scoped filtering detail but says nothing about how the optional account parameter selects credentials, result size, ordering, or pagination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the scope front-loaded, which is efficient. The trailing 'Read operation.' sentence is wasted space because it duplicates readOnlyHint=true already in the annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool with full annotation coverage and no output schema, the description is minimally adequate. It does not tell the agent what a returned team looks like or what identifiers it yields for downstream calls, which would be the useful addition here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single optional 'account' parameter is already documented in the schema as selecting private credentials and identity. The description adds no syntax, default, or behavior detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List teams') plus the scope ('the authenticated caller belongs to'), which is meaningfully more precise than the title. It does not, however, distinguish itself from similar listing siblings such as list_organizations or list_accounts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no named alternative. An agent must infer from the name alone when this tool is preferable to the other list_* tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_workbooksList workbooks and their saved flowsCRead-onlyIdempotent
List workbooks and their saved flows Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| user_id | No | The user ID for which to list workbooks. Required if project_id is not provided. | |
| project_id | No | The project ID for which to list workbooks. Required if user_id is not provided. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the 'Read operation.' sentence is a pure restatement of structured data and adds no new behavioral context. Nothing is said about scoping, credential selection via the account parameter, or result size.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very short and front-loaded, with the resource named in the first clause. The second line ('Read operation.') is arguably redundant given the annotations, which slightly dilutes efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless-required list tool with rich annotations and full schema coverage, most needs are met. With no output schema, though, the description could say more about the shape of the returned workbook/flow data, which it does only vaguely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the account, user_id and project_id parameters including their mutual exclusivity. The description adds no parameter meaning beyond that, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List workbooks and their saved flows'), which is clearly more informative than the bare name. However, it offers no differentiation from siblings like list_flows or list_agents, and doesn't clarify the relationship between workbooks and their saved flows.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus list_flows or other list endpoints, and no mention that the user_id/project_id parameters are mutually exclusive alternatives. The agent must infer usage entirely from the schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_permission_group_usersManage custom role usersADestructive
This endpoint allows organization administrators to add or remove users from a custom role (formerly "permission group"). Adding a user to a role does not remove them from any other role they belong to. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| action | No | The action to perform - either 'add' or 'remove' a user. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| user_id | No | Your user id -- you must be an organization admin to manage custom role users. | |
| group_id | No | The ID of the custom role to manage users for. | |
| user_email | No | The email address of the target user to add or remove. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| organization_id | No | The ID of the organization that the custom role belongs to. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructive=true and idempotent=false. The description adds genuinely non-obvious behavior: admin-only authorization, that an add does not revoke other role memberships, that explicit confirmation is required, and a warning against auto-resubmitting unknown outcomes. That goes meaningfully beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core capability, which is good, but the second paragraph mixes a relevant confirmation rule with generic boilerplate ("Runs can spend credits or trigger downstream actions") that is not specific to this account-operation tool and dilutes the message.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent mutation with no output schema, the description covers authorization, the confirmation gate, and the key membership side effect, which is enough for correct invocation. It leaves the relationship to manage_project_users and the account/payload precedence unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema. The description only echoes the add/remove semantics and admin requirement without adding format or usage detail for the nine parameters (e.g., account vs organization_id interplay), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (add or remove) and resource (users in a custom role, formerly "permission group") and scopes it to organization administrators. The analogous sibling manage_project_users is never referenced, so the agent must infer the distinction from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a prerequisite (must be an org admin) and a useful side-effect rule (adding does not remove from other roles), plus a confirmation requirement. However, it never states when to prefer this tool over manage_project_users or other role-related siblings, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_project_usersManage workspace usersBDestructive
This endpoint allows organization administrators to add or remove users from a workspace. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| action | No | The action to perform - either 'add' or 'remove' a user. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| user_id | No | Your user id -- you must be an organization admin to manage workspace users. | |
| is_admin | No | When adding a user, specify whether they should have admin privileges (default is false). | |
| user_email | No | The email address of the target user to add or remove. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| workspace_id | No | The ID of the workspace to manage users for. | |
| organization_id | No | The ID of the organization that the workspace belongs to. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true, and idempotentHint=false. The description adds auth requirements (admin role) and a warning about credits/downstream actions plus avoiding automatic resubmission of unknown outcomes—useful, but somewhat generic and not specific to this endpoint's actual side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences, front-loaded with the core purpose, then the confirmation requirement, then the caution. Efficient, though the last sentence is a generic boilerplate warning that could apply to many tools.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutation tool with nested payload, multiple input modes (flags, payload, file), no output schema, and destructive semantics, the description is adequate but thin. It doesn't explain the interplay of payload vs flags vs payload_file, which is a significant omission for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema fully documents all 10 parameters including the nested payload and file alternatives. The description adds no parameter-level detail beyond the schema, which is the baseline case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: adding or removing users from a workspace, scoped to organization administrators. It's clear, though it doesn't explicitly differentiate from the highly related sibling manage_permission_group_users.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the actor (organization administrators) and requires explicit confirmation, which implies usage context. But it doesn't state when to use this versus alternatives like manage_permission_group_users, nor when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
queue_session_messageQueue messageADestructive
Add a message to a session's queue instead of interrupting the agent. Queued messages are sent automatically, in order, when the agent finishes its current turn. A session's queue holds at most 20 messages.
To interrupt the current turn and send a queued message immediately, use Send queued message now. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| input | No | The message to queue. Cannot be empty. Also accepted as `message` for backwards compatibility. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| session_id | Yes | ID of the session to queue the message on. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare mutation, non-idempotency, and destructiveness, but the description adds behavioral detail they cannot carry: the 20-message capacity cap, automatic ordered delivery when the current turn ends, the confirmation requirement, and the credit/downstream-action warning. This is meaningful context beyond the annotation set, with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, followed by the alternative and then the confirmation warning; every sentence carries weight. The markdown link syntax in the routing sentence is slightly awkward but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with a nested payload object and no output schema, the description covers purpose, alternative, capacity, ordering, and the confirmation/credits caveat. The remaining parameter distinctions are fully handled by the 100%-covered schema, so nothing an agent needs is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (input, account, confirm, payload, session_id, payload_file) is already documented. The description only references 'the message to queue' and the capacity limit, adding little beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Add a message to a session's queue') and immediately frames it against the interrupting alternative, distinguishing it from send_message. An agent can tell it apart from the sibling queued-message tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the condition for use (queue instead of interrupting) and routes to the alternative (send_queued_message for immediate interruption), plus a confirmation precondition. When-to-use, when-not, and the alternative are all stated rather than inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_mcp_server_resourceRead MCP server resourceARead-onlyIdempotent
Read one resource by uri. Each content item is either text or a base64 blob, never both.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| uri | Yes | The resource `uri` from List MCP server resources. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| team_id | No | ||
| server_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive, so safety is covered; the description adds genuinely non-obvious return semantics — each content item is either `text` or a base64 `blob`, never both. It omits auth/credential behavior for the `account` parameter and any error/rate-limit context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and resource, then the return-shape caveat. The trailing 'Read operation.' is largely redundant with readOnlyHint=true, but the total footprint is small enough that it does not waste much space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no output schema and two undocumented parameters, the definition is on the thin side; it clarifies the payload shape but not what a 'resource' is relative to MCP prompts/tools, nor the team/account scoping needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: `uri` and `account` are documented while `team_id` and `server_id` have empty descriptions. The description explains `uri` and where to obtain it, but adds nothing for `server_id` or `account`, so half the parameters remain unexplained in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read one resource by `uri`'), and the singular 'one resource' implicitly contrasts with the sibling list_mcp_server_resources. It never names a sibling explicitly, so an agent must infer the list-then-read relationship rather than being told it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the schema note on `uri` ('from List MCP server resources') hints at the list-then-read workflow, but the description gives no explicit when-to-use or when-not guidance and does not distinguish this from get_mcp_server_prompt or call_mcp_tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rename_sessionRename sessionADestructive
Rename a session. name is the only mutable field; it is trimmed and must be between 1 and 256 characters after trimming.
Returns the full session, in the same shape as Retrieve session. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| session_id | Yes | ID of the session to rename. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=true, idempotent=false and openWorld=true; the description adds the confirmation gate, the trimming/normalization rule, and the return shape, which are not in the annotations. The credit/downstream-action sentence is fairly generic boilerplate that reads as policy text rather than tool-specific behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the operation and the key constraint, then the return shape. Efficient overall, though the closing sentence about credits and automatic resubmission is generic and not obviously tied to a rename.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description compensates by stating the return is the full session in retrieve_session's shape. The confirmation requirement and normalization rule are covered; the remaining gap is the unexplained relationship among `name`, `payload` and `payload_file`.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 83%, so the schema already documents `name`, `session_id`, `account`, `confirm` and the payload options; the description mostly repeats the trimming/1-256 rule. It does not explain how `payload`/`payload_file` relate to the body flags, which is the one parameter relationship an agent still has to guess at.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Rename a session') and immediately narrows scope by naming the only mutable field. It also anchors the return shape to the sibling `retrieve_session`, so an agent knows exactly what comes back without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives real when-to-use context: explicit confirmation is required for this exact operation, and the description warns against automatic resubmission of unknown outcomes. It does not enumerate alternatives (there is no competing rename sibling), so it falls just short of explicit when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_session_approvalsResolve approvalsADestructive
Answer pending asks on a session that is paused in the approval_required state — tool approvals, human input requests, and checkpoints.
List the pending asks with Retrieve session: each entry in pending_approvals carries the action_request_id to answer, and human_input asks include the questions to fill in via response.values. Resolutions are processed in order; the agent resumes once the pending asks are answered.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| session_id | Yes | ID of the session with pending approvals. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| approval_responses | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (destructiveHint=true, openWorldHint=true, idempotentHint=false) cover safety, and the description adds substantive context: credit spend, downstream actions, ordered processing of resolutions, and the resume-once-answered behavior. The explicit confirmation requirement maps to `confirm` and the warning about unknown outcomes is high-value operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences front-loaded with the purpose, then retrieval guidance, then safety warning. Dense but each sentence earns its place. Minor density concern but no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the workflow (retrieve then resolve), the ordered processing semantics, and the confirmation safety. No output schema exists; return values are not described, and the payload/payload_file alternatives are left to the schema. Adequate for a complex mutation tool with 6 params.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so the schema documents most parameters (action, reason, response.values, action_request_id). The description adds meaning to `action_request_id` and `response.values` sourcing and confirms the confirm/credits semantics, but doesn't clarify payload vs payload_file vs body flags alternation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (resolve/answer) and resource (pending asks on a paused session) with the exact precondition 'paused in the `approval_required` state'. Distinguishes its scope (tool approvals, human input, checkpoints) in a way no sibling replicates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent to first call Retrieve session to list pending asks and where to find `action_request_id` and `questions`. Also warns not to resubmit unknown outcomes automatically. Lacks an explicit 'when NOT to use' against siblings like send_message or queue_session_message, but the prerequisite condition is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retrieve_agentRetrieve agentARead-onlyIdempotent
Retrieve a single agent by ID.
In addition to regular agent IDs, agent_id accepts the reserved aliases gumball (your personal Gumball agent) and analytics (your analytics agent) on all agent-scoped endpoints. The alias resolves to your own copy of the platform agent, creating it on first use, and responses report the alias back as the agent's id.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| agent_id | Yes | ID of the agent to retrieve. Also accepts the reserved aliases `gumball` and `analytics`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely non-obvious behavior: passing the `gumball` or `analytics` alias resolves to the caller's own copy and 'creates it on first use' — a side effect not captured by any annotation — plus how the alias is echoed back in responses. It stops short of describing error behavior or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, followed by the alias note, with no wasted preamble. The trailing 'Read operation.' partially duplicates the readOnlyHint annotation and could be dropped, but the text is otherwise tight and well ordered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with 100% schema coverage and full annotations, the description covers the essentials and adds the alias-resolution behavior the agent could not otherwise know. With no output schema it also hints at the return shape ('responses report the alias back as the agent's id'), though it omits error and not-found behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, and the schema already documents both `account` and the alias-accepting `agent_id`. The description goes beyond the schema by explaining what the alias actually resolves to and the create-on-first-use semantics, adding real meaning to the `agent_id` parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource with scope: 'Retrieve a single agent by ID,' which implicitly separates it from list_agents and the create/update siblings. However, it never names an alternative tool, so differentiation is by inference rather than explicit contrast, keeping it just below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the singular 'by ID' framing, but there is no explicit when-to-use guidance, no statement of when to prefer this over list_agents or retrieve_agent_version, and no exclusions or prerequisites. The alias discussion is contextual but not routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retrieve_agent_versionRetrieve agent versionARead-onlyIdempotent
Retrieve one immutable agent version: its full configuration (composition) plus the structured changes relative to the version before it. Use it to export an agent's configuration or to audit what changed between versions.
changes is null for the first version of an agent, since there is no predecessor to diff against. Versions created before attachment snapshots were recorded report composition.complete: false (and changes.attachment_changes_complete: false); their skill_ids and knowledge_sources are null rather than empty. Skill file contents are never included, and this endpoint is read-only — it cannot restore or deploy a version.
Requires configuration access on the agent — callers limited to using the agent (no configuration access) get a 403.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| agent_id | Yes | ID of the agent the version belongs to. Also accepts the reserved aliases `gumball` and `analytics`. | |
| version_id | Yes | ID of the version to retrieve, from `GET /agents/{agent_id}/versions`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavioral context beyond the annotations: it explains that `changes` is null for the first version, that older versions report `composition.complete: false` and `changes.attachment_changes_complete: false`, that `skill_ids` and `knowledge_sources` may be null rather than empty, that skill file contents are never included, and that configuration access is required with a `403` otherwise.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and usage in the first two sentences, then uses short paragraphs for edge cases and access requirements. Every sentence contributes necessary information without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description explains the key returned fields (`composition`, `changes`), edge cases such as null values and incomplete snapshots, and the permission model. This is sufficient for an agent to call and interpret the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the input schema already documents all three parameters, including the meaning of `agent_id` aliases and where `version_id` comes from. The description does not add parameter-level syntax or format details beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb and resource: 'Retrieve one immutable agent version', and specifies the returned content as 'full configuration (`composition`) plus the structured `changes` relative to the version before it.' This clearly distinguishes it from listing versions or retrieving a whole agent, even without naming sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance: 'Use it to export an agent's configuration or to audit what changed between versions.' It also states a when-not boundary by noting the endpoint is read-only and 'cannot restore or deploy a version.' However, it does not name alternative tools such as `list_agent_versions` for related operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retrieve_evaluationRetrieve evaluationARead-onlyIdempotent
Retrieve a single evaluation result by ID. The evaluation must belong to the specified agent. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| agent_id | Yes | ID of the agent the evaluation belongs to. Also accepts the reserved aliases `gumball` and `analytics`. | |
| evaluation_id | Yes | ID of the evaluation to retrieve. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint and openWorldHint, so the safety profile is fully covered. The description's 'Read operation' simply restates readOnlyHint, adding no new behavioral detail; the only real addition is the agent-ownership constraint, and nothing is said about error behavior when ownership fails.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the primary action front-loaded and no wasted prose. The trailing 'Read operation' clause is redundant with the readOnlyHint annotation, which is a minor inefficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only fetch by ID with full schema coverage, complete annotations and no nested objects, the definition covers what an agent needs to call it correctly. The absence of an output schema means the return shape is undocumented, a small residual gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (including the account selector, aliases, and min lengths) are already documented in the schema. The description only reinforces the agent_id/evaluation_id pairing and adds no syntax or format detail, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Retrieve a single evaluation result by ID'), and the agent/ownership scoping sets it apart from list-style siblings like list_evaluations. It does not explicitly name the closest alternative (get_organization_evaluation), so sibling differentiation is only partial.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The constraint 'must belong to the specified agent' gives useful context for when this call is valid, but there is no explicit when-to-use guidance or exclusion versus get_organization_evaluation / list_evaluations. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retrieve_mcp_serverRetrieve an MCP serverBRead-onlyIdempotent
Return a single MCP server. The response populates allowed_tool_call_ids with the tool call IDs the caller is permitted to invoke on this server.
Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| team_id | No | Scope the lookup to a single team. | |
| server_id | Yes | Identifier of the MCP server to retrieve. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so "Read operation" adds nothing. The genuinely useful addition is the disclosure that the response populates allowed_tool_call_ids with the caller's permitted tool calls — a return-side behavior not covered by annotations or any output schema. It does not mention auth needs, rate limits, or error behavior, so a 3 fits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the purpose front-loaded, followed by the output detail. The trailing "Read operation" is redundant against readOnlyHint and slightly wastes space, but overall the entry is tight and well-ordered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-resource retrieve with 100% schema coverage and annotations carrying the safety profile, the description is nearly enough: it names the resource and one key response field. It omits any mention of the plural-list alternative, which is the main missing piece.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so account, team_id, and server_id are already documented in the schema. The description adds no parameter-level detail (e.g., whether account changes the scoping or what happens when team_id is omitted), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: "Return a single MCP server." The singular "single" implicitly distinguishes it from the sibling list_mcp_servers, so an agent can route without opening the schema. It stops short of explicitly naming the sibling as an alternative, so it earns a 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use or when-not guidance and never references the surrounding toolset (list_mcp_servers, list_mcp_server_tools). Retrieval-by-id is implied by the required server_id, but nothing steers an agent between this and the list siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retrieve_sessionRetrieve sessionARead-onlyIdempotent
Retrieve a session by ID, including its messages, current state, agent metadata, and participants. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| session_id | Yes | ID of the session to retrieve. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the appended 'Read operation' is largely redundant. The description does add useful disclosure that the payload includes messages, state, agent metadata, and participants, but nothing about auth, rate limits, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the primary purpose and returned payload front-loaded and no wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned content, which is the key missing structured information. It is largely complete for a read-by-ID tool, though it omits error/not-found behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both account and session_id are already documented in the schema. The description reinforces that retrieval is keyed on the ID but adds no syntax or format detail beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Retrieve) and resource (session) and enumerates what is returned (messages, state, agent metadata, participants), which distinguishes it from list_sessions. It could be sharper by explicitly contrasting with the sibling list_sessions, but the 'by ID' scope makes the difference clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'by ID' phrasing implies the tool is for when a specific session identifier is already known, but there is no explicit when-to-use routing against alternatives like list_sessions or get_run_details. Usage is merely implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
route_modelRoute a message to a modelARead-onlyIdempotent
Ask Gumloop Chew — the model router behind Auto — which model it would use for a given
message, and why. Chew is a router, not a model: router is always gumloop-chew, and the
concrete model it selected is route.model.
This endpoint returns a decision only. It does not run the selected model or its fallbacks. The routing judgement consumes credits and is bounded by the caller's model access.
Omit models to use the deduplicated union of Chew's lane chains, not every model the
caller may use. Restricted candidates can appear with status: "restricted"; they are never
selected or included in fallback_models.
Team scope requires actual team membership, even within the same organization. Personal API keys and OAuth are supported; this is not team-key-only. Read operation.
| Name | Required | Description | Default |
|---|---|---|---|
| agent | No | ||
| input | No | ||
| models | No | ||
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| history | No | ||
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| team_id | No | Scope model availability and credit attribution to a team the caller belongs to. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description still adds real context beyond them: the call consumes credits, it is bounded by the caller's model access, it never executes the selected model or its fallbacks, and restricted candidates surface with status "restricted" but are never selected. Return-shape detail (fallback_models) is thin, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then decision-only semantics, then parameter caveats. Every paragraph is topical, though the second paragraph's 'Chew is a router, not a model' partially restates the opening line, which costs a little density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a decision-only routing endpoint with no output schema, the description supplies the key return contract (router is always gumloop-chew, selected model at route.model, restricted status, fallback_models exclusion) plus auth/scope prerequisites. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% and 8 parameters exist, so the description must compensate. It clarifies `models` (omit -> lane-chain union rather than every available model) and `team_id` scoping, but leaves `account`, `payload`, and `payload_file` unexplained, and the `models` note largely repeats the schema's own 'Omit to use Chew's lane-chain union'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('ask the model router which model it would use') and immediately distinguishes itself from the neighbours: 'Chew is a router, not a model' and 'This endpoint returns a decision only. It does not run the selected model.' An agent can separate this from create_chat_completion or list_models without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage context: omit `models` to fall back to the lane-chain union, team scope requires actual membership, and personal API keys/OAuth are supported. It never explicitly names an alternative tool or states a when-not-to-use condition (e.g. 'use create_chat_completion to actually run a model'), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_evaluationsRun evaluationsADestructive
Grades up to 200 of the agent's finished sessions with its own evaluation configuration. Grading is asynchronous: each accepted session gets a result with status: queued; poll it with GET /agents/{agent_id}/evaluations/{evaluation_id} until it is completed or failed. A new result replaces the previous result for that session.
Sessions are skipped, not rejected, when they are unfinished, incognito, or not owned by this agent (ineligible), or already have a queued or running result (in_flight, with the existing result_id). The caller is charged one credit per queued session. Set dry_run: true to see the cost and skips without queuing anything.
Requires edit access on the agent and a plan with evaluations enabled. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| dry_run | No | Report cost and skipped sessions without queuing. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| agent_id | Yes | ID of the agent that owns the sessions. | |
| session_ids | No | Sessions to grade. Duplicates are rejected. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (destructiveHint/openWorldHint only flag risk). It discloses asynchronous queuing with status: queued, the polling endpoint to reach completed/failed, replacement of prior results, skip-not-reject semantics with ineligible/in_flight reasons and result_id, per-session credit cost, and the confirmation/anti-resubmission rule. This is unusually complete behavioral disclosure for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the verb, resource and cap in the first sentence, then layers async behavior, skip semantics, cost, and prerequisites in short paragraphs. Slightly dense across four paragraphs, and the closing confirmation boilerplate is the least information-dense part, but every block carries operational meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent mutation with no output schema, the description supplies the missing return information (queued status plus polling path), the cost model, permission prerequisites, and partial-success behavior. An agent has everything needed to invoke it correctly and safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so a 3 is the baseline, but the description adds genuine semantics: dry_run reports cost and skips without queuing, session_ids are skipped for unfinished/incognito/not-owned/already-in-flight sessions, and duplicates are rejected. This explains parameter consequences rather than restating the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (grades), the exact resource (the agent's finished sessions using its own evaluation configuration), and the hard cap (up to 200). It is clearly distinguishable from siblings like list_evaluations, retrieve_evaluation, get_evaluation_metrics and update_evaluation_config, which configure or read rather than execute grading.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong operating context: prerequisites (edit access on the agent, a plan with evaluations enabled), a recommended dry_run: true preview path, and an explicit confirmation requirement. It does not, however, name a sibling alternative for related tasks (e.g. configuring or reading evaluations), so the routing guidance is contextual rather than comparative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_organization_evaluationRun evaluation on sessionsADestructive
Grades up to 200 existing sessions with this evaluation. Grading is asynchronous: each accepted session gets a result with status: queued; poll it with GET /evaluations/{evaluation_id}/results/{result_id} until it is completed or failed.
Sessions are skipped, not rejected, when they are not completed sessions of an agent the evaluation covers (ineligible) or already have a queued or running result for this evaluation (in_flight, with the existing result_id). The caller is charged one credit per queued session. Set dry_run: true to see the cost and skips without queuing anything.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| dry_run | No | Report cost and skipped sessions without queuing. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| session_ids | No | Sessions to grade. Duplicates are rejected. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| evaluation_id | Yes | ID of the organization evaluation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Substantially exceeds the annotations: it discloses asynchronous grading, the queued status and the exact polling endpoint, skip-vs-reject semantics with the ineligible and in_flight reasons, the one-credit-per-queued-session charge, and a warning never to auto-resubmit unknown outcomes. This is consistent with destructiveHint=true and adds real operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded and the async/polling detail follows logically. Three short paragraphs with little filler, though the credit and confirmation warnings are slightly repetitive with schema descriptions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, nested, async, credit-spending tool with no output schema, the description supplies everything an agent needs: batching limit, async flow, polling path, skip conditions, cost model, dry-run preview and confirmation requirement. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents account, confirm, dry_run, session_ids, payload and payload_file. The description restates the dry_run effect and the 200-session cap but adds no syntax or semantics beyond the structured fields, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a specific verb and resource ('Grades up to 200 existing sessions with this evaluation'), so an agent immediately knows what the tool does. It does not, however, explicitly distinguish itself from the sibling run_evaluations, which a 5 would require.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear situational guidance: use dry_run:true to preview cost and skips without queuing, and confirmation is required for this exact account operation. It explains skip conditions (ineligible, in_flight) but never names an alternative sibling such as run_evaluations, so it stops short of full when/when-not/alternative coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_brainSearch Company BrainADestructive
Run a hybrid (semantic + keyword) search across the knowledge sources indexed in your Company Brain and return the most relevant, ranked snippets with citations.
Results are scoped to what the authenticated user can see: personal sources, plus any team and organization sources shared with them. Requires the Brain feature, which is available on the Pro and Enterprise plans. Each search consumes Gumloop credits. Explicit confirmation is required for this credit-consuming Brain search.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return. | |
| query | No | The natural-language search query. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| source_type | No | ||
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations mark this as non-read-only and destructive, and the description supplies the reason an agent needs to be careful: it consumes Gumloop credits and demands explicit confirmation. That is genuine context beyond the annotation flags. It still does not describe result volume, ranking behavior, or pagination, so it stops short of full disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and result format, then layers scoping, prerequisites, and cost in descending order of importance. Four sentences is slightly more than needed, and the inline documentation link is of marginal value to an agent, but nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with a nested payload object and no output schema, the description covers purpose, access scoping, plan gating, credit cost, and the confirmation gate. It leaves the payload/payload_file alternates and account semantics to the schema, which is acceptable given the high coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 86%, so the schema already documents query, limit, source_type, account, confirm, and payload. The description adds only that the query is hybrid semantic+keyword in nature, which mildly enriches the query parameter but does not explain the confirm or payload flags. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search), the mechanism (hybrid semantic + keyword), the resource (knowledge sources in the Company Brain), and the return shape (ranked snippets with citations). It is clearly distinguishable from siblings like list_brain_sources or get_brain_source, which browse rather than query.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives real conditions: results are scoped to what the authenticated user can see, the Brain feature must exist (Pro/Enterprise only), each call consumes credits, and explicit confirmation is required. It does not name an alternative tool for browsing indexed sources, so routing is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_messageSend messageADestructive
Append a user message to an existing session and resume the agent. The session must be idle, completed, failed, or approval_required; sending to a session that is processing or queued returns 409 interaction_not_in_terminal_state. To hand the agent a message while it is still busy, use the message queue instead.
Files uploaded via Upload session file can be attached to the message with attachments.
Sessions waiting on an approval
A session that stopped to ask you something is approval_required, and you have two ways to move it forward:
Answer the ask. Send the pending asks' responses to Resolve approvals. Use this to approve or reject a tool call, or to answer an Ask Question the agent raised. This endpoint rejects
approval_responseswith a400.Send a follow-up instead. Post a normal message here. It is appended to the session transcript and starts a new turn, leaving the pending ask unanswered. Use this when the answer no longer matters — for example to redirect the agent or drop the request it was asking about.
See Human in the Loop for how agents pause for approvals and questions.
Streaming the response
api.gumloop.com only serves the non-streaming response above. To stream agent output as it's produced, send the same request body (with stream: true) to the streaming host instead:
POST https://ws.gumloop.com/api/v1/sessions/{session_id}/messagesThe response is text/event-stream (Server-Sent Events). With the Python SDK, client.sessions.stream_message(session_id, input="...") routes to ws.gumloop.com automatically and yields parsed StreamEvent objects.
If you send stream: true to api.gumloop.com by mistake, the response is a 400 whose body contains the correct streaming host so you can retry against it.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| input | No | The next user message. Required. Also accepted as `message` for backwards compatibility. | |
| stream | No | Must be `false` (or omitted) when calling `api.gumloop.com`. Set to `true` only when calling `ws.gumloop.com` (see the streaming section above). | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| session_id | Yes | ID of the session to continue. | |
| attachments | No | ||
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true and idempotentHint=false, and the description reinforces and explains them: runs "can spend credits or trigger downstream actions," explicit confirmation is required for this exact account operation, and unknown outcomes must never be resubmitted automatically. It goes further with host-routing behavior (stream:true against api.gumloop.com yields a 400 containing the correct streaming host) that no annotation conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded — purpose, preconditions, and the sibling alternative come first, with approval and streaming details in clearly headed sections. However, the streaming subsection (curl example, SDK call, host-mismatch error) is long for a tool description and pushes the confirmation warning to the very end where it is easy to miss.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter, nested-object, non-idempotent mutation tool with no output schema, the definition covers the state machine, error codes, attachment provenance, streaming host split, and approval interactions. The only minor gap is that the non-streaming response shape is only referred to as "the response above" rather than described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 88%, so the baseline is 3, but the description adds real meaning: attachments must be stored paths returned by Upload session file for this session, and `stream` must be false on api.gumloop.com and true only on ws.gumloop.com. It also implicitly explains `confirm` via the confirmation requirement, though it never enumerates `account`, `payload`, or `payload_file` semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+effect: "Append a user message to an existing session and resume the agent." It explicitly distinguishes itself from sibling operations by naming the message queue for busy sessions and resolve_session_approvals for pending asks, so an agent can route correctly without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States precise preconditions (session must be idle/completed/failed/approval_required) and the failure mode for the wrong state (409 interaction_not_in_terminal_state). It names the alternative endpoint (message queue) for the busy case and lays out the two competing paths for approval_required sessions, including when a follow-up message is preferable to answering the ask.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_queued_messageSend queued message nowADestructive
Send a queued message immediately instead of waiting for the agent to finish its current turn. Any in-progress run is aborted, the queued message is appended to the transcript, and the agent starts processing it.
The response is the same envelope as Send message. Queued messages cannot be sent this way on incognito sessions, and a message that is currently being edited must have its edit finished or cancelled first. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| session_id | Yes | ID of the session the queued message belongs to. | |
| queued_message_id | Yes | ID of the queued message to send. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (destructive, non-idempotent, open-world), the description discloses the concrete side effects: any in-progress run is aborted, the message is appended to the transcript, and the agent starts processing. It also adds the confirmation requirement, credit-spend warning, and the 'never resubmit unknown outcomes' caution, which are not captured by the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then side effects, then caveats and the confirmation policy. Three short paragraphs, each earning its place; no filler, though the confirmation sentence is slightly redundant with the confirm param.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent mutation with no output schema, the description covers side effects, preconditions, and confirmation. It notes the response matches the 'Send message' envelope, addressing the return shape, though it does not detail error/abort outcomes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters, including confirm and account. The description reinforces the confirm requirement ('explicit confirmation is required for this exact account operation') but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (send), resource (queued message), and the distinguishing timing/scope: 'immediately instead of waiting for the agent to finish its current turn.' This clearly separates it from siblings like queue_session_message, update_queued_message, and send_message.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It establishes when to use it (to send now rather than wait for turn completion) and a clear when-not condition ('cannot be sent this way on incognito sessions'), plus the editing precondition. It stops short of explicitly naming alternative tools for the incognito/edit cases, leaving the agent to infer the fallback.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_organization_evaluation_targetsSet evaluation targetsADestructive
Replaces the full set of targets — who the evaluation grades. Targets expand to agents live: organization covers every agent in the organization, team every agent a team owns, user a member's personal agents, agent one agent. Removing the last target pauses an enabled evaluation; enabled in the response reflects that.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| targets | No | ||
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| evaluation_id | Yes | ID of the organization evaluation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing that removal of the last target pauses an enabled evaluation, that the response's `enabled` field reflects it, and that runs can spend credits or trigger downstream actions. These are non-obvious side effects an agent must know before calling a destructive, non-idempotent mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and follows with the target-expansion table and the confirmation warning; every sentence carries weight. Slightly dense in the target enumeration, but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation tool with no output schema and nested objects, the description covers replacement semantics, side effects, confirmation requirements, and a hint about the `enabled` return field. It does not explain `account`, `payload`, or `payload_file` behavior, leaving minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so the schema already documents most parameters (baseline 3). The description adds real value by spelling out what each target `type` expands to (`organization`/`team`/`user`/`agent`) and how removal affects evaluation state, enriching the enum semantics beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Replaces the full set of targets') and clarifies the target-expansion semantics, so an agent knows exactly what is being overwritten. It implicitly distinguishes itself from sibling update tools by emphasizing full replacement, but never names a sibling, which keeps it at a 4.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating context: explicit confirmation is required for this exact operation and unknown outcomes must not be auto-resubmitted. It does not, however, contrast itself against alternatives like update_organization_evaluation or update_evaluation_config, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_role_credit_limitSet custom role credit limitADestructive
This endpoint sets or clears the monthly credit limit of a custom role. The limit applies to each member of the role individually and takes effect immediately: member allowances are recalculated while preserving credits already used in the current billing cycle. Send "monthly_credit_limit": null to clear the role-level limit so members revert to the organization default. When a user belongs to multiple roles, the highest limit across their roles wins. Requests that do not change the stored value are no-ops. Changes are recorded in the organization audit trail.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| role_id | Yes | The ID of the custom role (the same ID used as `group_id` by the Manage custom role users endpoint). | |
| user_id | No | Your user id -- you must be an organization admin to manage custom role credit limits. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| organization_id | Yes | The ID of the organization the custom role belongs to. | |
| monthly_credit_limit | No | The monthly credit limit applied to each member of this role, or null to clear the role-level limit. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only flag it as a destructive, non-idempotent, open-world write. The description adds substantial context beyond that: immediate effect, per-member application, recalculation preserving already-used credits, highest-limit-wins across multiple roles, no-op behavior, and audit-trail recording. This is exactly the extra behavioral detail the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose and null semantics are front-loaded and each operational sentence earns its place. However, the closing boilerplate about 'runs can spend credits or trigger downstream actions' is generic and tangential to a credit-limit setter, slightly padding the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the definition is complete: it covers effect timing, recalculation semantics, role precedence, no-op behavior, audit recording, and confirmation. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning the schema lacks: the per-member scope of the limit and the multi-role precedence rule ('highest limit across their roles wins'). The null-clearing semantics are already in the schema, so this is additive rather than duplicative.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'sets or clears the monthly credit limit of a custom role'. This clearly distinguishes it from the read siblings get_role_credit_limit and list_role_credit_limits without needing to see their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the key usage condition ('send monthly_credit_limit: null to clear the role-level limit') and the confirmation requirement, giving clear context for when to invoke. It does not explicitly name the read alternatives (get_role_credit_limit/list_role_credit_limits), so routing is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_flowStart flow runADestructive
This endpoint is used to trigger a flow run via API Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| user_id | No | The id for the user initiating the flow. | |
| project_id | No | (Optional) The id of the project within which the flow is executed. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| saved_item_id | No | The id for the saved flow. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds value beyond that by disclosing credit spend, downstream side effects, the confirmation requirement, and non-idempotent retry behavior. It does not describe the return payload, but no output schema exists and annotations carry the rest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Short and front-loaded, with the operation stated first and the confirmation warning second. 'This endpoint is used to' is mild filler, but nothing is padded or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent, open-world mutation with a fully documented schema, the description covers the key risks (credits, downstream actions, confirmation, retries). It omits how to track the resulting run (e.g., get_run_details), which a fully complete definition would mention.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with 7 well-documented parameters (account, confirm, payload, payload_file, saved_item_id, user_id, project_id). The description adds nothing about parameter meaning, so the schema does the heavy lifting; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('trigger a flow run via API'), which is clear enough to act on. It does not differentiate from siblings like kill_flow, get_run_details, or list_flows, but the core action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives real usage conditions: explicit confirmation is required, runs can spend credits, and unknown outcomes must not be auto-resubmitted. It names no alternatives (e.g., kill_flow to stop, get_input_schema to inspect before running), so routing guidance is incomplete but the when/when-not is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_agentUpdate agentADestructive
Update an existing agent. Only fields included in the request body are changed; omitted fields are left untouched.
This endpoint edits document fields only. To attach or detach skills, use PATCH /agents/{agent_id}/skills; to manage MCP servers, use the agent MCP server endpoints.
is_active: false is not a pause switch. It retires the agent: the agent disappears from GET /agents, and both GET and PATCH /agents/{agent_id} return 404 afterwards, so you cannot set it back to true through the API. To stop an agent from running on its own while keeping it fully reachable, disable its triggers instead.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| tools | No | When provided, replaces the agent's tool list. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| team_id | No | When provided, transfers ownership of the agent to this team. | |
| agent_id | Yes | ID of the agent to update. Also accepts the reserved aliases `gumball` and `analytics`. | |
| metadata | No | ||
| is_active | No | ||
| resources | No | When provided, replaces the agent's resource list. | |
| model_name | No | ID of the LLM the agent runs on. Use `GET /models` to discover valid values. | |
| description | No | ||
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| system_prompt | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (destructiveHint/openWorldHint/idempotentHint) by disclosing the non-obvious irreversible semantic of is_active:false — the agent disappears from GET /agents and returns 404 with no API path back. It adds the credential/confirmation requirement and warns against auto-resubmitting unknown outcomes, which is exactly the behavioral context an agent needs before a destructive write.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then scope, then the irreversible-warning, then the confirmation rule — a sensible priority order. Slightly long, and the 'runs can spend credits' line is adjacent to but not strictly about this endpoint, keeping it just short of 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter, nested-payload mutation with no output schema, the description covers the dangerous parameter, sibling routing, and confirmation posture well. It omits mention of the payload/payload_file/body-flag exclusivity model, but that is fully documented in the schema, so what remains is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 64%, and the description adds real meaning for the highest-risk parameter (is_active retirement semantics) and for confirm ('explicit confirmation is required for this exact account operation'). It does not add semantics for payload vs body-flags vs payload_file mutual exclusivity in the prose, but that is covered by the schema, so this exceeds the coverage baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Update an existing agent') and immediately scopes it ('only fields included in the request body are changed'). It explicitly carves out what this endpoint is NOT for, naming the skill and MCP-server siblings that handle those cases, so an agent can distinguish it from update_agent_skills/attach_agent_mcp_server without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing to alternatives ('to attach or detach skills, use PATCH /agents/{agent_id}/skills'; 'to manage MCP servers, use...') and a clear when-not for is_active ('disable its triggers instead'). It also states the confirmation precondition, so the conditions selecting this tool are fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_agent_skillsAttach or detach agent skillsADestructive
Attach and/or detach skills on an agent using deltas. This is not a replace-list: skills you don't mention are left untouched.
The operation is idempotent. Re-attaching a skill that's already attached (or detaching one that isn't) is reported under
already_attached/already_detachedrather than failing.A skill ID may not appear in both
attachanddetach.Up to 100 unique skill IDs total (
attach+detach) per request.Attaching requires
INVOKEpermission on the skill. Detaching is permissive so stale attachments can always be removed. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| attach | No | Skill IDs to attach. Ignored if already attached. | |
| detach | No | Skill IDs to detach. Ignored if not attached. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| agent_id | Yes | ID of the agent whose skills to update. Also accepts the reserved aliases `gumball` and `analytics`. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description asserts 'The operation is idempotent,' but the annotations declare idempotentHint: false. This is a direct conflict on a named behavioral trait, so per the rubric the description contradicts the annotations. The otherwise-rich content (already_attached/already_detached reporting, permission asymmetry, credit/downstream warning) cannot offset a claim that contradicts the declared hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core delta semantics in the first sentence, then uses tight bullets for the edge cases and constraints. The closing confirmation/credit paragraph partially duplicates the `confirm` parameter description, which is the only real redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries return-value burden and does so by naming the already_attached/already_detached reporting fields, plus permission, limit, and confirmation requirements. It does not cover failure modes (e.g. invalid skill IDs, missing INVOKE permission) or the account/payload alternates, so it is strong but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine semantics beyond the schema: that attach/detach are deltas rather than a replacement list, that an ID may not appear in both, and that the combined total is capped at 100. It says little about account/confirm/payload/payload_file, keeping it below 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (attach/detach) and resource (skills on an agent), and explicitly disambiguates from a replace-list semantics. An agent can distinguish this from sibling tools like attach_agent_mcp_server or update_agent without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating context: delta semantics, mutual-exclusion rule, the 100-ID cap, permission requirements for attach vs detach, and the explicit-confirmation prerequisite. It stops short of naming an alternative tool (e.g. why to use this rather than update_agent or attach_agent_mcp_server), so it is not a full when/when-not/alternatives treatment.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_evaluation_configUpdate evaluation configADestructive
Partially update the evaluation configuration for an agent. Omitted fields keep their current value. Provided list fields (criteria, tags, data_points) replace that list entirely.
Requires Pro tier or above. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Tag vocabulary (replaces existing list). Max 50. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| enabled | No | Whether evaluations are enabled for this agent. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| agent_id | Yes | ID of the agent. Also accepts the reserved aliases `gumball` and `analytics`. | |
| criteria | No | ||
| sentiment | No | ||
| model_name | No | LLM model to use for evaluation. | |
| data_points | No | ||
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| include_auto_tags | No | Allow the evaluator to suggest tags beyond your predefined vocabulary. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructive=true, idempotent=false and openWorld=true, and the description adds genuinely non-redundant behavior: list fields are wholesale replaced rather than merged, a tier prerequisite gates the call, confirmation must be tied to the exact requested action, and unknown outcomes must not be resubmitted because runs can spend credits. That is exactly the extra context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the partial-update rule and the list-replacement rule before prerequisites. The closing sentence about credits and downstream actions is somewhat generic to the evaluation domain rather than specific to updating config, so it earns slightly less than a perfect score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter, nested, mutation tool with no output schema, the description covers purpose, partial semantics, list replacement, tier gating and confirmation. It leaves the payload/payload_file/body-flag mutual exclusivity and the agent_id aliases to the schema, which is acceptable given 75% coverage, but an agent assembling a call still relies on schema reading.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 12 parameters and 75% schema description coverage, the schema carries most parameter detail, but the description adds the crucial replace-vs-merge semantics for criteria, tags and data_points that the schema does not state. It does not touch payload/payload_file vs body-flag mutual exclusivity, leaving a gap, but the added list semantics are substantive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Partially update the evaluation configuration for an agent') and immediately qualifies the scope with 'partial', which cleanly separates it from the read-side siblings get_evaluation_config / retrieve_evaluation. The verb-plus-resource pairing leaves no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives real operating conditions: omitted fields keep current values, list fields replace entirely, Pro tier is required, and explicit confirmation is required for this exact account operation. It does not name alternative tools (e.g. get_evaluation_config for reading, run_evaluations for executing), so it stops short of explicit when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_organization_evaluationUpdate evaluationADestructive
Partial update. Only the fields you send change. config is merged field by field; a list you send (criteria, tags, data_points) replaces that list wholesale. description: null clears the description.
Setting enabled: true requires at least one criterion, tag, or data point (400 organization_evaluation_empty_rubric) and at least one covered agent (400 organization_evaluation_no_targets). Emptying the rubric of an enabled evaluation pauses it.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| evaluation_id | Yes | ID of the organization evaluation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and non-idempotent, and the description adds real value beyond that: the exact 400 error codes on invalid enable, the pause-on-empty-rubric behavior, the credit-spending warning, and the confirmation requirement. This is genuinely useful behavioral context not in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the partial-update semantics in the first two sentences, then operational constraints, then the confirmation warning. Dense but every sentence carries information. Slightly heavy on back-to-back clauses.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers mutation semantics, validation errors, side effects, and confirmation requirement for a destructive non-idempotent tool. No output schema, but a mutation returning the updated resource is expected, so return value isn't needed. Missing only a pointer to sibling alternatives.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema thoroughly documents each field (id semantics, tag name casing, frequency enum, etc.). The description adds merge semantics for config and null-clears for description, which is real added value but overlap is high; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (partial update) and resource, and immediately distinguishes semantics from sibling update tools (update_evaluation_config, set_organization_evaluation_targets) by clarifying that config merges field-by-field while lists replace wholesale and null clears description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear behavioral guidance: what happens when enabling, what happens when emptying the rubric, the confirmation requirement. Doesn't explicitly name sibling alternatives like update_evaluation_config, but the operational context is clear enough to select correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_queued_messageUpdate queued messageADestructive
Replace the content of a message that is still waiting in the session's queue. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| input | No | The new message content. Cannot be empty. Also accepted as `message` for backwards compatibility. | |
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| session_id | Yes | ID of the session the queued message belongs to. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| queued_message_id | Yes | ID of the queued message to update. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true/non-idempotent, but the description adds genuinely new context: an explicit confirmation requirement and the warning that runs can consume credits and trigger downstream actions. It stops short of describing what happens to the original content or any rate/error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, purpose front-loaded, no filler. The credit/downstream-action sentence is slightly generic boilerplate relative to a queue-content update, which keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description plus a fully described 7-parameter schema cover the call adequately, including the confirmation gate for a destructive, non-idempotent write. Only the return/confirmation behavior on success is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 7 parameters, including the payload/body-flag mutual exclusion and the account selector. The description adds no syntax, format, or constraint detail beyond that baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Replace the content of a message') plus a discriminating precondition ('still waiting in the session's queue'), which separates it from send_queued_message, delete_queued_message, and send_message without opening their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating rules: explicit confirmation is required, and unknown outcomes must never be resubmitted automatically. It does not, however, name the sibling alternative to use when the message has already been sent or should be discarded instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_skillUpdate skillADestructive
Replace a skill's files with a new upload. Reparses SKILL.md to update the skill's name, description, and metadata, and creates a new version. Maximum upload size is 10 MB.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| files | No | ||
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| skill_id | Yes | ID of the skill to update. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true; the description adds real value on top by disclosing the reparse/versioning behavior, the confirmation requirement, and the credit/downstream-action warning. The only gap is that it does not describe what happens to the prior version or how failures surface.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and effect, followed by size and safety constraints. The credit/downstream-action sentence is somewhat boilerplate but carries genuine operational meaning for a destructive, non-idempotent call.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema and 83% schema coverage, the description covers effect, versioning, confirmation, and retry safety. The size-limit inconsistency with the schema and the absence of any failure-mode description are the remaining gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 83%, so the schema already documents skill_id, account, confirm, payload, and payload_file. The description's only parameter-relevant claim — 'Maximum upload size is 10 MB' — actually conflicts with the schema, which caps each file and the total at 5 MiB, so it adds confusion rather than clarity. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Replace a skill's files with a new upload') plus the concrete side effects: reparse of SKILL.md, update of name/description/metadata, and new version creation. It does not distinguish itself from siblings like create_skill, delete_skill, or update_agent_skills, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides some operational guidance — explicit confirmation is required, and unknown outcomes must not be auto-resubmitted — which implies when it is safe to call. However, it never names an alternative (create_skill, update_agent_skills, delete_skill) or states the condition that selects this tool over them, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_brain_filesUpload filesADestructive
Upload up to 25 files as multipart/form-data parts named files. Accepted types are PDF, Word, PowerPoint, Excel, and text formats (.txt, .md, .html, .csv, .rtf), each up to 25 MB.
Indexing starts on its own after the upload: an active source indexes and bills immediately, a draft source runs a credit estimate instead (see Retrieve estimate). Poll List files until each file's status is indexed.
Files the upload policy refuses (unsupported type, too large, empty) are returned in rejected with a 201; the request is a 400 no_files_accepted only when every file was refused.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| files | No | ||
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| source_id | Yes | The source id. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing billing behavior (active bill immediately, draft estimates credits), the partial-rejection semantics (rejected array with 201 vs 400 no_files_accepted), and the confirmation requirement for a destructive/open-world write. This is real value-add over destructiveHint/openWorldHint alone. Minor deduction because the stated 25 MB per-file limit conflicts with the schema's 5 MiB, undermining trust in the behavioral claims.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and limits, then explains post-upload behavior and error semantics in scannable sentences. Slightly long, and the space spent on a size limit that contradicts the schema is wasted, but every other sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter, nested-object mutation tool with no output schema, the description covers outcome behavior (rejection arrays, status codes, polling) that the schema cannot convey. It leaves gaps around the payload/payload_file alternative body mechanisms and the account parameter, but the agent has enough to call the primary path correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so the baseline is 3, and the description does add genuine param context: the multipart part name `files` and the accepted file extensions. However, it silently omits source_id, account, payload, and payload_file, and its per-file size figure (25 MB) directly contradicts the schema's 'at most 5 MiB', which is worse than saying nothing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Upload up to 25 files as multipart/form-data parts named `files`', plus the accepted formats and the source-indexing consequence. An agent can tell it is the brain-source upload tool, but the description never distinguishes it from the sibling upload_files / upload_file / upload_session_file tools, so the sibling differentiation required for a 5 is absent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives useful context on what happens after the call (draft vs active indexing, poll List files until status is `indexed`) and that explicit confirmation is required. It stops short of stating when to choose this tool over upload_files or upload_session_file, so usage is implied rather than directed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_fileUpload fileBDestructive
Upload file Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| user_id | No | The user ID associated with the file. Required if project_id is not provided. | |
| file_name | No | The name of the file to be uploaded. | |
| file_path | No | Regular local file path, no symlink, at most 3 MiB. Encoded as native base64 file_content; cannot mix with file_content or payload routes. | |
| project_id | No | The project ID associated with the file. Required if user_id is not provided. | |
| file_content | No | Base64 encoded content of the file. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=true, idempotent=false, openWorld=true, readOnly=false, and the description goes beyond them by stating that confirmation is mandatory, that runs can spend credits or trigger downstream effects, and that unknown outcomes must not be resubmitted. These caveats meaningfully inform retry and safety behavior. It still omits the multi-route input handling and size limits, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The body is two tight sentences with no filler, and the critical constraint is front-loaded. The redundant 'Upload file' first line adds little, but the rest is efficient and well ordered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with nested objects and three alternative input routes and no output schema, the description covers the confirmation/credits safety context adequately but is silent on the input-route complexity that dominates correct invocation. It is minimally complete rather than fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and each parameter, including the mutually exclusive file_path/file_content/payload routes, is documented in the schema itself. The description adds nothing about parameters, which is acceptable given the schema does the work, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Upload file', which simply restates the tool name and title verbatim, so it does not independently convey purpose. It never distinguishes this single-file upload from close siblings like upload_files, upload_session_file, or upload_brain_files. The only added content is operational caution, not a clearer statement of what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives real usage guidance: explicit confirmation is required, the operation can spend credits or trigger downstream actions, and unknown outcomes must not be auto-resubmitted. However, it offers no when-to-use versus alternatives, and with upload_files and several scoped upload tools present, the agent gets no help choosing among them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_filesUpload multiple filesADestructive
Upload multiple files Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| files | No | ||
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| user_id | No | The user ID associated with the files. Required if project_id is not provided. | |
| project_id | No | The project ID associated with the files. Required if user_id is not provided. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, idempotentHint=false, openWorldHint=true, readOnlyHint=false. The description adds value beyond these by naming the concrete stakes ('runs can spend credits or trigger downstream actions') and the non-idempotency consequence ('never resubmit unknown outcomes automatically'), which explains WHY it is non-idempotent rather than just declaring it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The title line is repeated before the substantive guidance, which is slight redundancy, but the body is three front-loaded sentences with no filler. The key safety constraints (confirmation, no auto-resubmit) lead the content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-param, nested, destructive, non-idempotent mutation tool with no output schema, the description covers the two highest-risk behaviors (confirmation gating and non-resubmission). It does not explain return values or the payload vs. body-flags selection, but annotations plus rich schema cover most remaining ground.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 86%, so the schema already documents files, account, confirm, user_id/project_id, and payload. The description adds no parameter-specific meaning (e.g., the user_id vs project_id mutual requirement, or the payload/payload_file mutual exclusivity) beyond what the schema says. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Upload multiple files'), which clearly distinguishes it from the singular sibling upload_file and from upload_session_file/upload_brain_files. However, it doesn't clarify the target scope (e.g., account-level files vs. session/project files) beyond what the schema reveals.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a warning about resubmission ('never resubmit unknown outcomes automatically') and confirmation requirement, which implies caution-oriented usage. But it never names an alternative sibling (upload_file, upload_session_file) or states when this batch tool is preferable to the single-file variant.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_session_fileUpload session fileADestructive
Upload a file into a session's input namespace so it can be attached to a message. The response returns the stored path — pass it as file_name in the attachments array when sending a message on the same session.
Files are base64 encoded in the request body and limited to 200MB (decoded). Uploaded files are scoped to the session they were uploaded to and cannot be attached to messages on other sessions. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| file_name | No | Name of the file. Directory components are stripped; the base name is sanitized before storage. | |
| media_type | No | MIME type of the file. Echoed back in the response. | |
| session_id | Yes | ID of the session to upload the file to. | |
| file_content | No | Base64-encoded file contents. Maximum decoded size is 200MB. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, so the safety profile is covered. The description adds real context beyond that: 200MB decoded limit, base64 encoding, session scoping of stored files, mandatory confirmation, and a caution against automatic resubmission. The credits sentence is somewhat generic boilerplate, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short paragraphs, front-loaded with the action and the return-value contract. The trailing confirmation/credits sentences are somewhat generic and repeat a policy that the `confirm` parameter description already implies.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by stating what the response returns and how to use it. Limits, scoping, and confirmation are all covered; only the payload/payload_file/body-flag selection is left to the schema, which documents it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds cross-tool meaning: the returned stored path is consumed as `file_name` in the `attachments` array of send_message. That linkage is not derivable from this tool's own schema and is genuinely useful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (upload) and resource (file into a session's input namespace), and pins the purpose to a concrete downstream use (attaching to a message). An agent can distinguish it from sibling upload_file/upload_files/upload_brain_files because the session scoping and message-attachment framing are explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains when to use it (to get a path that can be attached to a message on the same session) and states the confirmation prerequisite. It does not explicitly name or exclude the sibling upload tools (upload_file, upload_files), so the routing guidance is clear but not complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
43 tool updates
v3.0.0- Changed
approve_brain_source1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
attach_agent_mcp_server1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
call_mcp_tools14 fields changed- added
Input schema / $defs / callsAdded value: +{ + "description": "Tool calls to execute. Dispatched concurrently; the batch is capped at 5.", + "items": { + "properties": { + "arguments": { + "description": "Arguments passed to the tool. Defaults to `{}`. Validated by the tool's `input_schema`.", + "type": "object" + }, + "ref": { + "description": "Caller-supplied identifier echoed back on the matching result. When omitted, Gumloop assigns the call's zero-based index in `calls` as its `ref`.", + "type": [ + "string", + "null" + ] + }, + "server_id": { + "type": "string" + }, + "tool_name": { + "type": "string" + } + }, + "required": [ + "server_id", + "tool_name" + ], + "type": "object" + }, + "maxItems": 5, + "minItems": 1, + "type": "array" +} - added
Input schema / properties / calls / $refAdded value: +"#/$defs/calls" - removed
Input schema / properties / calls / descriptionRemoved value: -"Tool calls to execute. Dispatched concurrently; the batch is capped at 5." - removed
Input schema / properties / calls / itemsRemoved value: -{ - "properties": { - "arguments": { - "description": "Arguments passed to the tool. Defaults to `{}`. Validated by the tool's `input_schema`.", - "type": "object" - }, - "ref": { - "description": "Caller-supplied identifier echoed back on the matching result. When omitted, Gumloop assigns the call's zero-based index in `calls` as its `ref`.", - "type": [ - "string", - "null" - ] - }, - "server_id": { - "type": "string" - }, - "tool_name": { - "type": "string" - } - }, - "required": [ - "server_id", - "tool_name" - ], - "type": "object" -} - removed
Input schema / properties / calls / maxItemsRemoved value: -5 - removed
Input schema / properties / calls / minItemsRemoved value: -1 - removed
Input schema / properties / calls / typeRemoved value: -"array" - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / payload / properties / calls / $refAdded value: +"#/$defs/calls" - removed
Input schema / properties / payload / properties / calls / descriptionRemoved value: -"Tool calls to execute. Dispatched concurrently; the batch is capped at 5." - removed
Input schema / properties / payload / properties / calls / itemsRemoved value: -{ - "properties": { - "arguments": { - "description": "Arguments passed to the tool. Defaults to `{}`. Validated by the tool's `input_schema`.", - "type": "object" - }, - "ref": { - "description": "Caller-supplied identifier echoed back on the matching result. When omitted, Gumloop assigns the call's zero-based index in `calls` as its `ref`.", - "type": [ - "string", - "null" - ] - }, - "server_id": { - "type": "string" - }, - "tool_name": { - "type": "string" - } - }, - "required": [ - "server_id", - "tool_name" - ], - "type": "object" -} - removed
Input schema / properties / payload / properties / calls / maxItemsRemoved value: -5 - removed
Input schema / properties / payload / properties / calls / minItemsRemoved value: -1 - removed
Input schema / properties / payload / properties / calls / typeRemoved value: -"array"
- Changed
cancel_session1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
create_agent19 fields changed- added
Input schema / $defs / is_activeAdded value: +{ + "default": true, + "description": "Whether the agent is active. Defaults to `true`.\n\nSetting this to `false` retires the agent: it disappears from `GET /agents`, and `GET`/`PATCH /agents/{agent_id}` return `404`, so it cannot be reactivated through the API. This is not a pause switch — to stop an agent from running while keeping it reachable, disable its triggers instead.\n", + "type": "boolean" +} - added
Input schema / $defs / skill_idsAdded value: +{ + "description": "IDs of skills to attach to the agent. Attachment happens inside the create transaction, so an invalid ID fails the whole request (no orphaned agent). Omit to attach none. The caller must hold `INVOKE` on each skill. After creation, manage skills with `PATCH /agents/{agent_id}/skills`.\n", + "items": { + "type": "string" + }, + "type": [ + "array", + "null" + ] +} - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / is_active / $refAdded value: +"#/$defs/is_active" - removed
Input schema / properties / is_active / defaultRemoved value: -true - removed
Input schema / properties / is_active / descriptionRemoved value: -"Whether the agent is active. Defaults to `true`.\n\nSetting this to `false` retires the agent: it disappears from `GET /agents`, and `GET`/`PATCH /agents/{agent_id}` return `404`, so it cannot be reactivated through the API. This is not a pause switch — to stop an agent from running while keeping it reachable, disable its triggers instead.\n" - removed
Input schema / properties / is_active / typeRemoved value: -"boolean" - added
Input schema / properties / payload / properties / is_active / $refAdded value: +"#/$defs/is_active" - removed
Input schema / properties / payload / properties / is_active / defaultRemoved value: -true - removed
Input schema / properties / payload / properties / is_active / descriptionRemoved value: -"Whether the agent is active. Defaults to `true`.\n\nSetting this to `false` retires the agent: it disappears from `GET /agents`, and `GET`/`PATCH /agents/{agent_id}` return `404`, so it cannot be reactivated through the API. This is not a pause switch — to stop an agent from running while keeping it reachable, disable its triggers instead.\n" - removed
Input schema / properties / payload / properties / is_active / typeRemoved value: -"boolean" - added
Input schema / properties / payload / properties / skill_ids / $refAdded value: +"#/$defs/skill_ids" - removed
Input schema / properties / payload / properties / skill_ids / descriptionRemoved value: -"IDs of skills to attach to the agent. Attachment happens inside the create transaction, so an invalid ID fails the whole request (no orphaned agent). Omit to attach none. The caller must hold `INVOKE` on each skill. After creation, manage skills with `PATCH /agents/{agent_id}/skills`.\n" - removed
Input schema / properties / payload / properties / skill_ids / itemsRemoved value: -{ - "type": "string" -} - removed
Input schema / properties / payload / properties / skill_ids / typeRemoved value: -[ - "array", - "null" -] - added
Input schema / properties / skill_ids / $refAdded value: +"#/$defs/skill_ids" - removed
Input schema / properties / skill_ids / descriptionRemoved value: -"IDs of skills to attach to the agent. Attachment happens inside the create transaction, so an invalid ID fails the whole request (no orphaned agent). Omit to attach none. The caller must hold `INVOKE` on each skill. After creation, manage skills with `PATCH /agents/{agent_id}/skills`.\n" - removed
Input schema / properties / skill_ids / itemsRemoved value: -{ - "type": "string" -} - removed
Input schema / properties / skill_ids / typeRemoved value: -[ - "array", - "null" -]
- Changed
create_brain_source1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
create_chat_completion31 fields changed- added
Input schema / $defs / image_configAdded value: +{ + "description": "Image-generation parameters (size, quality, aspect_ratio, background, output_format, partial_images). Optional. Image-generation models accept either `modalities: [\"image\"]` or `image_config` (or both); chat models ignore this field.", + "type": "object" +} - added
Input schema / $defs / messagesAdded value: +{ + "description": "Conversation history. Roles `system`, `developer`, `user`, `assistant`, and `tool`. User messages accept multipart `content` with `text` and `image_url` parts. Assistant messages can carry `tool_calls`; answer each one with a `tool` message whose `tool_call_id` matches the call's `id`.", + "items": { + "type": "object" + }, + "type": "array" +} - added
Input schema / $defs / response_formatAdded value: +{ + "description": "Constrain the response. `{type: \"json_object\"}` returns a JSON object; `{type: \"json_schema\", json_schema: {name, strict, schema}}` returns JSON matching the supplied schema.\n", + "type": "object" +} - added
Input schema / $defs / tool_choiceAdded value: +{ + "description": "`\"auto\"` lets the model choose and is the default when `tools` are sent. `\"none\"` disables tool calls, `\"required\"` forces a tool call, and `{\"type\": \"function\", \"function\": {\"name\": \"...\"}}` forces a specific tool.", + "oneOf": [ + { + "type": "string" + }, + { + "type": "object" + } + ] +} - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / image_config / $refAdded value: +"#/$defs/image_config" - removed
Input schema / properties / image_config / descriptionRemoved value: -"Image-generation parameters (size, quality, aspect_ratio, background, output_format, partial_images). Optional. Image-generation models accept either `modalities: [\"image\"]` or `image_config` (or both); chat models ignore this field." - removed
Input schema / properties / image_config / typeRemoved value: -"object" - added
Input schema / properties / messages / $refAdded value: +"#/$defs/messages" - removed
Input schema / properties / messages / descriptionRemoved value: -"Conversation history. Roles `system`, `developer`, `user`, `assistant`, and `tool`. User messages accept multipart `content` with `text` and `image_url` parts. Assistant messages can carry `tool_calls`; answer each one with a `tool` message whose `tool_call_id` matches the call's `id`." - removed
Input schema / properties / messages / itemsRemoved value: -{ - "type": "object" -} - removed
Input schema / properties / messages / typeRemoved value: -"array" - added
Input schema / properties / payload / properties / image_config / $refAdded value: +"#/$defs/image_config" - removed
Input schema / properties / payload / properties / image_config / descriptionRemoved value: -"Image-generation parameters (size, quality, aspect_ratio, background, output_format, partial_images). Optional. Image-generation models accept either `modalities: [\"image\"]` or `image_config` (or both); chat models ignore this field." - removed
Input schema / properties / payload / properties / image_config / typeRemoved value: -"object" - added
Input schema / properties / payload / properties / messages / $refAdded value: +"#/$defs/messages" - removed
Input schema / properties / payload / properties / messages / descriptionRemoved value: -"Conversation history. Roles `system`, `developer`, `user`, `assistant`, and `tool`. User messages accept multipart `content` with `text` and `image_url` parts. Assistant messages can carry `tool_calls`; answer each one with a `tool` message whose `tool_call_id` matches the call's `id`." - removed
Input schema / properties / payload / properties / messages / itemsRemoved value: -{ - "type": "object" -} - removed
Input schema / properties / payload / properties / messages / typeRemoved value: -"array" - added
Input schema / properties / payload / properties / response_format / $refAdded value: +"#/$defs/response_format" - removed
Input schema / properties / payload / properties / response_format / descriptionRemoved value: -"Constrain the response. `{type: \"json_object\"}` returns a JSON object; `{type: \"json_schema\", json_schema: {name, strict, schema}}` returns JSON matching the supplied schema.\n" - removed
Input schema / properties / payload / properties / response_format / typeRemoved value: -"object" - added
Input schema / properties / payload / properties / tool_choice / $refAdded value: +"#/$defs/tool_choice" - removed
Input schema / properties / payload / properties / tool_choice / descriptionRemoved value: -"`\"auto\"` lets the model choose and is the default when `tools` are sent. `\"none\"` disables tool calls, `\"required\"` forces a tool call, and `{\"type\": \"function\", \"function\": {\"name\": \"...\"}}` forces a specific tool." - removed
Input schema / properties / payload / properties / tool_choice / oneOfRemoved value: -[ - { - "type": "string" - }, - { - "type": "object" - } -] - added
Input schema / properties / response_format / $refAdded value: +"#/$defs/response_format" - removed
Input schema / properties / response_format / descriptionRemoved value: -"Constrain the response. `{type: \"json_object\"}` returns a JSON object; `{type: \"json_schema\", json_schema: {name, strict, schema}}` returns JSON matching the supplied schema.\n" - removed
Input schema / properties / response_format / typeRemoved value: -"object" - added
Input schema / properties / tool_choice / $refAdded value: +"#/$defs/tool_choice" - removed
Input schema / properties / tool_choice / descriptionRemoved value: -"`\"auto\"` lets the model choose and is the default when `tools` are sent. `\"none\"` disables tool calls, `\"required\"` forces a tool call, and `{\"type\": \"function\", \"function\": {\"name\": \"...\"}}` forces a specific tool." - removed
Input schema / properties / tool_choice / oneOfRemoved value: -[ - { - "type": "string" - }, - { - "type": "object" - } -]
- Changed
create_organization_evaluation1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
create_session10 fields changed- added
Input schema / $defs / nameAdded value: +{ + "description": "Optional display name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so an empty or whitespace-only value is rejected. A name you supply is kept; the automatic title only fills in sessions created without one. Rename the session later with `PATCH /sessions/{session_id}`.", + "maxLength": 256, + "type": "string" +} - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / name / $refAdded value: +"#/$defs/name" - removed
Input schema / properties / name / descriptionRemoved value: -"Optional display name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so an empty or whitespace-only value is rejected. A name you supply is kept; the automatic title only fills in sessions created without one. Rename the session later with `PATCH /sessions/{session_id}`." - removed
Input schema / properties / name / maxLengthRemoved value: -256 - removed
Input schema / properties / name / typeRemoved value: -"string" - added
Input schema / properties / payload / properties / name / $refAdded value: +"#/$defs/name" - removed
Input schema / properties / payload / properties / name / descriptionRemoved value: -"Optional display name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so an empty or whitespace-only value is rejected. A name you supply is kept; the automatic title only fills in sessions created without one. Rename the session later with `PATCH /sessions/{session_id}`." - removed
Input schema / properties / payload / properties / name / maxLengthRemoved value: -256 - removed
Input schema / properties / payload / properties / name / typeRemoved value: -"string"
- Changed
create_skill14 fields changed- added
Input schema / $defs / filesAdded value: +{ + "description": "Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks.", + "items": { + "minLength": 1, + "type": "string" + }, + "maxItems": 25, + "minItems": 1, + "type": "array" +} - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / files / $refAdded value: +"#/$defs/files" - removed
Input schema / properties / files / descriptionRemoved value: -"Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks." - removed
Input schema / properties / files / itemsRemoved value: -{ - "minLength": 1, - "type": "string" -} - removed
Input schema / properties / files / maxItemsRemoved value: -25 - removed
Input schema / properties / files / minItemsRemoved value: -1 - removed
Input schema / properties / files / typeRemoved value: -"array" - added
Input schema / properties / payload / properties / files / $refAdded value: +"#/$defs/files" - removed
Input schema / properties / payload / properties / files / descriptionRemoved value: -"Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks." - removed
Input schema / properties / payload / properties / files / itemsRemoved value: -{ - "minLength": 1, - "type": "string" -} - removed
Input schema / properties / payload / properties / files / maxItemsRemoved value: -25 - removed
Input schema / properties / payload / properties / files / minItemsRemoved value: -1 - removed
Input schema / properties / payload / properties / files / typeRemoved value: -"array"
- Changed
delete_brain_file1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
delete_brain_source1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
delete_organization_evaluation1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
delete_queued_message1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
delete_skill1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
detach_agent_mcp_server1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
export_data75 fields changed- added
Input schema / $defs / category_filterAdded value: +{ + "description": "**Only applicable when `data_type` is `\"credit_logs\"`.** Optional filter to export only credit logs matching a specific category (e.g., `\"PIPELINE_RUN\"`, `\"AGENT_RUN\"`, `\"CUSTOM_NODE_RD\"`, `\"EXTERNAL_GUMCP_CALL\"`). When omitted, all categories are included.\n", + "type": "string" +} - added
Input schema / $defs / data_typeAdded value: +{ + "default": "workflows", + "description": "The type of data to export. Use `\"workflows\"` to export workflow run data, `\"agents\"` to export agent configuration data, `\"agent_interactions\"` to export agent interaction data, `\"credit_logs\"` to export credit transaction history, `\"interaction_evaluations\"` to export completed chat evaluations, or `\"gumstack\"` to export Gumstack tool call activity. Defaults to `\"workflows\"`. Audit-log exports are available in the Gumloop app only, not through this endpoint.\n", + "enum": [ + "workflows", + "agents", + "agent_interactions", + "credit_logs", + "interaction_evaluations", + "gumstack" + ], + "type": "string" +} - added
Input schema / $defs / entity_idsAdded value: +{ + "description": "An optional array of specific entity IDs to filter the export. For workflow exports (`data_type: \"workflows\"`), these are workbook IDs. For agent exports (`data_type: \"agents\"`) and agent interaction exports (`data_type: \"agent_interactions\"`), these are agent IDs. When provided, only data for the specified entities will be included. **Not applicable when `data_type` is `\"credit_logs\"`.**\n", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / $defs / export_fieldsAdded value: +{ + "description": "Array of fields to include in the export. The available fields depend on the `data_type`.\n\n**Workflow fields** (when `data_type` is `\"workflows\"`):\n- `workbook_id` - The workbook identifier\n- `workbook_name` - The workbook name\n- `workbook_created_ts` - Workbook creation timestamp\n- `user_id` - The user identifier\n- `user_email` - The user's email address\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n- `run_id` - The flow run identifier\n- `credit_cost` - Credits consumed by the run\n- `pl_run_created_ts` - Flow run creation timestamp\n- `pl_run_finished_ts` - Flow run completion timestamp\n- `pipeline` - The full pipeline configuration (JSON)\n\n**Agent fields** (when `data_type` is `\"agents\"`):\n- `agent_id` - The agent identifier\n- `agent_name` - The agent name\n- `agent_description` - The agent description\n- `agent_model` - The model used by the agent\n- `agent_system_prompt` - The agent's system prompt\n- `agent_created_ts` - Agent creation timestamp\n- `agent_tools` - Tools configured for the agent (JSON)\n- `agent_metadata` - Additional agent metadata (JSON)\n- `agent_evaluations_enabled` - Whether the agent's own Evaluations are turned on (`true`/`false`)\n- `creator_email` - Email of the user who created the agent\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n- `folder_id` - The identifier of the folder containing the agent (empty when the agent is not in a folder)\n- `folder_name` - The name of the folder containing the agent (empty when the agent is not in a folder)\n\n**Agent interaction fields** (when `data_type` is `\"agent_interactions\"`):\n- `interaction_id` - Unique identifier for the chat session\n- `agent_id` - The agent identifier\n- `agent_name` - The agent name\n- `interaction_type` - Type of interaction (e.g., chat, slack, api, triggered)\n- `interaction_name` - Display name of the chat session\n- `trigger_type` - For triggered interactions, the specific trigger type (e.g., time_based, polling_new_record_salesforce). Null for non-triggered interactions.\n- `interaction_created_ts` - Chat session creation timestamp\n- `user_email` - Email of the user who initiated the chat\n- `credit_cost` - Total credits consumed (LLM + tool + flow)\n- `llm_credit_cost` - Credits consumed by LLM calls only\n- `tool_credit_cost` - Credits consumed by tool calls\n- `flow_credit_cost` - Credits consumed by pipeline runs\n- `message_count` - Number of messages in the conversation\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n\n**Interaction evaluation fields** (when `data_type` is `\"interaction_evaluations\"`):\n- `evaluation_id` - Unique identifier for the evaluation row\n- `interaction_id` - The chat session that was evaluated (one chat can have several evaluation rows)\n- `agent_id` - The identifier of the agent that was evaluated\n- `organization_evaluation_id` - The organization-level evaluation this row belongs to (empty for agent-level rubrics)\n- `evaluation_created_ts` - When the evaluation reached its final state\n- `status` - The evaluation's state\n- `grade` - The grade the evaluation produced\n- `call_outcome` - The outcome the evaluation recorded for the conversation\n- `sentiment` - The sentiment the evaluation recorded\n- `error_code` - Error code when the evaluation could not complete\n- `summary` - The evaluation's written summary\n- `evaluation_model` - The model that ran the evaluation\n- `credit_cost` - Credits consumed by the evaluation\n- `user_email` - Email of the user whose chat was evaluated\n\n**Credit log fields** (when `data_type` is `\"credit_logs\"`):\n- `user_email` - Email of the user associated with the credit log entry\n- `permission_group_id` - Custom role ID(s) the user belongs to (semicolon-separated if multiple) (disabled by default)\n- `permission_group_name` - Custom role name(s) the user belongs to (semicolon-separated if multiple) (disabled by default)\n- `timestamp` - When the credit transaction occurred\n- `category` - The category of the credit log (e.g., PIPELINE_RUN, AGENT_RUN)\n- `type` - The specific type of credit charge\n- `name` - Display name of the credit log entry\n- `amount` - Number of credits charged or adjusted\n- `balance` - Credit balance after the transaction\n- `log_id` - Unique identifier for the credit log entry\n- `correlation_id` - Join key to the related run or interaction (disabled by default)\n- `balance_scope` - Whether the balance is organization- or user-scoped (disabled by default)\n- `project_id` - The project identifier (disabled by default)\n\nNot all combinations of selected fields are guaranteed to produce data for every row.\n", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / $defs / export_levelAdded value: +{ + "default": "organization", + "description": "The scope of the export. Use `\"organization\"` to export across the entire organization, or `\"workspace\"` to export from a single workspace (requires exactly one ID in `workspace_ids`). Defaults to `\"organization\"`. **Not applicable when `data_type` is `\"credit_logs\"`.**\n", + "enum": [ + "workspace", + "organization" + ], + "type": "string" +} - added
Input schema / $defs / include_all_workspacesAdded value: +{ + "default": false, + "description": "Whether to include all workspaces in the organization. When `true`, also sets `include_personal_workspaces` to `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**", + "type": "boolean" +} - added
Input schema / $defs / include_personal_workspacesAdded value: +{ + "default": false, + "description": "Whether to include personal workspaces in the export. Ignored if `include_all_workspaces` is `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**", + "type": "boolean" +} - added
Input schema / $defs / workspace_idsAdded value: +{ + "description": "An optional array of workspace IDs to include in the export. When `export_level` is `\"workspace\"`, exactly one workspace ID is required. Ignored if `include_all_workspaces` is `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / category_filter / $refAdded value: +"#/$defs/category_filter" - removed
Input schema / properties / category_filter / descriptionRemoved value: -"**Only applicable when `data_type` is `\"credit_logs\"`.** Optional filter to export only credit logs matching a specific category (e.g., `\"PIPELINE_RUN\"`, `\"AGENT_RUN\"`, `\"CUSTOM_NODE_RD\"`, `\"EXTERNAL_GUMCP_CALL\"`). When omitted, all categories are included.\n" - removed
Input schema / properties / category_filter / typeRemoved value: -"string" - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / data_type / $refAdded value: +"#/$defs/data_type" - removed
Input schema / properties / data_type / defaultRemoved value: -"workflows" - removed
Input schema / properties / data_type / descriptionRemoved value: -"The type of data to export. Use `\"workflows\"` to export workflow run data, `\"agents\"` to export agent configuration data, `\"agent_interactions\"` to export agent interaction data, `\"credit_logs\"` to export credit transaction history, `\"interaction_evaluations\"` to export completed chat evaluations, or `\"gumstack\"` to export Gumstack tool call activity. Defaults to `\"workflows\"`. Audit-log exports are available in the Gumloop app only, not through this endpoint.\n" - removed
Input schema / properties / data_type / enumRemoved value: -[ - "workflows", - "agents", - "agent_interactions", - "credit_logs", - "interaction_evaluations", - "gumstack" -] - removed
Input schema / properties / data_type / typeRemoved value: -"string" - added
Input schema / properties / entity_ids / $refAdded value: +"#/$defs/entity_ids" - removed
Input schema / properties / entity_ids / descriptionRemoved value: -"An optional array of specific entity IDs to filter the export. For workflow exports (`data_type: \"workflows\"`), these are workbook IDs. For agent exports (`data_type: \"agents\"`) and agent interaction exports (`data_type: \"agent_interactions\"`), these are agent IDs. When provided, only data for the specified entities will be included. **Not applicable when `data_type` is `\"credit_logs\"`.**\n" - removed
Input schema / properties / entity_ids / itemsRemoved value: -{ - "type": "string" -} - removed
Input schema / properties / entity_ids / typeRemoved value: -"array" - added
Input schema / properties / export_fields / $refAdded value: +"#/$defs/export_fields" - removed
Input schema / properties / export_fields / descriptionRemoved value: -"Array of fields to include in the export. The available fields depend on the `data_type`.\n\n**Workflow fields** (when `data_type` is `\"workflows\"`):\n- `workbook_id` - The workbook identifier\n- `workbook_name` - The workbook name\n- `workbook_created_ts` - Workbook creation timestamp\n- `user_id` - The user identifier\n- `user_email` - The user's email address\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n- `run_id` - The flow run identifier\n- `credit_cost` - Credits consumed by the run\n- `pl_run_created_ts` - Flow run creation timestamp\n- `pl_run_finished_ts` - Flow run completion timestamp\n- `pipeline` - The full pipeline configuration (JSON)\n\n**Agent fields** (when `data_type` is `\"agents\"`):\n- `agent_id` - The agent identifier\n- `agent_name` - The agent name\n- `agent_description` - The agent description\n- `agent_model` - The model used by the agent\n- `agent_system_prompt` - The agent's system prompt\n- `agent_created_ts` - Agent creation timestamp\n- `agent_tools` - Tools configured for the agent (JSON)\n- `agent_metadata` - Additional agent metadata (JSON)\n- `agent_evaluations_enabled` - Whether the agent's own Evaluations are turned on (`true`/`false`)\n- `creator_email` - Email of the user who created the agent\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n- `folder_id` - The identifier of the folder containing the agent (empty when the agent is not in a folder)\n- `folder_name` - The name of the folder containing the agent (empty when the agent is not in a folder)\n\n**Agent interaction fields** (when `data_type` is `\"agent_interactions\"`):\n- `interaction_id` - Unique identifier for the chat session\n- `agent_id` - The agent identifier\n- `agent_name` - The agent name\n- `interaction_type` - Type of interaction (e.g., chat, slack, api, triggered)\n- `interaction_name` - Display name of the chat session\n- `trigger_type` - For triggered interactions, the specific trigger type (e.g., time_based, polling_new_record_salesforce). Null for non-triggered interactions.\n- `interaction_created_ts` - Chat session creation timestamp\n- `user_email` - Email of the user who initiated the chat\n- `credit_cost` - Total credits consumed (LLM + tool + flow)\n- `llm_credit_cost` - Credits consumed by LLM calls only\n- `tool_credit_cost` - Credits consumed by tool calls\n- `flow_credit_cost` - Credits consumed by pipeline runs\n- `message_count` - Number of messages in the conversation\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n\n**Interaction evaluation fields** (when `data_type` is `\"interaction_evaluations\"`):\n- `evaluation_id` - Unique identifier for the evaluation row\n- `interaction_id` - The chat session that was evaluated (one chat can have several evaluation rows)\n- `agent_id` - The identifier of the agent that was evaluated\n- `organization_evaluation_id` - The organization-level evaluation this row belongs to (empty for agent-level rubrics)\n- `evaluation_created_ts` - When the evaluation reached its final state\n- `status` - The evaluation's state\n- `grade` - The grade the evaluation produced\n- `call_outcome` - The outcome the evaluation recorded for the conversation\n- `sentiment` - The sentiment the evaluation recorded\n- `error_code` - Error code when the evaluation could not complete\n- `summary` - The evaluation's written summary\n- `evaluation_model` - The model that ran the evaluation\n- `credit_cost` - Credits consumed by the evaluation\n- `user_email` - Email of the user whose chat was evaluated\n\n**Credit log fields** (when `data_type` is `\"credit_logs\"`):\n- `user_email` - Email of the user associated with the credit log entry\n- `permission_group_id` - Custom role ID(s) the user belongs to (semicolon-separated if multiple) (disabled by default)\n- `permission_group_name` - Custom role name(s) the user belongs to (semicolon-separated if multiple) (disabled by default)\n- `timestamp` - When the credit transaction occurred\n- `category` - The category of the credit log (e.g., PIPELINE_RUN, AGENT_RUN)\n- `type` - The specific type of credit charge\n- `name` - Display name of the credit log entry\n- `amount` - Number of credits charged or adjusted\n- `balance` - Credit balance after the transaction\n- `log_id` - Unique identifier for the credit log entry\n- `correlation_id` - Join key to the related run or interaction (disabled by default)\n- `balance_scope` - Whether the balance is organization- or user-scoped (disabled by default)\n- `project_id` - The project identifier (disabled by default)\n\nNot all combinations of selected fields are guaranteed to produce data for every row.\n" - removed
Input schema / properties / export_fields / itemsRemoved value: -{ - "type": "string" -} - removed
Input schema / properties / export_fields / typeRemoved value: -"array" - added
Input schema / properties / export_level / $refAdded value: +"#/$defs/export_level" - removed
Input schema / properties / export_level / defaultRemoved value: -"organization" - removed
Input schema / properties / export_level / descriptionRemoved value: -"The scope of the export. Use `\"organization\"` to export across the entire organization, or `\"workspace\"` to export from a single workspace (requires exactly one ID in `workspace_ids`). Defaults to `\"organization\"`. **Not applicable when `data_type` is `\"credit_logs\"`.**\n" - removed
Input schema / properties / export_level / enumRemoved value: -[ - "workspace", - "organization" -] - removed
Input schema / properties / export_level / typeRemoved value: -"string" - added
Input schema / properties / include_all_workspaces / $refAdded value: +"#/$defs/include_all_workspaces" - removed
Input schema / properties / include_all_workspaces / defaultRemoved value: -false - removed
Input schema / properties / include_all_workspaces / descriptionRemoved value: -"Whether to include all workspaces in the organization. When `true`, also sets `include_personal_workspaces` to `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**" - removed
Input schema / properties / include_all_workspaces / typeRemoved value: -"boolean" - added
Input schema / properties / include_personal_workspaces / $refAdded value: +"#/$defs/include_personal_workspaces" - removed
Input schema / properties / include_personal_workspaces / defaultRemoved value: -false - removed
Input schema / properties / include_personal_workspaces / descriptionRemoved value: -"Whether to include personal workspaces in the export. Ignored if `include_all_workspaces` is `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**" - removed
Input schema / properties / include_personal_workspaces / typeRemoved value: -"boolean" - added
Input schema / properties / payload / properties / category_filter / $refAdded value: +"#/$defs/category_filter" - removed
Input schema / properties / payload / properties / category_filter / descriptionRemoved value: -"**Only applicable when `data_type` is `\"credit_logs\"`.** Optional filter to export only credit logs matching a specific category (e.g., `\"PIPELINE_RUN\"`, `\"AGENT_RUN\"`, `\"CUSTOM_NODE_RD\"`, `\"EXTERNAL_GUMCP_CALL\"`). When omitted, all categories are included.\n" - removed
Input schema / properties / payload / properties / category_filter / typeRemoved value: -"string" - added
Input schema / properties / payload / properties / data_type / $refAdded value: +"#/$defs/data_type" - removed
Input schema / properties / payload / properties / data_type / defaultRemoved value: -"workflows" - removed
Input schema / properties / payload / properties / data_type / descriptionRemoved value: -"The type of data to export. Use `\"workflows\"` to export workflow run data, `\"agents\"` to export agent configuration data, `\"agent_interactions\"` to export agent interaction data, `\"credit_logs\"` to export credit transaction history, `\"interaction_evaluations\"` to export completed chat evaluations, or `\"gumstack\"` to export Gumstack tool call activity. Defaults to `\"workflows\"`. Audit-log exports are available in the Gumloop app only, not through this endpoint.\n" - removed
Input schema / properties / payload / properties / data_type / enumRemoved value: -[ - "workflows", - "agents", - "agent_interactions", - "credit_logs", - "interaction_evaluations", - "gumstack" -] - removed
Input schema / properties / payload / properties / data_type / typeRemoved value: -"string" - added
Input schema / properties / payload / properties / entity_ids / $refAdded value: +"#/$defs/entity_ids" - removed
Input schema / properties / payload / properties / entity_ids / descriptionRemoved value: -"An optional array of specific entity IDs to filter the export. For workflow exports (`data_type: \"workflows\"`), these are workbook IDs. For agent exports (`data_type: \"agents\"`) and agent interaction exports (`data_type: \"agent_interactions\"`), these are agent IDs. When provided, only data for the specified entities will be included. **Not applicable when `data_type` is `\"credit_logs\"`.**\n" - removed
Input schema / properties / payload / properties / entity_ids / itemsRemoved value: -{ - "type": "string" -} - removed
Input schema / properties / payload / properties / entity_ids / typeRemoved value: -"array" - added
Input schema / properties / payload / properties / export_fields / $refAdded value: +"#/$defs/export_fields" - removed
Input schema / properties / payload / properties / export_fields / descriptionRemoved value: -"Array of fields to include in the export. The available fields depend on the `data_type`.\n\n**Workflow fields** (when `data_type` is `\"workflows\"`):\n- `workbook_id` - The workbook identifier\n- `workbook_name` - The workbook name\n- `workbook_created_ts` - Workbook creation timestamp\n- `user_id` - The user identifier\n- `user_email` - The user's email address\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n- `run_id` - The flow run identifier\n- `credit_cost` - Credits consumed by the run\n- `pl_run_created_ts` - Flow run creation timestamp\n- `pl_run_finished_ts` - Flow run completion timestamp\n- `pipeline` - The full pipeline configuration (JSON)\n\n**Agent fields** (when `data_type` is `\"agents\"`):\n- `agent_id` - The agent identifier\n- `agent_name` - The agent name\n- `agent_description` - The agent description\n- `agent_model` - The model used by the agent\n- `agent_system_prompt` - The agent's system prompt\n- `agent_created_ts` - Agent creation timestamp\n- `agent_tools` - Tools configured for the agent (JSON)\n- `agent_metadata` - Additional agent metadata (JSON)\n- `agent_evaluations_enabled` - Whether the agent's own Evaluations are turned on (`true`/`false`)\n- `creator_email` - Email of the user who created the agent\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n- `folder_id` - The identifier of the folder containing the agent (empty when the agent is not in a folder)\n- `folder_name` - The name of the folder containing the agent (empty when the agent is not in a folder)\n\n**Agent interaction fields** (when `data_type` is `\"agent_interactions\"`):\n- `interaction_id` - Unique identifier for the chat session\n- `agent_id` - The agent identifier\n- `agent_name` - The agent name\n- `interaction_type` - Type of interaction (e.g., chat, slack, api, triggered)\n- `interaction_name` - Display name of the chat session\n- `trigger_type` - For triggered interactions, the specific trigger type (e.g., time_based, polling_new_record_salesforce). Null for non-triggered interactions.\n- `interaction_created_ts` - Chat session creation timestamp\n- `user_email` - Email of the user who initiated the chat\n- `credit_cost` - Total credits consumed (LLM + tool + flow)\n- `llm_credit_cost` - Credits consumed by LLM calls only\n- `tool_credit_cost` - Credits consumed by tool calls\n- `flow_credit_cost` - Credits consumed by pipeline runs\n- `message_count` - Number of messages in the conversation\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n\n**Interaction evaluation fields** (when `data_type` is `\"interaction_evaluations\"`):\n- `evaluation_id` - Unique identifier for the evaluation row\n- `interaction_id` - The chat session that was evaluated (one chat can have several evaluation rows)\n- `agent_id` - The identifier of the agent that was evaluated\n- `organization_evaluation_id` - The organization-level evaluation this row belongs to (empty for agent-level rubrics)\n- `evaluation_created_ts` - When the evaluation reached its final state\n- `status` - The evaluation's state\n- `grade` - The grade the evaluation produced\n- `call_outcome` - The outcome the evaluation recorded for the conversation\n- `sentiment` - The sentiment the evaluation recorded\n- `error_code` - Error code when the evaluation could not complete\n- `summary` - The evaluation's written summary\n- `evaluation_model` - The model that ran the evaluation\n- `credit_cost` - Credits consumed by the evaluation\n- `user_email` - Email of the user whose chat was evaluated\n\n**Credit log fields** (when `data_type` is `\"credit_logs\"`):\n- `user_email` - Email of the user associated with the credit log entry\n- `permission_group_id` - Custom role ID(s) the user belongs to (semicolon-separated if multiple) (disabled by default)\n- `permission_group_name` - Custom role name(s) the user belongs to (semicolon-separated if multiple) (disabled by default)\n- `timestamp` - When the credit transaction occurred\n- `category` - The category of the credit log (e.g., PIPELINE_RUN, AGENT_RUN)\n- `type` - The specific type of credit charge\n- `name` - Display name of the credit log entry\n- `amount` - Number of credits charged or adjusted\n- `balance` - Credit balance after the transaction\n- `log_id` - Unique identifier for the credit log entry\n- `correlation_id` - Join key to the related run or interaction (disabled by default)\n- `balance_scope` - Whether the balance is organization- or user-scoped (disabled by default)\n- `project_id` - The project identifier (disabled by default)\n\nNot all combinations of selected fields are guaranteed to produce data for every row.\n" - removed
Input schema / properties / payload / properties / export_fields / itemsRemoved value: -{ - "type": "string" -} - removed
Input schema / properties / payload / properties / export_fields / typeRemoved value: -"array" - added
Input schema / properties / payload / properties / export_level / $refAdded value: +"#/$defs/export_level" - removed
Input schema / properties / payload / properties / export_level / defaultRemoved value: -"organization" - removed
Input schema / properties / payload / properties / export_level / descriptionRemoved value: -"The scope of the export. Use `\"organization\"` to export across the entire organization, or `\"workspace\"` to export from a single workspace (requires exactly one ID in `workspace_ids`). Defaults to `\"organization\"`. **Not applicable when `data_type` is `\"credit_logs\"`.**\n" - removed
Input schema / properties / payload / properties / export_level / enumRemoved value: -[ - "workspace", - "organization" -] - removed
Input schema / properties / payload / properties / export_level / typeRemoved value: -"string" - added
Input schema / properties / payload / properties / include_all_workspaces / $refAdded value: +"#/$defs/include_all_workspaces" - removed
Input schema / properties / payload / properties / include_all_workspaces / defaultRemoved value: -false - removed
Input schema / properties / payload / properties / include_all_workspaces / descriptionRemoved value: -"Whether to include all workspaces in the organization. When `true`, also sets `include_personal_workspaces` to `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**" - removed
Input schema / properties / payload / properties / include_all_workspaces / typeRemoved value: -"boolean" - added
Input schema / properties / payload / properties / include_personal_workspaces / $refAdded value: +"#/$defs/include_personal_workspaces" - removed
Input schema / properties / payload / properties / include_personal_workspaces / defaultRemoved value: -false - removed
Input schema / properties / payload / properties / include_personal_workspaces / descriptionRemoved value: -"Whether to include personal workspaces in the export. Ignored if `include_all_workspaces` is `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**" - removed
Input schema / properties / payload / properties / include_personal_workspaces / typeRemoved value: -"boolean" - added
Input schema / properties / payload / properties / workspace_ids / $refAdded value: +"#/$defs/workspace_ids" - removed
Input schema / properties / payload / properties / workspace_ids / descriptionRemoved value: -"An optional array of workspace IDs to include in the export. When `export_level` is `\"workspace\"`, exactly one workspace ID is required. Ignored if `include_all_workspaces` is `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**" - removed
Input schema / properties / payload / properties / workspace_ids / itemsRemoved value: -{ - "type": "string" -} - removed
Input schema / properties / payload / properties / workspace_ids / typeRemoved value: -"array" - added
Input schema / properties / workspace_ids / $refAdded value: +"#/$defs/workspace_ids" - removed
Input schema / properties / workspace_ids / descriptionRemoved value: -"An optional array of workspace IDs to include in the export. When `export_level` is `\"workspace\"`, exactly one workspace ID is required. Ignored if `include_all_workspaces` is `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**" - removed
Input schema / properties / workspace_ids / itemsRemoved value: -{ - "type": "string" -} - removed
Input schema / properties / workspace_ids / typeRemoved value: -"array"
- Changed
import_browser_profile_cookies1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
kill_flow1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
manage_permission_group_users1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
manage_project_users1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
queue_session_message1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
rename_session12 fields changed- added
Input schema / $defs / nameAdded value: +{ + "description": "New name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so a whitespace-only value is rejected.", + "maxLength": 256, + "minLength": 1, + "type": "string" +} - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / name / $refAdded value: +"#/$defs/name" - removed
Input schema / properties / name / descriptionRemoved value: -"New name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so a whitespace-only value is rejected." - removed
Input schema / properties / name / maxLengthRemoved value: -256 - removed
Input schema / properties / name / minLengthRemoved value: -1 - removed
Input schema / properties / name / typeRemoved value: -"string" - added
Input schema / properties / payload / properties / name / $refAdded value: +"#/$defs/name" - removed
Input schema / properties / payload / properties / name / descriptionRemoved value: -"New name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so a whitespace-only value is rejected." - removed
Input schema / properties / payload / properties / name / maxLengthRemoved value: -256 - removed
Input schema / properties / payload / properties / name / minLengthRemoved value: -1 - removed
Input schema / properties / payload / properties / name / typeRemoved value: -"string"
- Changed
resolve_session_approvals14 fields changed- added
Input schema / $defs / approval_responsesAdded value: +{ + "description": "Answers to pending asks. Each `action_request_id` may appear at most once.", + "items": { + "properties": { + "action": { + "description": "Whether to approve or reject the pending ask.", + "enum": [ + "accept", + "reject" + ], + "type": "string" + }, + "action_request_id": { + "description": "ID of the pending ask, from `pending_approvals` on Retrieve session.", + "type": "string" + }, + "reason": { + "description": "Optional reason recorded with the resolution.", + "maxLength": 1000, + "type": "string" + }, + "response": { + "description": "For `human_input` asks — form answers keyed by question name.", + "properties": { + "values": { + "description": "Map of question name to answer.", + "type": "object" + } + }, + "type": "object" + } + }, + "required": [ + "action_request_id", + "action" + ], + "type": "object" + }, + "maxItems": 20, + "minItems": 1, + "type": "array" +} - added
Input schema / properties / approval_responses / $refAdded value: +"#/$defs/approval_responses" - removed
Input schema / properties / approval_responses / descriptionRemoved value: -"Answers to pending asks. Each `action_request_id` may appear at most once." - removed
Input schema / properties / approval_responses / itemsRemoved value: -{ - "properties": { - "action": { - "description": "Whether to approve or reject the pending ask.", - "enum": [ - "accept", - "reject" - ], - "type": "string" - }, - "action_request_id": { - "description": "ID of the pending ask, from `pending_approvals` on Retrieve session.", - "type": "string" - }, - "reason": { - "description": "Optional reason recorded with the resolution.", - "maxLength": 1000, - "type": "string" - }, - "response": { - "description": "For `human_input` asks — form answers keyed by question name.", - "properties": { - "values": { - "description": "Map of question name to answer.", - "type": "object" - } - }, - "type": "object" - } - }, - "required": [ - "action_request_id", - "action" - ], - "type": "object" -} - removed
Input schema / properties / approval_responses / maxItemsRemoved value: -20 - removed
Input schema / properties / approval_responses / minItemsRemoved value: -1 - removed
Input schema / properties / approval_responses / typeRemoved value: -"array" - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / payload / properties / approval_responses / $refAdded value: +"#/$defs/approval_responses" - removed
Input schema / properties / payload / properties / approval_responses / descriptionRemoved value: -"Answers to pending asks. Each `action_request_id` may appear at most once." - removed
Input schema / properties / payload / properties / approval_responses / itemsRemoved value: -{ - "properties": { - "action": { - "description": "Whether to approve or reject the pending ask.", - "enum": [ - "accept", - "reject" - ], - "type": "string" - }, - "action_request_id": { - "description": "ID of the pending ask, from `pending_approvals` on Retrieve session.", - "type": "string" - }, - "reason": { - "description": "Optional reason recorded with the resolution.", - "maxLength": 1000, - "type": "string" - }, - "response": { - "description": "For `human_input` asks — form answers keyed by question name.", - "properties": { - "values": { - "description": "Map of question name to answer.", - "type": "object" - } - }, - "type": "object" - } - }, - "required": [ - "action_request_id", - "action" - ], - "type": "object" -} - removed
Input schema / properties / payload / properties / approval_responses / maxItemsRemoved value: -20 - removed
Input schema / properties / payload / properties / approval_responses / minItemsRemoved value: -1 - removed
Input schema / properties / payload / properties / approval_responses / typeRemoved value: -"array"
- Changed
route_model44 fields changed- added
Input schema / $defs / agentAdded value: +{ + "additionalProperties": false, + "description": "Optional context about the agent the message is for. Sharper context produces a sharper route.", + "properties": { + "description": { + "maxLength": 2000, + "type": "string" + }, + "name": { + "maxLength": 200, + "type": "string" + }, + "system_prompt": { + "maxLength": 20000, + "type": "string" + } + }, + "type": "object" +} - added
Input schema / $defs / historyAdded value: +{ + "description": "Prior turns, oldest first, for context.", + "items": { + "additionalProperties": false, + "properties": { + "content": { + "description": "Message text, as a string or text parts. `input` and `message` are accepted as aliases.", + "oneOf": [ + { + "type": "string" + }, + { + "items": { + "type": "object" + }, + "type": "array" + } + ] + }, + "model": { + "description": "The model that produced this turn. Assistant messages only.", + "maxLength": 200, + "type": "string" + }, + "role": { + "enum": [ + "user", + "assistant" + ], + "type": "string" + } + }, + "required": [ + "role", + "content" + ], + "type": "object" + }, + "maxItems": 20, + "type": "array" +} - added
Input schema / $defs / inputAdded value: +{ + "description": "The message to route. `message` is accepted as an alias. A string or an array of `{type: text, text: ...}` parts.", + "oneOf": [ + { + "type": "string" + }, + { + "items": { + "type": "object" + }, + "type": "array" + } + ] +} - added
Input schema / $defs / modelsAdded value: +{ + "description": "Candidate model IDs to choose between. IDs must be nonempty, unique after remapping, and registered. Omit to use Chew's lane-chain union.", + "items": { + "type": "string" + }, + "maxItems": 50, + "minItems": 1, + "type": "array", + "uniqueItems": true +} - added
Input schema / properties / agent / $refAdded value: +"#/$defs/agent" - removed
Input schema / properties / agent / additionalPropertiesRemoved value: -false - removed
Input schema / properties / agent / descriptionRemoved value: -"Optional context about the agent the message is for. Sharper context produces a sharper route." - removed
Input schema / properties / agent / propertiesRemoved value: -{ - "description": { - "maxLength": 2000, - "type": "string" - }, - "name": { - "maxLength": 200, - "type": "string" - }, - "system_prompt": { - "maxLength": 20000, - "type": "string" - } -} - removed
Input schema / properties / agent / typeRemoved value: -"object" - added
Input schema / properties / history / $refAdded value: +"#/$defs/history" - removed
Input schema / properties / history / descriptionRemoved value: -"Prior turns, oldest first, for context." - removed
Input schema / properties / history / itemsRemoved value: -{ - "additionalProperties": false, - "properties": { - "content": { - "description": "Message text, as a string or text parts. `input` and `message` are accepted as aliases.", - "oneOf": [ - { - "type": "string" - }, - { - "items": { - "type": "object" - }, - "type": "array" - } - ] - }, - "model": { - "description": "The model that produced this turn. Assistant messages only.", - "maxLength": 200, - "type": "string" - }, - "role": { - "enum": [ - "user", - "assistant" - ], - "type": "string" - } - }, - "required": [ - "role", - "content" - ], - "type": "object" -} - removed
Input schema / properties / history / maxItemsRemoved value: -20 - removed
Input schema / properties / history / typeRemoved value: -"array" - added
Input schema / properties / input / $refAdded value: +"#/$defs/input" - removed
Input schema / properties / input / descriptionRemoved value: -"The message to route. `message` is accepted as an alias. A string or an array of `{type: text, text: ...}` parts." - removed
Input schema / properties / input / oneOfRemoved value: -[ - { - "type": "string" - }, - { - "items": { - "type": "object" - }, - "type": "array" - } -] - added
Input schema / properties / models / $refAdded value: +"#/$defs/models" - removed
Input schema / properties / models / descriptionRemoved value: -"Candidate model IDs to choose between. IDs must be nonempty, unique after remapping, and registered. Omit to use Chew's lane-chain union." - removed
Input schema / properties / models / itemsRemoved value: -{ - "type": "string" -} - removed
Input schema / properties / models / maxItemsRemoved value: -50 - removed
Input schema / properties / models / minItemsRemoved value: -1 - removed
Input schema / properties / models / typeRemoved value: -"array" - removed
Input schema / properties / models / uniqueItemsRemoved value: -true - added
Input schema / properties / payload / properties / agent / $refAdded value: +"#/$defs/agent" - removed
Input schema / properties / payload / properties / agent / additionalPropertiesRemoved value: -false - removed
Input schema / properties / payload / properties / agent / descriptionRemoved value: -"Optional context about the agent the message is for. Sharper context produces a sharper route." - removed
Input schema / properties / payload / properties / agent / propertiesRemoved value: -{ - "description": { - "maxLength": 2000, - "type": "string" - }, - "name": { - "maxLength": 200, - "type": "string" - }, - "system_prompt": { - "maxLength": 20000, - "type": "string" - } -} - removed
Input schema / properties / payload / properties / agent / typeRemoved value: -"object" - added
Input schema / properties / payload / properties / history / $refAdded value: +"#/$defs/history" - removed
Input schema / properties / payload / properties / history / descriptionRemoved value: -"Prior turns, oldest first, for context." - removed
Input schema / properties / payload / properties / history / itemsRemoved value: -{ - "additionalProperties": false, - "properties": { - "content": { - "description": "Message text, as a string or text parts. `input` and `message` are accepted as aliases.", - "oneOf": [ - { - "type": "string" - }, - { - "items": { - "type": "object" - }, - "type": "array" - } - ] - }, - "model": { - "description": "The model that produced this turn. Assistant messages only.", - "maxLength": 200, - "type": "string" - }, - "role": { - "enum": [ - "user", - "assistant" - ], - "type": "string" - } - }, - "required": [ - "role", - "content" - ], - "type": "object" -} - removed
Input schema / properties / payload / properties / history / maxItemsRemoved value: -20 - removed
Input schema / properties / payload / properties / history / typeRemoved value: -"array" - added
Input schema / properties / payload / properties / input / $refAdded value: +"#/$defs/input" - removed
Input schema / properties / payload / properties / input / descriptionRemoved value: -"The message to route. `message` is accepted as an alias. A string or an array of `{type: text, text: ...}` parts." - removed
Input schema / properties / payload / properties / input / oneOfRemoved value: -[ - { - "type": "string" - }, - { - "items": { - "type": "object" - }, - "type": "array" - } -] - added
Input schema / properties / payload / properties / models / $refAdded value: +"#/$defs/models" - removed
Input schema / properties / payload / properties / models / descriptionRemoved value: -"Candidate model IDs to choose between. IDs must be nonempty, unique after remapping, and registered. Omit to use Chew's lane-chain union." - removed
Input schema / properties / payload / properties / models / itemsRemoved value: -{ - "type": "string" -} - removed
Input schema / properties / payload / properties / models / maxItemsRemoved value: -50 - removed
Input schema / properties / payload / properties / models / minItemsRemoved value: -1 - removed
Input schema / properties / payload / properties / models / typeRemoved value: -"array" - removed
Input schema / properties / payload / properties / models / uniqueItemsRemoved value: -true
- Changed
run_evaluations1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
run_organization_evaluation1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
search_brain12 fields changed- added
Input schema / $defs / source_typeAdded value: +{ + "description": "Restrict results to specific source types. Omit to search every source you can access. Valid values: `notion`, `google_drive`, `slack`, `github`, `confluence`, `direct_file_uploads`, `gumloop_artifacts`.\n", + "items": { + "enum": [ + "notion", + "google_drive", + "slack", + "github", + "confluence", + "direct_file_uploads", + "gumloop_artifacts" + ], + "type": "string" + }, + "minItems": 1, + "type": [ + "array", + "null" + ] +} - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / payload / properties / source_type / $refAdded value: +"#/$defs/source_type" - removed
Input schema / properties / payload / properties / source_type / descriptionRemoved value: -"Restrict results to specific source types. Omit to search every source you can access. Valid values: `notion`, `google_drive`, `slack`, `github`, `confluence`, `direct_file_uploads`, `gumloop_artifacts`.\n" - removed
Input schema / properties / payload / properties / source_type / itemsRemoved value: -{ - "enum": [ - "notion", - "google_drive", - "slack", - "github", - "confluence", - "direct_file_uploads", - "gumloop_artifacts" - ], - "type": "string" -} - removed
Input schema / properties / payload / properties / source_type / minItemsRemoved value: -1 - removed
Input schema / properties / payload / properties / source_type / typeRemoved value: -[ - "array", - "null" -] - added
Input schema / properties / source_type / $refAdded value: +"#/$defs/source_type" - removed
Input schema / properties / source_type / descriptionRemoved value: -"Restrict results to specific source types. Omit to search every source you can access. Valid values: `notion`, `google_drive`, `slack`, `github`, `confluence`, `direct_file_uploads`, `gumloop_artifacts`.\n" - removed
Input schema / properties / source_type / itemsRemoved value: -{ - "enum": [ - "notion", - "google_drive", - "slack", - "github", - "confluence", - "direct_file_uploads", - "gumloop_artifacts" - ], - "type": "string" -} - removed
Input schema / properties / source_type / minItemsRemoved value: -1 - removed
Input schema / properties / source_type / typeRemoved value: -[ - "array", - "null" -]
- Changed
send_message12 fields changed- added
Input schema / $defs / attachmentsAdded value: +{ + "description": "Files to attach to the message. Each `file_name` must be a stored path returned by Upload session file for this session.", + "items": { + "properties": { + "file_name": { + "description": "Stored path returned by Upload session file.", + "type": "string" + }, + "media_type": { + "description": "MIME type of the file.", + "type": [ + "string", + "null" + ] + } + }, + "required": [ + "file_name" + ], + "type": "object" + }, + "maxItems": 10, + "type": "array" +} - added
Input schema / properties / attachments / $refAdded value: +"#/$defs/attachments" - removed
Input schema / properties / attachments / descriptionRemoved value: -"Files to attach to the message. Each `file_name` must be a stored path returned by Upload session file for this session." - removed
Input schema / properties / attachments / itemsRemoved value: -{ - "properties": { - "file_name": { - "description": "Stored path returned by Upload session file.", - "type": "string" - }, - "media_type": { - "description": "MIME type of the file.", - "type": [ - "string", - "null" - ] - } - }, - "required": [ - "file_name" - ], - "type": "object" -} - removed
Input schema / properties / attachments / maxItemsRemoved value: -10 - removed
Input schema / properties / attachments / typeRemoved value: -"array" - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / payload / properties / attachments / $refAdded value: +"#/$defs/attachments" - removed
Input schema / properties / payload / properties / attachments / descriptionRemoved value: -"Files to attach to the message. Each `file_name` must be a stored path returned by Upload session file for this session." - removed
Input schema / properties / payload / properties / attachments / itemsRemoved value: -{ - "properties": { - "file_name": { - "description": "Stored path returned by Upload session file.", - "type": "string" - }, - "media_type": { - "description": "MIME type of the file.", - "type": [ - "string", - "null" - ] - } - }, - "required": [ - "file_name" - ], - "type": "object" -} - removed
Input schema / properties / payload / properties / attachments / maxItemsRemoved value: -10 - removed
Input schema / properties / payload / properties / attachments / typeRemoved value: -"array"
- Changed
send_queued_message1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
set_organization_evaluation_targets5 fields changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / payload / $defs / EvaluationTarget / $refAdded value: +"#/$defs/EvaluationTarget" - removed
Input schema / properties / payload / $defs / EvaluationTarget / propertiesRemoved value: -{ - "id": { - "description": "Team, user, or agent ID. Omitted for `organization`; responses return the organization ID.", - "type": "string" - }, - "type": { - "description": "What the target expands to. `user` covers a member's personal agents.", - "enum": [ - "organization", - "team", - "user", - "agent" - ], - "type": "string" - } -} - removed
Input schema / properties / payload / $defs / EvaluationTarget / requiredRemoved value: -[ - "type" -] - removed
Input schema / properties / payload / $defs / EvaluationTarget / typeRemoved value: -"object"
- Changed
set_role_credit_limit1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
start_flow1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
update_agent8 fields changed- added
Input schema / $defs / is_activeAdded value: +{ + "description": "Setting this to `false` retires the agent: it disappears from `GET /agents`, and `GET`/`PATCH /agents/{agent_id}` return `404`, so it cannot be reactivated through the API. This is not a pause switch — to stop an agent from running while keeping it reachable, disable its triggers instead.\n", + "type": [ + "boolean", + "null" + ] +} - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / is_active / $refAdded value: +"#/$defs/is_active" - removed
Input schema / properties / is_active / descriptionRemoved value: -"Setting this to `false` retires the agent: it disappears from `GET /agents`, and `GET`/`PATCH /agents/{agent_id}` return `404`, so it cannot be reactivated through the API. This is not a pause switch — to stop an agent from running while keeping it reachable, disable its triggers instead.\n" - removed
Input schema / properties / is_active / typeRemoved value: -[ - "boolean", - "null" -] - added
Input schema / properties / payload / properties / is_active / $refAdded value: +"#/$defs/is_active" - removed
Input schema / properties / payload / properties / is_active / descriptionRemoved value: -"Setting this to `false` retires the agent: it disappears from `GET /agents`, and `GET`/`PATCH /agents/{agent_id}` return `404`, so it cannot be reactivated through the API. This is not a pause switch — to stop an agent from running while keeping it reachable, disable its triggers instead.\n" - removed
Input schema / properties / payload / properties / is_active / typeRemoved value: -[ - "boolean", - "null" -]
- Changed
update_agent_skills1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
update_evaluation_config19 fields changed- added
Input schema / $defs / criteriaAdded value: +{ + "description": "Quality criteria to check (replaces existing list). Max 30.", + "items": { + "properties": { + "name": { + "type": "string" + }, + "priority": { + "enum": [ + "needs_review", + "needs_attention" + ], + "type": "string" + }, + "prompt": { + "description": "True/false statement to evaluate.", + "type": "string" + }, + "type": { + "enum": [ + "prohibited_action", + "prohibited_words", + "voice_tone", + "other" + ], + "type": "string" + } + }, + "type": "object" + }, + "type": "array" +} - added
Input schema / $defs / data_pointsAdded value: +{ + "description": "Data points to extract (replaces existing list). Max 40.", + "items": { + "properties": { + "data_type": { + "enum": [ + "string", + "boolean", + "integer", + "number" + ], + "type": "string" + }, + "description": { + "type": "string" + }, + "name": { + "type": "string" + } + }, + "type": "object" + }, + "type": "array" +} - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / criteria / $refAdded value: +"#/$defs/criteria" - removed
Input schema / properties / criteria / descriptionRemoved value: -"Quality criteria to check (replaces existing list). Max 30." - removed
Input schema / properties / criteria / itemsRemoved value: -{ - "properties": { - "name": { - "type": "string" - }, - "priority": { - "enum": [ - "needs_review", - "needs_attention" - ], - "type": "string" - }, - "prompt": { - "description": "True/false statement to evaluate.", - "type": "string" - }, - "type": { - "enum": [ - "prohibited_action", - "prohibited_words", - "voice_tone", - "other" - ], - "type": "string" - } - }, - "type": "object" -} - removed
Input schema / properties / criteria / typeRemoved value: -"array" - added
Input schema / properties / data_points / $refAdded value: +"#/$defs/data_points" - removed
Input schema / properties / data_points / descriptionRemoved value: -"Data points to extract (replaces existing list). Max 40." - removed
Input schema / properties / data_points / itemsRemoved value: -{ - "properties": { - "data_type": { - "enum": [ - "string", - "boolean", - "integer", - "number" - ], - "type": "string" - }, - "description": { - "type": "string" - }, - "name": { - "type": "string" - } - }, - "type": "object" -} - removed
Input schema / properties / data_points / typeRemoved value: -"array" - added
Input schema / properties / payload / properties / criteria / $refAdded value: +"#/$defs/criteria" - removed
Input schema / properties / payload / properties / criteria / descriptionRemoved value: -"Quality criteria to check (replaces existing list). Max 30." - removed
Input schema / properties / payload / properties / criteria / itemsRemoved value: -{ - "properties": { - "name": { - "type": "string" - }, - "priority": { - "enum": [ - "needs_review", - "needs_attention" - ], - "type": "string" - }, - "prompt": { - "description": "True/false statement to evaluate.", - "type": "string" - }, - "type": { - "enum": [ - "prohibited_action", - "prohibited_words", - "voice_tone", - "other" - ], - "type": "string" - } - }, - "type": "object" -} - removed
Input schema / properties / payload / properties / criteria / typeRemoved value: -"array" - added
Input schema / properties / payload / properties / data_points / $refAdded value: +"#/$defs/data_points" - removed
Input schema / properties / payload / properties / data_points / descriptionRemoved value: -"Data points to extract (replaces existing list). Max 40." - removed
Input schema / properties / payload / properties / data_points / itemsRemoved value: -{ - "properties": { - "data_type": { - "enum": [ - "string", - "boolean", - "integer", - "number" - ], - "type": "string" - }, - "description": { - "type": "string" - }, - "name": { - "type": "string" - } - }, - "type": "object" -} - removed
Input schema / properties / payload / properties / data_points / typeRemoved value: -"array"
- Changed
update_organization_evaluation1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
update_queued_message1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
update_skill14 fields changed- added
Input schema / $defs / filesAdded value: +{ + "description": "Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks.", + "items": { + "minLength": 1, + "type": "string" + }, + "maxItems": 25, + "minItems": 1, + "type": "array" +} - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / files / $refAdded value: +"#/$defs/files" - removed
Input schema / properties / files / descriptionRemoved value: -"Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks." - removed
Input schema / properties / files / itemsRemoved value: -{ - "minLength": 1, - "type": "string" -} - removed
Input schema / properties / files / maxItemsRemoved value: -25 - removed
Input schema / properties / files / minItemsRemoved value: -1 - removed
Input schema / properties / files / typeRemoved value: -"array" - added
Input schema / properties / payload / properties / files / $refAdded value: +"#/$defs/files" - removed
Input schema / properties / payload / properties / files / descriptionRemoved value: -"Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks." - removed
Input schema / properties / payload / properties / files / itemsRemoved value: -{ - "minLength": 1, - "type": "string" -} - removed
Input schema / properties / payload / properties / files / maxItemsRemoved value: -25 - removed
Input schema / properties / payload / properties / files / minItemsRemoved value: -1 - removed
Input schema / properties / payload / properties / files / typeRemoved value: -"array"
- Changed
upload_brain_files14 fields changed- added
Input schema / $defs / filesAdded value: +{ + "description": "Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks.", + "items": { + "minLength": 1, + "type": "string" + }, + "maxItems": 25, + "minItems": 1, + "type": "array" +} - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / files / $refAdded value: +"#/$defs/files" - removed
Input schema / properties / files / descriptionRemoved value: -"Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks." - removed
Input schema / properties / files / itemsRemoved value: -{ - "minLength": 1, - "type": "string" -} - removed
Input schema / properties / files / maxItemsRemoved value: -25 - removed
Input schema / properties / files / minItemsRemoved value: -1 - removed
Input schema / properties / files / typeRemoved value: -"array" - added
Input schema / properties / payload / properties / files / $refAdded value: +"#/$defs/files" - removed
Input schema / properties / payload / properties / files / descriptionRemoved value: -"Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks." - removed
Input schema / properties / payload / properties / files / itemsRemoved value: -{ - "minLength": 1, - "type": "string" -} - removed
Input schema / properties / payload / properties / files / maxItemsRemoved value: -25 - removed
Input schema / properties / payload / properties / files / minItemsRemoved value: -1 - removed
Input schema / properties / payload / properties / files / typeRemoved value: -"array"
- Changed
upload_file1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
- Changed
upload_files8 fields changed- added
Input schema / $defs / filesAdded value: +{ + "items": { + "properties": { + "file_content": { + "description": "Base64 encoded content of the file.", + "format": "byte", + "type": "string" + }, + "file_name": { + "description": "The name of the file to be uploaded.", + "type": "string" + } + }, + "type": "object" + }, + "type": "array" +} - changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action." - added
Input schema / properties / files / $refAdded value: +"#/$defs/files" - removed
Input schema / properties / files / itemsRemoved value: -{ - "properties": { - "file_content": { - "description": "Base64 encoded content of the file.", - "format": "byte", - "type": "string" - }, - "file_name": { - "description": "The name of the file to be uploaded.", - "type": "string" - } - }, - "type": "object" -} - removed
Input schema / properties / files / typeRemoved value: -"array" - added
Input schema / properties / payload / properties / files / $refAdded value: +"#/$defs/files" - removed
Input schema / properties / payload / properties / files / itemsRemoved value: -{ - "properties": { - "file_content": { - "description": "Base64 encoded content of the file.", - "format": "byte", - "type": "string" - }, - "file_name": { - "description": "The name of the file to be uploaded.", - "type": "string" - } - }, - "type": "object" -} - removed
Input schema / properties / payload / properties / files / typeRemoved value: -"array"
- Changed
upload_session_file1 field changed- changed
Input schema / properties / confirm / descriptionPrevious value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
92 tool updates
v2.0.1- First observed
approve_brain_source - First observed
attach_agent_mcp_server - First observed
call_mcp_tools - First observed
cancel_session - First observed
create_agent - First observed
create_brain_source - First observed
create_chat_completion - First observed
create_organization_evaluation - First observed
create_session - First observed
create_skill - First observed
delete_brain_file - First observed
delete_brain_source - First observed
delete_organization_evaluation - First observed
delete_queued_message - First observed
delete_skill - First observed
detach_agent_mcp_server - First observed
download_artifact_file - First observed
download_file - First observed
download_files - First observed
download_skill_file - First observed
export_data - First observed
get_brain_source - First observed
get_brain_source_estimate - First observed
get_evaluation_config - First observed
get_evaluation_metrics - First observed
get_evaluation_options - First observed
get_export_status - First observed
get_input_schema - First observed
get_mcp_server_prompt - First observed
get_organization_audit_logs - First observed
get_organization_evaluation - First observed
get_organization_evaluation_metrics - First observed
get_organization_evaluation_result - First observed
get_role_credit_limit - First observed
get_run_details - First observed
get_run_history - First observed
import_browser_profile_cookies - First observed
kill_flow - First observed
list_accounts - First observed
list_agent_mcp_servers - First observed
list_agent_versions - First observed
list_agents - First observed
list_artifacts - First observed
list_brain_files - First observed
list_brain_sources - First observed
list_browser_profiles - First observed
list_evaluations - First observed
list_flows - First observed
list_mcp_server_prompts - First observed
list_mcp_server_resources - First observed
list_mcp_server_tools - First observed
list_mcp_servers - First observed
list_models - First observed
list_organization_evaluation_results - First observed
list_organization_evaluations - First observed
list_organizations - First observed
list_queued_messages - First observed
list_role_credit_limits - First observed
list_sessions - First observed
list_skills - First observed
list_teams - First observed
list_workbooks - First observed
manage_permission_group_users - First observed
manage_project_users - First observed
queue_session_message - First observed
read_mcp_server_resource - First observed
rename_session - First observed
resolve_session_approvals - First observed
retrieve_agent - First observed
retrieve_agent_version - First observed
retrieve_evaluation - First observed
retrieve_mcp_server - First observed
retrieve_session - First observed
route_model - First observed
run_evaluations - First observed
run_organization_evaluation - First observed
search_brain - First observed
send_message - First observed
send_queued_message - First observed
set_organization_evaluation_targets - First observed
set_role_credit_limit - First observed
start_flow - First observed
update_agent - First observed
update_agent_skills - First observed
update_evaluation_config - First observed
update_organization_evaluation - First observed
update_queued_message - First observed
update_skill - First observed
upload_brain_files - First observed
upload_file - First observed
upload_files - First observed
upload_session_file
TDQS
Scored across 92 tools
Many tools have overlapping or indistinguishable purposes across the 92-tool surface, such as list_flows vs list_workbooks vs get_run_history, or upload_file vs upload_files vs upload_session_file vs upload_brain_files. with such a large flat set, boundaries between agent, session, flow, evaluation, brain, and skill tools are unclear without deep cross-referencing.
Most tools follow a consistent verb_noun snake_case pattern (list_agents, create_skill, delete_brain_source), with only minor deviations like route_model and call_mcp_tools. The convention is largely predictable.
92 tools is excessive for a single MCP server and likely buries the core operations. Many are thin variants (get_evaluation_metrics vs get_organization_evaluation_metrics) that could be consolidated.
Coverage is broad and deep across flows, agents, sessions, evaluations, brain, skills, and MCP, with both read and write operations. minor CRUD gaps exist but the surface is mostly complete.
Maintenance
Related MCP Connectors
- mcp-serverOAuthcom.make
Give your AI agents the tools to build, manage, and run automation workflows.
Your org's AI agents, tasks, runs, search, and brain files as MCP tools and resources.
Hosted AI agents and workflows with app OAuth, human approval gates, and a run ledger.
- SkilderOAuthai.skilder
One place to build, share, and govern the skills and tools your AI agents use at work.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to interact with Databricks workspaces programmatically, providing comprehensive tools for cluster management, notebook operations, job orchestration, Unity Catalog data governance, user management, permissions control, and FinOps cost analytics.492 npmMIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to manage n8n workflow automation instances through tools for workflow CRUD operations, execution monitoring, and webhook triggering. It facilitates programmatic interaction with n8n instances via the n8n API with AI-optimized descriptions and error handling.62MIT
- FlicenseBqualityFmaintenanceEnables AI assistants to manage GoClaw AI gateway infrastructure through a comprehensive suite of 66 tools for managing agents, sessions, and configurations. It provides real-time gateway context and guided workflows with enterprise-grade security features like audit logging and secret scrubbing.6610-
- AlicenseNot gradedqualityDmaintenanceProvides 70 tools to interact with the Brainbase API, enabling management of workers, chat/voice deployments, flows, resources, and more via natural language.1MIT