Skip to main content
Glama

Gumloop MCP Server & CLI

npm CI License YouTube X LinkedIn

Gumloop MCP server and CLI for Codex and AI agents. 92 tools for current flows, agents, sessions, Brain, skills, artifacts and administration, with private accounts and explicit operation approval. One shared implementation supplies both binaries and a desktop bundle.

Built and maintained by Navid Moazzez. Built on Slipway, which turns one definition of each tool into the MCP server and the CLI. The complete guide is on navid.me.

The terminal illustrates real command names and approval flow. It is not a recording of a provider account run. Gumloop already has official CLI and hosted MCP products; their supported platform, authentication and workflows are compared below.

Requires Node 22+ and eligible Gumloop API access for account operations. Validation: fixture tests, schema validation and protocol/artifact discovery are separate from provider-account and desktop GUI outcomes, which are not claimed. Section 7 has the measured token costs.

Two ways to use it

Command line

npm install -g @thenavidm/gumloop-mcp-cli@latest
gumloop-cli
gumloop-cli list-agents --agent
gumloop-cli schema start-flow
gumloop-cli start-flow --payload-file /absolute/private/approved-flow.json --account work --confirm --agent

MCP server, for your AI app

codex mcp add gumloop -- npx -y @thenavidm/gumloop-mcp-cli@latest

Configure private credentials first. Ask: “Inspect this existing run and return its state; do not start another run.” Read INSTALL.md for complete client/OS routes.

Which one

Where you work

Surface

Codex or shell agent

Shared CLI, local MCP or both

Desktop chat

Compatible local MCP/desktop archive

Scripts/CI

CLI or MCP client

Remote-only client

Official hosted MCP

Related MCP server: n8n Manager for AI Agents

Features

Capability

CLI command

MCP tool

Saved flows and inputs

list-flows / get-input-schema

list_flows / get_input_schema

Run and inspect

start-flow / get-run-details

start_flow / get_run_details

Agents and sessions

list-agents / create-session / retrieve-session

list_agents / create_session / retrieve_session

Skills and Brain

create-skill / search-brain

create_skill / search_brain

Private file delivery

download-artifact-file / download-files

download_artifact_file / download_files

Explicit account selection

list-accounts / --account

list_accounts / account

Diagnosis/setup

doctor / login

CLI utilities

Contents

Number

Section

What it covers

1

What you can ask it

What you can ask it

2

Quick install

Quick install

3

Set up Gumloop access

Set up Gumloop access

4

Connect your client

Connect your client

5

Check it works

Check it works

6

Output, flags and exit codes

Output, flags and exit codes

7

MCP or CLI and token cost

MCP or CLI and token cost

8

Every tool and argument

Every tool and argument

9

Flows, sessions and account workflows

Flows, sessions and account workflows

10

Jobs, pagination and private files

Jobs, pagination and private files

11

Several private accounts

Several private accounts

12

Writing safely

Writing safely

13

How it works

How it works

14

Your data

Your data

15

Environment variables

Environment variables

16

Updates and removal

Updates and removal

17

Troubleshooting

Troubleshooting

18

API coverage and comparisons

API coverage and comparisons

19

Versions

Versions

20

FAQ

FAQ

1. What you can ask it

  • Inspect accessible agents, versions, teams and models.

  • List saved flows and their exact input schema before approving a run.

  • Start only the selected flow, keep its returned run ID and read its state.

  • Create a reviewed agent session, then inspect its approval requirements.

  • Approve only the selected session checkpoint; do not approve a whole conversation implicitly.

  • Upload a selected regular file and save an existing run download privately.

  • Review organization usage, audit records and applicable role credit limits.

Actual discovery supplies 92 tools: 50 reads and 42 confirmation-gated operations, from all 91 reviewed current REST operations plus list_accounts. The shared handlers cover MCP and native Node CLI. This is protocol/schema evidence; account outcomes and desktop GUI installation remain separate.

2. Quick install

npm install -g @thenavidm/gumloop-mcp-cli@latest
gumloop-cli --version
gumloop-cli login
gumloop-cli doctor
gumloop-cli tools

Node 22+ is required for manual installation. The versioned desktop archive bundles production dependencies for a compatible host. Read INSTALL.md before configuring credentials. After private setup:

codex mcp add gumloop -- npx -y @thenavidm/gumloop-mcp-cli@latest
codex mcp list

3. Set up Gumloop access

Private credentials and permissions

  1. Open the intended account's Connectors settings. Select Gumloop API Key and create the personal or team key needed for your task.

  2. Read Profile Settings for your user ID. A team ID is optional; older flow endpoints call it project_id. Set GUMLOOP_USER_ID and, when appropriate, GUMLOOP_TEAM_ID in private settings.

  3. Save the credential as a token-only regular file outside repositories. Set GUMLOOP_TOKEN_FILE to its absolute path; GUMLOOP_API_KEY is the alternative. Do not paste real keys into chat or command examples.

  4. Run gumloop-cli doctor, then doctor --network. The network check reads accessible agent metadata without printing account content.

  5. Inspect the exact schema and selected resource before confirming an operation. Runs can spend credits and invoke connected downstream services.

Current authentication docs require Pro or above for API keys. Personal keys act as their owner; team keys can act as team members through the requested user identity. Workflow-step credential settings decide personal/team credentials when project_id is supplied. Local profiles do not bypass account, role or team permissions.

Requests use Authorization: Bearer, with the selected profile user ID as x-auth-key, following the current SDK. Ordinary endpoints use https://api.gumloop.com/api/v1; non-streaming chat uses the separately documented https://ws.gumloop.com/api/v1. No arbitrary host override or redirects are accepted.

For POSIX, use a private 0700 directory and regular 0600 token-only file, no symlink, at most 64 KB. Windows requires a user-only ACL; POSIX checks do not verify it. File credentials override environment credentials and are cached until restart. The package does not read an automatic .env, use the OS keychain or start browser login.

OAuth and rotation

An already authorized OAuth access token can be supplied through the same private Bearer credential path. This wrapper does not register an OAuth app, accept refresh tokens, refresh access tokens or manage consent. OAuth docs describe invite-only client registration, authorization code with PKCE S256 and gumloop_api/userinfo scopes. gumloop_api requires Pro or above; requesting userinfo alone is not API permission. Use an expiry-aware issuer or the official CLI's supported OAuth/keychain workflow when automatic refresh is needed.

Rotate/revoke the intended grant through its provider controls, update private settings and restart. Removing a local package does not revoke provider credentials, undo runs or remove hosted data.

Plans, credits and rate limits

The AGPL wrapper is free; Gumloop account access and processing are billed separately. Credit documentation describes variable charges for model work, connector calls, compute and orchestration; Brain searches can also incur charges. Local read-only mode is an operation policy, not a guarantee of zero provider charges. Credit-consuming Brain search is confirmation-gated here.

Agent concurrency limits are organization-wide, currently 25 on Pro and 100 on Enterprise, with customizable Enterprise limits. Pro rejects excess agent work; Enterprise can queue it. Webhook triggers separately document 100 requests/minute per trigger. These are distinct limits; this wrapper's default 150 ms account/process pacing reserves no provider capacity.

GET 429 retries require an explicit Retry-After at most ten seconds, default two/max five retries. Missing/longer delays return exit 7. Mutations, paid Brain search and network timeouts never retry automatically. Local JSON request cap is 5 MiB; responses/downloads are capped at 10 MiB. These local caps are separate from vendor storage, upload and pagination limits.

4. Connect your client

INSTALL.md gives Codex-first setup, optional Claude Code, desktop archive/manual configuration, Cursor, VS Code/Copilot, Windsurf, Zed, Gemini CLI, Docker and Cline. The local transport is stdio. Configure private credentials in the actual client environment; GUI apps and remote containers do not inherit every shell setting.

The published official gumloop 0.5.2 CLI and current provider docs explicitly refuse native Windows and recommend WSL or the Python SDK. This wrapper targets Node 22+ on native Windows, macOS and Linux. Compatible desktop-host availability and GUI installation are separate from CLI portability. Remote-only clients can use the official https://mcp.gumloop.com/gumloop/mcp endpoint.

npm ships SKILL.md but does not register it automatically. Install it through the client's supported skill location. A connection/client approval and this wrapper's confirm argument are separate controls.

5. Check it works

gumloop-cli --version
gumloop-cli doctor
gumloop-cli doctor --network
gumloop-cli list-accounts --agent
gumloop-cli list-agents --agent

Discovery, schemas/help and account labels need no provider key. Network doctor reads GET /agents and prints only diagnostic status. This validates that request, not every account operation. Full discovery exposes 92 tools; read-only discovery exposes 50. Missing credentials exits 10, invalid input or refused operations exit 2. First validate an existing resource; do not start billable automation merely to test installation.

6. Output, flags and exit codes

Tool results go to stdout. Errors are JSON on stderr. JSON operations return structured results, so --select can retain nested fields. Private download operations return file metadata; binary bytes and signed results stay in the requested private file.

gumloop-cli start-flow --help
gumloop-cli schema start-flow
gumloop-cli list-agents --agent --select agents

Flag

What it does

--json

JSON output

--compact

Single-line JSON

--agent

Compact JSON and no prompts; never confirms a write

--select a,b.c

Keep selected fields; dotted paths descend and arrays are traversed

--confirm

Confirm the requested flow, session, paid search or account operation

--no-input, --no-color, --yes

Automation switches; none overrides the spending guard

Global output flags apply to tool commands. doctor has its own --network option and returns a JSON diagnostic.

Exit code

Meaning

What a script should do

0

Success

Read stdout

1

Unexpected error

Report it with the command that caused it

2

Usage, invalid input, a refused write, an unknown command or a hidden write

Fix the input or confirm only the requested action

3

Job or local upload file not found

Check the ID/path

4

Authentication or entitlement rejected

Check private credential settings and permissions

5

API or network failure

Inspect an accepted job before another paid submission

7

Rate limited

Wait; do not loop over paid submissions

10

Credentials not configured

Complete local setup

The underscore spelling also works. start_flow and start-flow call the same tool. Nested objects use quoted JSON. Arrays of objects use repeated flags, one JSON object at a time.

7. MCP or CLI and token cost

MCP and CLI use the same catalogue, schemas, handlers and write guard: Slipway builds the MCP server, over stdio or --http, and the CLI from each tool's one definition. Shell scripts can select fields with --select after receipt; this does not change upstream result size or billing.

Measured on 2026-10-05 against 2.0.2, with Claude Code 2.1.286 on Claude Opus 5.5 (one short prompt with and without the server connected, the difference read from the API's own usage figures) and Codex 0.159.3 on gpt-6.1-sol:

Cost

2.0.2

3.0.0

Claude Code, every tool loaded, every message

61,936

56,695

Claude Code's default, tool search, every message

1,486

1,488

SKILL.md, read once

3,150

3,210

Codex over the CLI, one task, median of five

107,376

83,661

Codex over MCP, the same task, median of five

78,443

78,249

The task was "find the command that starts a flow run, and the flags it requires". Every tool loaded costs less because parts that several tools repeated are written once. Over the CLI, every 3.0.0 run asked which, whose answer carries the command's help, where every 2.0.2 run read the command list and then the help: one request fewer. Over MCP, 3.0.0 cost slightly less. SKILL.md costs 60 more because it says how approval works over MCP and what exit codes 1 and 2 cover.

Tool-list bytes or characters divided by four are not API usage, and no other offering was measured.

Provider credits and client-model tokens are separate.

8. Every tool and argument

Every route and argument below comes from actual stdio discovery and the reviewed current schema. Use schema COMMAND for exact inline nested objects and unions.

Tool

Route

Mode

start_flow

POST /api/v1/start_pipeline

Confirm exact operation

kill_flow

POST /api/v1/kill_pipeline

Confirm exact operation

get_run_details

GET /api/v1/get_pl_run

Read

list_workbooks

GET /api/v1/list_workbooks

Read

list_flows

GET /api/v1/list_saved_items

Read

get_input_schema

GET /api/v1/get_inputs

Read

get_run_history

GET /api/v1/get_plrun_saved_item_map

Read

download_file

POST /api/v1/download_file

Read

download_files

POST /api/v1/download_files

Read

upload_file

POST /api/v1/upload_file

Confirm exact operation

upload_files

POST /api/v1/upload_files

Confirm exact operation

get_organization_audit_logs

GET /api/v1/get_audit_logs

Read

manage_project_users

POST /api/v1/manage_workspace_users

Confirm exact operation

manage_permission_group_users

POST /api/v1/manage_permission_group_users

Confirm exact operation

list_role_credit_limits

GET /api/v1/organizations/{organization_id}/roles/credit-limits

Read

get_role_credit_limit

GET /api/v1/organizations/{organization_id}/roles/{role_id}/credit-limit

Read

set_role_credit_limit

PUT /api/v1/organizations/{organization_id}/roles/{role_id}/credit-limit

Confirm exact operation

export_data

POST /api/v1/export_data

Confirm exact operation

get_export_status

GET /api/v1/export_status

Read

list_agents

GET /api/v1/agents

Read

create_agent

POST /api/v1/agents

Confirm exact operation

retrieve_agent

GET /api/v1/agents/{agent_id}

Read

update_agent

PATCH /api/v1/agents/{agent_id}

Confirm exact operation

create_chat_completion

POST /api/v1/chat/completions

Confirm exact operation

list_models

GET /api/v1/models

Read

route_model

POST /api/v1/models/route

Read

list_agent_versions

GET /api/v1/agents/{agent_id}/versions

Read

retrieve_agent_version

GET /api/v1/agents/{agent_id}/versions/{version_id}

Read

update_agent_skills

PATCH /api/v1/agents/{agent_id}/skills

Confirm exact operation

list_agent_mcp_servers

GET /api/v1/agents/{agent_id}/mcp-servers

Read

attach_agent_mcp_server

PUT /api/v1/agents/{agent_id}/mcp-servers/{server_id}

Confirm exact operation

detach_agent_mcp_server

DELETE /api/v1/agents/{agent_id}/mcp-servers/{server_id}

Confirm exact operation

list_sessions

GET /api/v1/agents/{agent_id}/sessions

Read

create_session

POST /api/v1/agents/{agent_id}/sessions

Confirm exact operation

retrieve_session

GET /api/v1/sessions/{session_id}

Read

rename_session

PATCH /api/v1/sessions/{session_id}

Confirm exact operation

send_message

POST /api/v1/sessions/{session_id}/messages

Confirm exact operation

cancel_session

POST /api/v1/sessions/{session_id}/cancel

Confirm exact operation

upload_session_file

POST /api/v1/sessions/{session_id}/files

Confirm exact operation

resolve_session_approvals

POST /api/v1/sessions/{session_id}/approvals

Confirm exact operation

list_queued_messages

GET /api/v1/sessions/{session_id}/queue

Read

queue_session_message

POST /api/v1/sessions/{session_id}/queue

Confirm exact operation

update_queued_message

PATCH /api/v1/sessions/{session_id}/queue/{queued_message_id}

Confirm exact operation

delete_queued_message

DELETE /api/v1/sessions/{session_id}/queue/{queued_message_id}

Confirm exact operation

send_queued_message

POST /api/v1/sessions/{session_id}/queue/{queued_message_id}/send

Confirm exact operation

list_mcp_servers

GET /api/v1/mcp/servers

Read

retrieve_mcp_server

GET /api/v1/mcp/servers/{server_id}

Read

list_mcp_server_tools

GET /api/v1/mcp/servers/{server_id}/tools

Read

list_mcp_server_resources

GET /api/v1/mcp/servers/{server_id}/resources

Read

read_mcp_server_resource

GET /api/v1/mcp/servers/{server_id}/resources/read

Read

list_mcp_server_prompts

GET /api/v1/mcp/servers/{server_id}/prompts

Read

get_mcp_server_prompt

POST /api/v1/mcp/servers/{server_id}/prompts/get

Read

call_mcp_tools

POST /api/v1/mcp/tools/call

Confirm exact operation

search_brain

POST /api/v1/brain/search

Confirm exact operation

list_brain_sources

GET /api/v1/brain/sources

Read

create_brain_source

POST /api/v1/brain/sources

Confirm exact operation

get_brain_source

GET /api/v1/brain/sources/{source_id}

Read

delete_brain_source

DELETE /api/v1/brain/sources/{source_id}

Confirm exact operation

list_brain_files

GET /api/v1/brain/sources/{source_id}/files

Read

upload_brain_files

POST /api/v1/brain/sources/{source_id}/files

Confirm exact operation

delete_brain_file

DELETE /api/v1/brain/sources/{source_id}/files/{file_id}

Confirm exact operation

get_brain_source_estimate

GET /api/v1/brain/sources/{source_id}/estimate

Read

approve_brain_source

POST /api/v1/brain/sources/{source_id}/approve

Confirm exact operation

list_skills

GET /api/v1/skills

Read

create_skill

POST /api/v1/skills

Confirm exact operation

update_skill

PATCH /api/v1/skills/{skill_id}

Confirm exact operation

delete_skill

DELETE /api/v1/skills/{skill_id}

Confirm exact operation

download_skill_file

GET /api/v1/skills/{skill_id}/download

Read

list_artifacts

GET /api/v1/agents/{agent_id}/artifacts

Read

download_artifact_file

GET /api/v1/artifacts/{artifact_id}/download

Read

list_browser_profiles

GET /api/v1/browser-profiles

Read

import_browser_profile_cookies

POST /api/v1/browser-profiles/{profile_id}/cookies

Confirm exact operation

list_teams

GET /api/v1/teams

Read

list_evaluations

GET /api/v1/agents/{agent_id}/evaluations

Read

run_evaluations

POST /api/v1/agents/{agent_id}/evaluations/run

Confirm exact operation

get_evaluation_metrics

GET /api/v1/agents/{agent_id}/evaluations/metrics

Read

retrieve_evaluation

GET /api/v1/agents/{agent_id}/evaluations/{evaluation_id}

Read

get_evaluation_config

GET /api/v1/agents/{agent_id}/evaluation-config

Read

update_evaluation_config

PATCH /api/v1/agents/{agent_id}/evaluation-config

Confirm exact operation

list_organizations

GET /api/v1/organizations

Read

get_evaluation_options

GET /api/v1/evaluation-options

Read

list_organization_evaluations

GET /api/v1/evaluations

Read

create_organization_evaluation

POST /api/v1/evaluations

Confirm exact operation

get_organization_evaluation

GET /api/v1/evaluations/{evaluation_id}

Read

update_organization_evaluation

PATCH /api/v1/evaluations/{evaluation_id}

Confirm exact operation

delete_organization_evaluation

DELETE /api/v1/evaluations/{evaluation_id}

Confirm exact operation

set_organization_evaluation_targets

PUT /api/v1/evaluations/{evaluation_id}/targets

Confirm exact operation

run_organization_evaluation

POST /api/v1/evaluations/{evaluation_id}/run

Confirm exact operation

list_organization_evaluation_results

GET /api/v1/evaluations/{evaluation_id}/results

Read

get_organization_evaluation_result

GET /api/v1/evaluations/{evaluation_id}/results/{result_id}

Read

get_organization_evaluation_metrics

GET /api/v1/evaluations/{evaluation_id}/metrics

Read

list_accounts

Local, no network

Read

start_flow

gumloop-cli start-flow

Argument

Required

Type

Details

user_id

No; body and guard rules apply

string

The id for the user initiating the flow.

project_id

No; body and guard rules apply

string

(Optional) The id of the project within which the flow is executed.

saved_item_id

No; body and guard rules apply

string

The id for the saved flow.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: saved_item_id, user_id. Profile defaults fill supported user/team identity fields before body validation.

kill_flow

gumloop-cli kill-flow

Argument

Required

Type

Details

run_id

No; body and guard rules apply

string

The ID of the pipeline run to kill.

user_id

No; body and guard rules apply

string

The user ID. Required if project_id is not provided.

project_id

No; body and guard rules apply

string

The project ID. Required if user_id is not provided.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: run_id. Profile defaults fill supported user/team identity fields before body validation.

get_run_details

gumloop-cli get-run-details

Argument

Required

Type

Details

run_id

Yes

string

ID of the flow run to retrieve

user_id

No; body and guard rules apply

string

The id for the user initiating the flow. Required if project_id is not provided.

project_id

No; body and guard rules apply

string

The id of the project within which the flow is executed. Required if user_id is not provided.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

list_workbooks

gumloop-cli list-workbooks

Argument

Required

Type

Details

user_id

No; body and guard rules apply

string

The user ID for which to list workbooks. Required if project_id is not provided.

project_id

No; body and guard rules apply

string

The project ID for which to list workbooks. Required if user_id is not provided.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

list_flows

gumloop-cli list-flows

Argument

Required

Type

Details

user_id

No; body and guard rules apply

string

The user ID to for which to list items. Required if project_id is not provided.

project_id

No; body and guard rules apply

string

The project ID for which to list items. Required if user_id is not provided.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

get_input_schema

gumloop-cli get-input-schema

Argument

Required

Type

Details

saved_item_id

Yes

string

The ID of the saved item for which to retrieve input schemas.

user_id

No; body and guard rules apply

string

User ID that created the flow. Required if project_id is not provided.

project_id

No; body and guard rules apply

string

Project ID that the flow is under. Required if user_id is not provided.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

get_run_history

gumloop-cli get-run-history

Argument

Required

Type

Details

workbook_id

No; body and guard rules apply

string

The ID of the workbook to retrieve run history for. Required if saved_item_id is not provided.

saved_item_id

No; body and guard rules apply

string

The ID of the saved item to retrieve run history for. Required if workbook_id is not provided.

user_id

No; body and guard rules apply

string

The user ID. Required if project_id is not provided.

project_id

No; body and guard rules apply

string

The project ID. Required if user_id is not provided.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

download_file

gumloop-cli download-file

Argument

Required

Type

Details

file_name

No; body and guard rules apply

string

The name of the file to download.

run_id

No; body and guard rules apply

string

The ID of the flow run associated with the file.

saved_item_id

No; body and guard rules apply

string

The saved item ID associated with the file.

user_id

No; body and guard rules apply

string

Optional. The user ID associated with the flow run.

project_id

No; body and guard rules apply

string

Optional. The project ID associated with the flow run.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

output_file

Yes

string

New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.

output_file must be new and absolute in a private directory, reserved exclusively before the API call. Binary bytes or signed JSON are kept out of model output. This wrapper never follows a signed external download URL.

download_files

gumloop-cli download-files

Argument

Required

Type

Details

file_names

No; body and guard rules apply

array

An array of file names to download. Items: string.

run_id

No; body and guard rules apply

string

The ID of the flow run associated with the files.

user_id

No; body and guard rules apply

string

The user ID associated with the files. Required if project_id is not provided.

project_id

No; body and guard rules apply

string

The project ID associated with the files. Required if user_id is not provided.

saved_item_id

No; body and guard rules apply

string

Optional. The saved item ID associated with the files.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

output_file

Yes

string

New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.

output_file must be new and absolute in a private directory, reserved exclusively before the API call. Binary bytes or signed JSON are kept out of model output. This wrapper never follows a signed external download URL.

upload_file

gumloop-cli upload-file

Argument

Required

Type

Details

file_name

No; body and guard rules apply

string

The name of the file to be uploaded.

file_content

No; body and guard rules apply

string

Base64 encoded content of the file. format: byte.

user_id

No; body and guard rules apply

string

The user ID associated with the file. Required if project_id is not provided.

project_id

No; body and guard rules apply

string

The project ID associated with the file. Required if user_id is not provided.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

file_path

No; body and guard rules apply

string

Regular local file path, no symlink, at most 3 MiB. Encoded as native base64 file_content; cannot mix with file_content or payload routes. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.

upload_files

gumloop-cli upload-files

Argument

Required

Type

Details

files

No; body and guard rules apply

array

See the full input schema. Items: object.

user_id

No; body and guard rules apply

string

The user ID associated with the files. Required if project_id is not provided.

project_id

No; body and guard rules apply

string

The project ID associated with the files. Required if user_id is not provided.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.

get_organization_audit_logs

gumloop-cli get-organization-audit-logs

Argument

Required

Type

Details

organization_id

Yes

string

The ID of the organization to retrieve audit logs for.

user_id

No; body and guard rules apply

string

Your user id -- you must be an organization admin to retrieve organization logs.

start_time

Yes

string

Start timestamp for log filtering (ISO format). format: date-time.

end_time

Yes

string

End timestamp for log filtering (ISO format). format: date-time.

event_types

No; body and guard rules apply

string

Comma-separated list of event types to filter by (e.g. user_sign_in,credential_retrieval). The singular event_type param accepts a single value.

user_ids

No; body and guard rules apply

string

Comma-separated list of user IDs whose events should be returned.

ip_addresses

No; body and guard rules apply

string

Comma-separated list of source IP addresses to filter by. The singular ip_address param accepts a single value.

workspace_ids

No; body and guard rules apply

string

Comma-separated list of workspace (team) IDs to filter by. The singular workspace_id param accepts a single value.

entity_ids

No; body and guard rules apply

string

Comma-separated list of entity IDs (agents, workbooks, files) to filter by. The singular entity_id param accepts a single value.

page

No; body and guard rules apply

integer

Page number for pagination. default: 1.

page_size

No; body and guard rules apply

integer

Number of records per page. default: 50.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

manage_project_users

gumloop-cli manage-project-users

Argument

Required

Type

Details

organization_id

No; body and guard rules apply

string

The ID of the organization that the workspace belongs to.

user_id

No; body and guard rules apply

string

Your user id -- you must be an organization admin to manage workspace users.

workspace_id

No; body and guard rules apply

string

The ID of the workspace to manage users for.

action

No; body and guard rules apply

string

The action to perform - either 'add' or 'remove' a user. Values: add, remove.

user_email

No; body and guard rules apply

string

The email address of the target user to add or remove.

is_admin

No; body and guard rules apply

boolean

When adding a user, specify whether they should have admin privileges (default is false).

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: organization_id, user_id, workspace_id, action, user_email. Profile defaults fill supported user/team identity fields before body validation.

manage_permission_group_users

gumloop-cli manage-permission-group-users

Argument

Required

Type

Details

organization_id

No; body and guard rules apply

string

The ID of the organization that the custom role belongs to.

user_id

No; body and guard rules apply

string

Your user id -- you must be an organization admin to manage custom role users.

group_id

No; body and guard rules apply

string

The ID of the custom role to manage users for.

action

No; body and guard rules apply

string

The action to perform - either 'add' or 'remove' a user. Values: add, remove.

user_email

No; body and guard rules apply

string

The email address of the target user to add or remove.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: organization_id, user_id, group_id, action, user_email. Profile defaults fill supported user/team identity fields before body validation.

list_role_credit_limits

gumloop-cli list-role-credit-limits

Argument

Required

Type

Details

organization_id

Yes

string

The ID of the organization. minLength: 1.

user_id

No; body and guard rules apply

string

Your user id -- you must be an organization admin to manage custom role credit limits.

page_size

No; body and guard rules apply

integer

Number of roles per page. maximum: 100. default: 20.

cursor

No; body and guard rules apply

string

Opaque cursor from a previous response's next_cursor; omit for the first page.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

get_role_credit_limit

gumloop-cli get-role-credit-limit

Argument

Required

Type

Details

organization_id

Yes

string

The ID of the organization the custom role belongs to. minLength: 1.

role_id

Yes

string

The ID of the custom role (the same ID used as group_id by the Manage custom role users endpoint). minLength: 1.

user_id

No; body and guard rules apply

string

Your user id -- you must be an organization admin to manage custom role credit limits.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

set_role_credit_limit

gumloop-cli set-role-credit-limit

Argument

Required

Type

Details

organization_id

Yes

string

The ID of the organization the custom role belongs to. minLength: 1.

role_id

Yes

string

The ID of the custom role (the same ID used as group_id by the Manage custom role users endpoint). minLength: 1.

monthly_credit_limit

No; body and guard rules apply

integer/null

The monthly credit limit applied to each member of this role, or null to clear the role-level limit. minimum: 0. maximum: 1000000000.

user_id

No; body and guard rules apply

string

Your user id -- you must be an organization admin to manage custom role credit limits.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: monthly_credit_limit, user_id. Profile defaults fill supported user/team identity fields before body validation.

export_data

gumloop-cli export-data

Argument

Required

Type

Details

user_id

No; body and guard rules apply

string

The ID of the user requesting the export.

data_type

No; body and guard rules apply

string

The type of data to export. Use "workflows" to export workflow run data, "agents" to export agent configuration data, "agent_interactions" to export agent interaction data, "credit_logs" to export credit transaction history, "interaction_evaluations" to export completed chat evaluations, or "gumstack" to export Gumstack tool call activity. Defaults to "workflows". Audit-log exports are available in the Gumloop app only, not through this endpoint. Values: workflows, agents, agent_interactions, credit_logs, interaction_evaluations, gumstack. default: workflows.

category_filter

No; body and guard rules apply

string

Only applicable when data_type is "credit_logs". Optional filter to export only credit logs matching a specific category (e.g., "PIPELINE_RUN", "AGENT_RUN", "CUSTOM_NODE_RD", "EXTERNAL_GUMCP_CALL"). When omitted, all categories are included.

export_level

No; body and guard rules apply

string

The scope of the export. Use "organization" to export across the entire organization, or "workspace" to export from a single workspace (requires exactly one ID in workspace_ids). Defaults to "organization". Not applicable when data_type is "credit_logs". Values: workspace, organization. default: organization.

export_fields

No; body and guard rules apply

array

Array of fields to include in the export. The available fields depend on the data_type. Workflow fields (when data_type is "workflows"): - workbook_id - The workbook identifier - workbook_name - The workbook name - workbook_created_ts - Workbook creation timestamp - user_id - The user identifier - user_email - The user's email address - workspace_id - The workspace identifier - workspace_name - The workspace name - run_id - The flow run identifier - credit_cost - Credits consumed by the run - pl_run_created_ts - Flow run creation timestamp - pl_run_finished_ts - Flow run completion timestamp - pipeline - The full pipeline configuration (JSON) Agent fields (when data_type is "agents"): - agent_id - The agent identifier - agent_name - The agent name - agent_description - The agent description - agent_model - The model used by the agent - agent_system_prompt - The agent's system prompt - agent_created_ts - Agent creation timestamp - agent_tools - Tools configured for the agent (JSON) - agent_metadata - Additional agent metadata (JSON) - agent_evaluations_enabled - Whether the agent's own Evaluations are turned on (true/false) - creator_email - Email of the user who created the agent - workspace_id - The workspace identifier - workspace_name - The workspace name - folder_id - The identifier of the folder containing the agent (empty when the agent is not in a folder) - folder_name - The name of the folder containing the agent (empty when the agent is not in a folder) Agent interaction fields (when data_type is "agent_interactions"): - interaction_id - Unique identifier for the chat session - agent_id - The agent identifier - agent_name - The agent name - interaction_type - Type of interaction (e.g., chat, slack, api, triggered) - interaction_name - Display name of the chat session - trigger_type - For triggered interactions, the specific trigger type (e.g., time_based, polling_new_record_salesforce). Null for non-triggered interactions. - interaction_created_ts - Chat session creation timestamp - user_email - Email of the user who initiated the chat - credit_cost - Total credits consumed (LLM + tool + flow) - llm_credit_cost - Credits consumed by LLM calls only - tool_credit_cost - Credits consumed by tool calls - flow_credit_cost - Credits consumed by pipeline runs - message_count - Number of messages in the conversation - workspace_id - The workspace identifier - workspace_name - The workspace name Interaction evaluation fields (when data_type is "interaction_evaluations"): - evaluation_id - Unique identifier for the evaluation row - interaction_id - The chat session that was evaluated (one chat can have several evaluation rows) - agent_id - The identifier of the agent that was evaluated - organization_evaluation_id - The organization-level evaluation this row belongs to (empty for agent-level rubrics) - evaluation_created_ts - When the evaluation reached its final state - status - The evaluation's state - grade - The grade the evaluation produced - call_outcome - The outcome the evaluation recorded for the conversation - sentiment - The sentiment the evaluation recorded - error_code - Error code when the evaluation could not complete - summary - The evaluation's written summary - evaluation_model - The model that ran the evaluation - credit_cost - Credits consumed by the evaluation - user_email - Email of the user whose chat was evaluated Credit log fields (when data_type is "credit_logs"): - user_email - Email of the user associated with the credit log entry - permission_group_id - Custom role ID(s) the user belongs to (semicolon-separated if multiple) (disabled by default) - permission_group_name - Custom role name(s) the user belongs to (semicolon-separated if multiple) (disabled by default) - timestamp - When the credit transaction occurred - category - The category of the credit log (e.g., PIPELINE_RUN, AGENT_RUN) - type - The specific type of credit charge - name - Display name of the credit log entry - amount - Number of credits charged or adjusted - balance - Credit balance after the transaction - log_id - Unique identifier for the credit log entry - correlation_id - Join key to the related run or interaction (disabled by default) - balance_scope - Whether the balance is organization- or user-scoped (disabled by default) - project_id - The project identifier (disabled by default) Not all combinations of selected fields are guaranteed to produce data for every row. Items: string.

start_date

No; body and guard rules apply

string

Start date for the export in ISO 8601 format (e.g., 2025-01-01T00:00:00Z). format: date-time.

end_date

No; body and guard rules apply

string

End date for the export in ISO 8601 format (e.g., 2025-12-31T23:59:59Z). format: date-time.

include_all_workspaces

No; body and guard rules apply

boolean

Whether to include all workspaces in the organization. When true, also sets include_personal_workspaces to true. Not applicable when data_type is "credit_logs". default: False.

include_personal_workspaces

No; body and guard rules apply

boolean

Whether to include personal workspaces in the export. Ignored if include_all_workspaces is true. Not applicable when data_type is "credit_logs". default: False.

workspace_ids

No; body and guard rules apply

array

An optional array of workspace IDs to include in the export. When export_level is "workspace", exactly one workspace ID is required. Ignored if include_all_workspaces is true. Not applicable when data_type is "credit_logs". Items: string.

entity_ids

No; body and guard rules apply

array

An optional array of specific entity IDs to filter the export. For workflow exports (data_type: "workflows"), these are workbook IDs. For agent exports (data_type: "agents") and agent interaction exports (data_type: "agent_interactions"), these are agent IDs. When provided, only data for the specified entities will be included. Not applicable when data_type is "credit_logs". Items: string.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: user_id, export_fields, start_date, end_date. Profile defaults fill supported user/team identity fields before body validation.

get_export_status

gumloop-cli get-export-status

Argument

Required

Type

Details

user_id

No; body and guard rules apply

string

The ID of the user requesting the export status.

data_export_id

Yes

string

The unique identifier of the data export job to check (returned by the Export data endpoint).

download

No; body and guard rules apply

boolean

Set to true to download the export file directly when the export is completed. When true and the export state is COMPLETED, the response will be a CSV file download instead of JSON. default: False.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

output_file

Yes

string

New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. minLength: 1.

output_file must be new and absolute in a private directory, reserved exclusively before the API call. Binary bytes or signed JSON are kept out of model output. This wrapper never follows a signed external download URL.

list_agents

gumloop-cli list-agents

Argument

Required

Type

Details

team_id

No; body and guard rules apply

string

Scope the listing to a single team. When omitted, returns agents owned by the authenticated user.

search

No; body and guard rules apply

string

Case-insensitive substring match against the agent name.

creator

No; body and guard rules apply

string

Filter to agents created by this user ID.

has_triggers

No; body and guard rules apply

boolean

When true, only returns agents that have at least one active trigger configured.

tool

No; body and guard rules apply

string

Filter to agents that use the specified MCP server as a tool.

flow

No; body and guard rules apply

string

Filter to agents that use the specified saved flow as a tool.

sort_order

No; body and guard rules apply

string

Sort order for the listing. Defaults to newest first. Values: newest, oldest, name_asc, name_desc. default: newest.

include_last_used

No; body and guard rules apply

boolean

When true, populates last_used_at on each agent with the timestamp of its most recent session.

include_last_updated

No; body and guard rules apply

boolean

When true, populates last_updated_at on each agent with the timestamp of its most recent configuration change.

page_size

No; body and guard rules apply

integer

Number of agents per page. Sending page_size or cursor opts into cursor pagination; requests that send neither return the full list. minimum: 1. maximum: 100. default: 20.

cursor

No; body and guard rules apply

string

Opaque cursor from a previous response's next_cursor. Pass it to fetch the next page.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

create_agent

gumloop-cli create-agent

Argument

Required

Type

Details

name

No; body and guard rules apply

string

Display name for the agent.

model_name

No; body and guard rules apply

string

ID of the LLM the agent runs on. Use GET /models to discover valid values.

description

No; body and guard rules apply

string/null

See the full input schema.

system_prompt

No; body and guard rules apply

string/null

See the full input schema.

tools

No; body and guard rules apply

array

Tools the agent can call. Each tool is an object whose shape depends on the tool type. default: []. Items: object.

resources

No; body and guard rules apply

array

Resources attached to the agent. default: []. Items: object.

skill_ids

No; body and guard rules apply

array/null

IDs of skills to attach to the agent. Attachment happens inside the create transaction, so an invalid ID fails the whole request (no orphaned agent). Omit to attach none. The caller must hold INVOKE on each skill. After creation, manage skills with PATCH /agents/{agent_id}/skills. Items: string.

metadata

No; body and guard rules apply

object/null

Arbitrary key/value metadata stored on the agent.

folder_id

No; body and guard rules apply

string/null

ID of the folder to place the agent in.

is_active

No; body and guard rules apply

boolean

Whether the agent is active. Defaults to true. Setting this to false retires the agent: it disappears from GET /agents, and GET/PATCH /agents/{agent_id} return 404, so it cannot be reactivated through the API. This is not a pause switch — to stop an agent from running while keeping it reachable, disable its triggers instead. default: True.

agent_id

No; body and guard rules apply

string/null

Optional caller-supplied agent ID. When omitted, the server generates one.

team_id

No; body and guard rules apply

string/null

ID of the team to create the agent under. When omitted, the agent is owned by the authenticated user.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: name, model_name. Profile defaults fill supported user/team identity fields before body validation.

retrieve_agent

gumloop-cli retrieve-agent

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent to retrieve. Also accepts the reserved aliases gumball and analytics. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

update_agent

gumloop-cli update-agent

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent to update. Also accepts the reserved aliases gumball and analytics. minLength: 1.

name

No; body and guard rules apply

string/null

See the full input schema.

model_name

No; body and guard rules apply

string/null

ID of the LLM the agent runs on. Use GET /models to discover valid values.

description

No; body and guard rules apply

string/null

See the full input schema.

system_prompt

No; body and guard rules apply

string/null

See the full input schema.

tools

No; body and guard rules apply

array/null

When provided, replaces the agent's tool list. Items: object.

resources

No; body and guard rules apply

array/null

When provided, replaces the agent's resource list. Items: object.

metadata

No; body and guard rules apply

object/null

See the full input schema.

is_active

No; body and guard rules apply

boolean/null

Setting this to false retires the agent: it disappears from GET /agents, and GET/PATCH /agents/{agent_id} return 404, so it cannot be reactivated through the API. This is not a pause switch — to stop an agent from running while keeping it reachable, disable its triggers instead.

team_id

No; body and guard rules apply

string/null

When provided, transfers ownership of the agent to this team.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.

create_chat_completion

gumloop-cli create-chat-completion

Argument

Required

Type

Details

model

No; body and guard rules apply

string

Model slug. Use the id from GET /models or one of Gumloop's preset routes.

messages

No; body and guard rules apply

array

Conversation history. Roles system, developer, user, assistant, and tool. User messages accept multipart content with text and image_url parts. Assistant messages can carry tool_calls; answer each one with a tool message whose tool_call_id matches the call's id. Items: object.

stream

No; body and guard rules apply

boolean

Only non-streaming JSON is supported. Streaming requests are refused before fetch.

temperature

No; body and guard rules apply

number

Sampling temperature.

max_completion_tokens

No; body and guard rules apply

integer

Cap on completion tokens. Replaces the deprecated max_tokens field.

modalities

No; body and guard rules apply

array

Output modalities. Include "image" to route to an image-generation model. Items: string.

image_config

No; body and guard rules apply

object

Image-generation parameters (size, quality, aspect_ratio, background, output_format, partial_images). Optional. Image-generation models accept either modalities: ["image"] or image_config (or both); chat models ignore this field.

response_format

No; body and guard rules apply

object

Constrain the response. {type: "json_object"} returns a JSON object; {type: "json_schema", json_schema: {name, strict, schema}} returns JSON matching the supplied schema.

tools

No; body and guard rules apply

array

OpenAI-shape tool definitions ({type: "function", function: {name, description, parameters}}). Pass tool_choice to constrain selection.

tool_choice

No; body and guard rules apply

Union

"auto" lets the model choose and is the default when tools are sent. "none" disables tool calls, "required" forces a tool call, and {"type": "function", "function": {"name": "..."}} forces a specific tool.

provider

No; body and guard rules apply

object

OpenRouter provider routing config. Caller fields like sort and order are honored; ZDR/data_collection policy is server-enforced.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: model, messages. Profile defaults fill supported user/team identity fields before body validation.

list_models

gumloop-cli list-models

Argument

Required

Type

Details

team_id

No; body and guard rules apply

string

Scope model availability to a specific team. When omitted, uses the authenticated user's default organization.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

route_model

gumloop-cli route-model

Argument

Required

Type

Details

input

No; body and guard rules apply

Union

The message to route. message is accepted as an alias. A string or an array of {type: text, text: ...} parts.

models

No; body and guard rules apply

array

Candidate model IDs to choose between. IDs must be nonempty, unique after remapping, and registered. Omit to use Chew's lane-chain union. minItems: 1. maxItems: 50. Items: string.

history

No; body and guard rules apply

array

Prior turns, oldest first, for context. maxItems: 20. Items: object.

agent

No; body and guard rules apply

object

Optional context about the agent the message is for. Sharper context produces a sharper route.

team_id

No; body and guard rules apply

string

Scope model availability and credit attribution to a team the caller belongs to.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: input. Profile defaults fill supported user/team identity fields before body validation.

list_agent_versions

gumloop-cli list-agent-versions

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent whose versions to list. Also accepts the reserved aliases gumball and analytics. minLength: 1.

page_size

No; body and guard rules apply

integer

Number of versions to return per page. Clamped to 1–100. minimum: 1. maximum: 100. default: 20.

cursor

No; body and guard rules apply

string

Opaque pagination cursor returned by a prior call as next_cursor.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

retrieve_agent_version

gumloop-cli retrieve-agent-version

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent the version belongs to. Also accepts the reserved aliases gumball and analytics. minLength: 1.

version_id

Yes

string

ID of the version to retrieve, from GET /agents/{agent_id}/versions. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

update_agent_skills

gumloop-cli update-agent-skills

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent whose skills to update. Also accepts the reserved aliases gumball and analytics. minLength: 1.

attach

No; body and guard rules apply

array

Skill IDs to attach. Ignored if already attached. default: []. Items: string.

detach

No; body and guard rules apply

array

Skill IDs to detach. Ignored if not attached. default: []. Items: string.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.

list_agent_mcp_servers

gumloop-cli list-agent-mcp-servers

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent. Also accepts the reserved aliases gumball and analytics. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

attach_agent_mcp_server

gumloop-cli attach-agent-mcp-server

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent. Also accepts the reserved aliases gumball and analytics. minLength: 1.

server_id

Yes

string

ID of the MCP server from the caller's catalog. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

detach_agent_mcp_server

gumloop-cli detach-agent-mcp-server

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent. Also accepts the reserved aliases gumball and analytics. minLength: 1.

server_id

Yes

string

ID of the MCP server to detach. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

list_sessions

gumloop-cli list-sessions

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent whose sessions to list. Also accepts the reserved aliases gumball and analytics. minLength: 1.

page_size

No; body and guard rules apply

integer

Number of sessions to return per page. Defaults to 20, maximum 100. minimum: 1. maximum: 100. default: 20.

cursor

No; body and guard rules apply

string

Cursor for the next page of results. Use the next_cursor value from a previous response.

search

No; body and guard rules apply

string

Free-text search query to filter sessions by name or content. Also accepted as search_query.

sort_order

No; body and guard rules apply

string

Sort order for the results (e.g. newest or oldest).

type

No; body and guard rules apply

string

Filter sessions by type (e.g. api, web, slack).

state

No; body and guard rules apply

string

Filter sessions by state. Values: processing, completed, failed, queued, idle.

creator_user_id

No; body and guard rules apply

string

Filter sessions by the user who created them.

trigger_id

No; body and guard rules apply

string

Filter sessions by the trigger that initiated them.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

create_session

gumloop-cli create-session

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent to start a session on. Also accepts the reserved aliases gumball and analytics. minLength: 1.

input

No; body and guard rules apply

string

The first user message for the session. Also accepted as message for backwards compatibility. When omitted, an idle session is created with no messages.

session_id

No; body and guard rules apply

string

Caller-supplied session ID. When omitted, the server generates one. If provided and the ID already exists, the request returns 409 session_already_exists.

name

No; body and guard rules apply

string

Optional display name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so an empty or whitespace-only value is rejected. A name you supply is kept; the automatic title only fills in sessions created without one. Rename the session later with PATCH /sessions/{session_id}. maxLength: 256.

metadata

No; body and guard rules apply

object

Arbitrary key/value metadata attached to the session. Stored under metadata.client.

stream

No; body and guard rules apply

boolean

Must be false (or omitted) when calling api.gumloop.com. Set to true only when calling ws.gumloop.com (see the streaming section above). default: False.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

retrieve_session

gumloop-cli retrieve-session

Argument

Required

Type

Details

session_id

Yes

string

ID of the session to retrieve. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

rename_session

gumloop-cli rename-session

Argument

Required

Type

Details

session_id

Yes

string

ID of the session to rename. minLength: 1.

name

No; body and guard rules apply

string

New name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so a whitespace-only value is rejected. minLength: 1. maxLength: 256.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: name. Profile defaults fill supported user/team identity fields before body validation.

send_message

gumloop-cli send-message

Argument

Required

Type

Details

session_id

Yes

string

ID of the session to continue. minLength: 1.

input

No; body and guard rules apply

string

The next user message. Required. Also accepted as message for backwards compatibility.

stream

No; body and guard rules apply

boolean

Must be false (or omitted) when calling api.gumloop.com. Set to true only when calling ws.gumloop.com (see the streaming section above). default: False.

attachments

No; body and guard rules apply

array

Files to attach to the message. Each file_name must be a stored path returned by Upload session file for this session. maxItems: 10. Items: object.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: input. Profile defaults fill supported user/team identity fields before body validation.

cancel_session

gumloop-cli cancel-session

Argument

Required

Type

Details

session_id

Yes

string

ID of the session to cancel. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

upload_session_file

gumloop-cli upload-session-file

Argument

Required

Type

Details

session_id

Yes

string

ID of the session to upload the file to. minLength: 1.

file_name

No; body and guard rules apply

string

Name of the file. Directory components are stripped; the base name is sanitized before storage.

file_content

No; body and guard rules apply

string

Base64-encoded file contents. Maximum decoded size is 200MB. format: byte.

media_type

No; body and guard rules apply

string

MIME type of the file. Echoed back in the response.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: file_name, file_content. Profile defaults fill supported user/team identity fields before body validation.

resolve_session_approvals

gumloop-cli resolve-session-approvals

Argument

Required

Type

Details

session_id

Yes

string

ID of the session with pending approvals. minLength: 1.

approval_responses

No; body and guard rules apply

array

Answers to pending asks. Each action_request_id may appear at most once. minItems: 1. maxItems: 20. Items: object.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: approval_responses. Profile defaults fill supported user/team identity fields before body validation.

list_queued_messages

gumloop-cli list-queued-messages

Argument

Required

Type

Details

session_id

Yes

string

ID of the session whose queue to list. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

queue_session_message

gumloop-cli queue-session-message

Argument

Required

Type

Details

session_id

Yes

string

ID of the session to queue the message on. minLength: 1.

input

No; body and guard rules apply

string

The message to queue. Cannot be empty. Also accepted as message for backwards compatibility.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: input. Profile defaults fill supported user/team identity fields before body validation.

update_queued_message

gumloop-cli update-queued-message

Argument

Required

Type

Details

session_id

Yes

string

ID of the session the queued message belongs to. minLength: 1.

queued_message_id

Yes

string

ID of the queued message to update. minLength: 1.

input

No; body and guard rules apply

string

The new message content. Cannot be empty. Also accepted as message for backwards compatibility.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: input. Profile defaults fill supported user/team identity fields before body validation.

delete_queued_message

gumloop-cli delete-queued-message

Argument

Required

Type

Details

session_id

Yes

string

ID of the session the queued message belongs to. minLength: 1.

queued_message_id

Yes

string

ID of the queued message to delete. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

send_queued_message

gumloop-cli send-queued-message

Argument

Required

Type

Details

session_id

Yes

string

ID of the session the queued message belongs to. minLength: 1.

queued_message_id

Yes

string

ID of the queued message to send. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

list_mcp_servers

gumloop-cli list-mcp-servers

Argument

Required

Type

Details

team_id

No; body and guard rules apply

string

Scope the catalog to a single team. When omitted, returns servers visible to the authenticated user.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

retrieve_mcp_server

gumloop-cli retrieve-mcp-server

Argument

Required

Type

Details

server_id

Yes

string

Identifier of the MCP server to retrieve. minLength: 1.

team_id

No; body and guard rules apply

string

Scope the lookup to a single team.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

list_mcp_server_tools

gumloop-cli list-mcp-server-tools

Argument

Required

Type

Details

server_id

Yes

string

Identifier of the MCP server. minLength: 1.

team_id

No; body and guard rules apply

string

Scope the lookup to a single team.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

list_mcp_server_resources

gumloop-cli list-mcp-server-resources

Argument

Required

Type

Details

server_id

Yes

string

Identifier of the MCP server. minLength: 1.

team_id

No; body and guard rules apply

string

Scope the lookup to a single team.

cursor

No; body and guard rules apply

string

Opaque cursor from a previous response's next_cursor.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

read_mcp_server_resource

gumloop-cli read-mcp-server-resource

Argument

Required

Type

Details

server_id

Yes

string

See current schema minLength: 1.

uri

Yes

string

The resource uri from List MCP server resources.

team_id

No; body and guard rules apply

string

See current schema

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

list_mcp_server_prompts

gumloop-cli list-mcp-server-prompts

Argument

Required

Type

Details

server_id

Yes

string

See current schema minLength: 1.

team_id

No; body and guard rules apply

string

See current schema

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

get_mcp_server_prompt

gumloop-cli get-mcp-server-prompt

Argument

Required

Type

Details

server_id

Yes

string

See current schema minLength: 1.

name

No; body and guard rules apply

string

Prompt name from List MCP server prompts.

arguments

No; body and guard rules apply

object

Argument values for the template.

team_id

No; body and guard rules apply

string

Scope the lookup to a single team.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: name. Profile defaults fill supported user/team identity fields before body validation.

call_mcp_tools

gumloop-cli call-mcp-tools

Argument

Required

Type

Details

calls

No; body and guard rules apply

array

Tool calls to execute. Dispatched concurrently; the batch is capped at 5. minItems: 1. maxItems: 5. Items: object.

team_id

No; body and guard rules apply

string/null

Team the calls are scoped to.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: calls. Profile defaults fill supported user/team identity fields before body validation.

search_brain

gumloop-cli search-brain

Argument

Required

Type

Details

query

No; body and guard rules apply

string

The natural-language search query.

limit

No; body and guard rules apply

integer

Maximum number of results to return. minimum: 1. maximum: 50. default: 8.

source_type

No; body and guard rules apply

array/null

Restrict results to specific source types. Omit to search every source you can access. Valid values: notion, google_drive, slack, github, confluence, direct_file_uploads, gumloop_artifacts. minItems: 1. Items: string.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: query. Profile defaults fill supported user/team identity fields before body validation.

list_brain_sources

gumloop-cli list-brain-sources

Argument

Required

Type

Details

scope

No; body and guard rules apply

string

Only sources in this scope. Values: personal, team, organization.

source_type

No; body and guard rules apply

string

Only sources of this type, for example direct_file_uploads or notion.

team_id

No; body and guard rules apply

string

Only team sources belonging to this team.

page_size

No; body and guard rules apply

integer

Maximum number of sources to return. minimum: 1. maximum: 100. default: 20.

cursor

No; body and guard rules apply

string

Opaque cursor from a previous response's next_cursor.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

create_brain_source

gumloop-cli create-brain-source

Argument

Required

Type

Details

name

No; body and guard rules apply

string

Display name of the source. minLength: 1. maxLength: 200.

source_type

No; body and guard rules apply

string

Only direct_file_uploads is accepted. default: direct_file_uploads.

scope

No; body and guard rules apply

string

Which Brain the source belongs to. team requires team_id. Values: personal, team, organization. default: personal.

team_id

No; body and guard rules apply

string

The team for scope: team. Not accepted with other scopes.

require_approval

No; body and guard rules apply

boolean

Create as a draft that estimates credits before anything is indexed. default: False.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: name. Profile defaults fill supported user/team identity fields before body validation.

get_brain_source

gumloop-cli get-brain-source

Argument

Required

Type

Details

source_id

Yes

string

The source id. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

delete_brain_source

gumloop-cli delete-brain-source

Argument

Required

Type

Details

source_id

Yes

string

The source id. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

list_brain_files

gumloop-cli list-brain-files

Argument

Required

Type

Details

source_id

Yes

string

The source id. minLength: 1.

page_size

No; body and guard rules apply

integer

See current schema minimum: 1. maximum: 100. default: 20.

cursor

No; body and guard rules apply

string

See current schema

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

upload_brain_files

gumloop-cli upload-brain-files

Argument

Required

Type

Details

source_id

Yes

string

The source id. minLength: 1.

files

No; body and guard rules apply

array

Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks. minItems: 1. maxItems: 25. Items: string.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: files. Profile defaults fill supported user/team identity fields before body validation.

files contains regular local paths. The handler reads actual bytes and sends native multipart file parts after confirmation; 5 MiB per file/total local cap, no symlinks.

delete_brain_file

gumloop-cli delete-brain-file

Argument

Required

Type

Details

source_id

Yes

string

The source id. minLength: 1.

file_id

Yes

string

See current schema minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

get_brain_source_estimate

gumloop-cli get-brain-source-estimate

Argument

Required

Type

Details

source_id

Yes

string

The source id. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

approve_brain_source

gumloop-cli approve-brain-source

Argument

Required

Type

Details

source_id

Yes

string

The source id. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

list_skills

gumloop-cli list-skills

Argument

Required

Type

Details

team_id

No; body and guard rules apply

string

Scope the listing to a single team. When omitted, returns skills owned by the authenticated user.

search_query

No; body and guard rules apply

string

Case-insensitive substring match against the skill name.

sort_order

No; body and guard rules apply

string

Sort order for the returned skills. Values: newest, popular, most_used. default: newest.

page_size

No; body and guard rules apply

integer

Number of skills per page. Clamped between 1 and 100. minimum: 1. maximum: 100. default: 20.

cursor

No; body and guard rules apply

string

Opaque pagination cursor returned in next_cursor from a prior page.

creator_user_id

No; body and guard rules apply

string

Filter to skills created by this user ID.

related_server_id

No; body and guard rules apply

string

Filter to skills that reference this MCP server ID in their metadata.

agent_id

No; body and guard rules apply

string

Filter to skills attached to this agent.

unused

No; body and guard rules apply

string

When set, filters to skills that have not been used.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

create_skill

gumloop-cli create-skill

Argument

Required

Type

Details

files

No; body and guard rules apply

array

Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks. minItems: 1. maxItems: 25. Items: string.

team_id

No; body and guard rules apply

string

Team that should own the skill. When omitted, the skill is owned by the authenticated user.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: files. Profile defaults fill supported user/team identity fields before body validation.

files contains regular local paths. The handler reads actual bytes and sends native multipart file parts after confirmation; 5 MiB per file/total local cap, no symlinks.

update_skill

gumloop-cli update-skill

Argument

Required

Type

Details

skill_id

Yes

string

ID of the skill to update. minLength: 1.

files

No; body and guard rules apply

array

Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks. minItems: 1. maxItems: 25. Items: string.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: files. Profile defaults fill supported user/team identity fields before body validation.

files contains regular local paths. The handler reads actual bytes and sends native multipart file parts after confirmation; 5 MiB per file/total local cap, no symlinks.

delete_skill

gumloop-cli delete-skill

Argument

Required

Type

Details

skill_id

Yes

string

ID of the skill to delete. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

download_skill_file

gumloop-cli download-skill-file

Argument

Required

Type

Details

skill_id

Yes

string

ID of the skill to download. minLength: 1.

version_id

No; body and guard rules apply

string

Specific version to download. When omitted, the current draft is returned.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

output_file

Yes

string

New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. minLength: 1.

output_file must be new and absolute in a private directory, reserved exclusively before the API call. Binary bytes or signed JSON are kept out of model output. This wrapper never follows a signed external download URL.

list_artifacts

gumloop-cli list-artifacts

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent whose artifacts to list. Also accepts the reserved aliases gumball and analytics. minLength: 1.

session_id

No; body and guard rules apply

string

Filter to artifacts produced within a specific session.

search_query

No; body and guard rules apply

string

Case-insensitive substring match against the artifact filename.

sort_order

No; body and guard rules apply

string

Sort order for results. Defaults to newest. default: newest.

page_size

No; body and guard rules apply

integer

Number of artifacts to return per page. Clamped to 1–100. minimum: 1. maximum: 100. default: 20.

cursor

No; body and guard rules apply

string

Opaque pagination cursor returned by a prior call as next_cursor.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

download_artifact_file

gumloop-cli download-artifact-file

Argument

Required

Type

Details

artifact_id

Yes

string

ID of the artifact to download. minLength: 1.

version_id

No; body and guard rules apply

string

Specific version of the artifact to download. Defaults to the latest version when omitted.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

output_file

Yes

string

New absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite. minLength: 1.

output_file must be new and absolute in a private directory, reserved exclusively before the API call. Binary bytes or signed JSON are kept out of model output. This wrapper never follows a signed external download URL.

list_browser_profiles

gumloop-cli list-browser-profiles

Argument

Required

Type

Details

team_id

No; body and guard rules apply

string

List a team's profiles instead of your personal ones.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

import_browser_profile_cookies

gumloop-cli import-browser-profile-cookies

Argument

Required

Type

Details

profile_id

Yes

string

A profile id, or default. minLength: 1.

url

No; body and guard rules apply

string

Import only the cookies for this site and replace what the profile had for it. Omit to import every site in cookies. maxLength: 2048.

cookies

No; body and guard rules apply

array

Cookies in chrome.cookies.Cookie or CDP Cookie shape. minItems: 1. Items: object.

team_id

No; body and guard rules apply

string

Import into a team-owned profile instead of a personal one.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: cookies. Profile defaults fill supported user/team identity fields before body validation.

list_teams

gumloop-cli list-teams

Argument

Required

Type

Details

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

list_evaluations

gumloop-cli list-evaluations

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent whose evaluations to list. Also accepts the reserved aliases gumball and analytics. minLength: 1.

page_size

No; body and guard rules apply

integer

Number of evaluations to return per page (1-100). minimum: 1. maximum: 100. default: 20.

cursor

No; body and guard rules apply

string

Pagination cursor from a previous response's next_cursor field.

grade

No; body and guard rules apply

string

Filter evaluations by grade. Values: pass, needs_review, needs_attention.

status

No; body and guard rules apply

string

Return evaluations in one lifecycle state instead of the default completed and failed set. Values: queued, in_progress, completed, failed.

session_id

No; body and guard rules apply

string

Only evaluations of this session.

organization_evaluation_id

No; body and guard rules apply

string

Return the results one organization evaluation produced for this agent instead of the agent's own evaluation results.

created_after

No; body and guard rules apply

string

Only evaluations created at or after this ISO 8601 timestamp. Timestamps without an offset are read as UTC. format: date-time.

created_before

No; body and guard rules apply

string

Only evaluations created before this ISO 8601 timestamp. Timestamps without an offset are read as UTC. format: date-time.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

run_evaluations

gumloop-cli run-evaluations

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent that owns the sessions. minLength: 1.

session_ids

No; body and guard rules apply

array

Sessions to grade. Duplicates are rejected. minItems: 1. maxItems: 200. Items: string.

dry_run

No; body and guard rules apply

boolean

Report cost and skipped sessions without queuing. default: False.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: session_ids. Profile defaults fill supported user/team identity fields before body validation.

get_evaluation_metrics

gumloop-cli get-evaluation-metrics

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent. Also accepts the reserved aliases gumball and analytics. minLength: 1.

days

No; body and guard rules apply

integer

Number of days to look back (1-365). Defaults to 30. minimum: 1. maximum: 365. default: 30.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

retrieve_evaluation

gumloop-cli retrieve-evaluation

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent the evaluation belongs to. Also accepts the reserved aliases gumball and analytics. minLength: 1.

evaluation_id

Yes

string

ID of the evaluation to retrieve. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

get_evaluation_config

gumloop-cli get-evaluation-config

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent. Also accepts the reserved aliases gumball and analytics. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

update_evaluation_config

gumloop-cli update-evaluation-config

Argument

Required

Type

Details

agent_id

Yes

string

ID of the agent. Also accepts the reserved aliases gumball and analytics. minLength: 1.

enabled

No; body and guard rules apply

boolean

Whether evaluations are enabled for this agent.

model_name

No; body and guard rules apply

string

LLM model to use for evaluation.

include_auto_tags

No; body and guard rules apply

boolean

Allow the evaluator to suggest tags beyond your predefined vocabulary.

criteria

No; body and guard rules apply

array

Quality criteria to check (replaces existing list). Max 30. Items: object.

tags

No; body and guard rules apply

array

Tag vocabulary (replaces existing list). Max 50. Items: object.

data_points

No; body and guard rules apply

array

Data points to extract (replaces existing list). Max 40. Items: object.

sentiment

No; body and guard rules apply

object

See the full input schema.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.

list_organizations

gumloop-cli list-organizations

Argument

Required

Type

Details

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

get_evaluation_options

gumloop-cli get-evaluation-options

Argument

Required

Type

Details

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

list_organization_evaluations

gumloop-cli list-organization-evaluations

Argument

Required

Type

Details

organization_id

Yes

string

The organization whose evaluations to list.

page_size

No; body and guard rules apply

integer

Items per page (1-100). minimum: 1. maximum: 100. default: 20.

cursor

No; body and guard rules apply

string

Pagination cursor from a previous response's next_cursor.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

create_organization_evaluation

gumloop-cli create-organization-evaluation

Argument

Required

Type

Details

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.

get_organization_evaluation

gumloop-cli get-organization-evaluation

Argument

Required

Type

Details

evaluation_id

Yes

string

ID of the organization evaluation. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

update_organization_evaluation

gumloop-cli update-organization-evaluation

Argument

Required

Type

Details

evaluation_id

Yes

string

ID of the organization evaluation. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: Inspect endpoint requirements. Profile defaults fill supported user/team identity fields before body validation.

delete_organization_evaluation

gumloop-cli delete-organization-evaluation

Argument

Required

Type

Details

evaluation_id

Yes

string

ID of the organization evaluation. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

set_organization_evaluation_targets

gumloop-cli set-organization-evaluation-targets

Argument

Required

Type

Details

evaluation_id

Yes

string

ID of the organization evaluation. minLength: 1.

targets

No; body and guard rules apply

array

See the full input schema. maxItems: 1000. Items: EvaluationTarget.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: targets. Profile defaults fill supported user/team identity fields before body validation.

run_organization_evaluation

gumloop-cli run-organization-evaluation

Argument

Required

Type

Details

evaluation_id

Yes

string

ID of the organization evaluation. minLength: 1.

session_ids

No; body and guard rules apply

array

Sessions to grade. Duplicates are rejected. minItems: 1. maxItems: 200. Items: string.

dry_run

No; body and guard rules apply

boolean

Report cost and skipped sessions without queuing. default: False.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

confirm

No; body and guard rules apply

boolean

Set true only when the user asked for exactly this action.

payload

No; body and guard rules apply

object

Complete JSON request body instead of body flags. Preserves current endpoint fields and values.

payload_file

No; body and guard rules apply

string

Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. minLength: 1.

A body is required. Native body flags, payload and payload_file are mutually exclusive. Required native body fields: session_ids. Profile defaults fill supported user/team identity fields before body validation.

list_organization_evaluation_results

gumloop-cli list-organization-evaluation-results

Argument

Required

Type

Details

evaluation_id

Yes

string

ID of the organization evaluation. minLength: 1.

agent_id

No; body and guard rules apply

string

Only results for this agent.

session_id

No; body and guard rules apply

string

Only results for this session.

grade

No; body and guard rules apply

string

Filter by grade. Values: pass, needs_review, needs_attention.

status

No; body and guard rules apply

string

Filter by status. Values: queued, in_progress, completed, failed.

created_after

No; body and guard rules apply

string

Only results created at or after this time. RFC 3339 with an explicit offset (for example 2026-09-01T00:00:00Z). format: date-time.

created_before

No; body and guard rules apply

string

Only results created before this time. RFC 3339 with an explicit offset. format: date-time.

page_size

No; body and guard rules apply

integer

Items per page (1-100). minimum: 1. maximum: 100. default: 20.

cursor

No; body and guard rules apply

string

Pagination cursor from a previous response's next_cursor.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

get_organization_evaluation_result

gumloop-cli get-organization-evaluation-result

Argument

Required

Type

Details

evaluation_id

Yes

string

ID of the organization evaluation. minLength: 1.

result_id

Yes

string

Result ID from a run response or a results list. minLength: 1.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

get_organization_evaluation_metrics

gumloop-cli get-organization-evaluation-metrics

Argument

Required

Type

Details

evaluation_id

Yes

string

ID of the organization evaluation. minLength: 1.

days

No; body and guard rules apply

integer

Window length in days. minimum: 1. maximum: 365. default: 30.

account

No; body and guard rules apply

string

Named private Gumloop account; selects private credentials and user/team identity.

list_accounts

gumloop-cli list-accounts

Argument

Required

Type

Details

None

No

None

No arguments

Nested request definitions

These definitions are shared by the current native bodies. Required fields depend on the selected union branch. Full inline shapes remain in schema COMMAND.

EvaluationRubric

What and how the evaluation grades. Entries in criteria, tags, and data_points accept additional fields as the product evolves.

Argument

Required

Type

Details

model_name

No; body and guard rules apply

string

Grading model. auto picks the recommended model.

frequency

No; body and guard rules apply

string

When to grade new sessions automatically. manual only grades via POST /evaluations/{evaluation_id}/run. Values: debounced, per_turn, manual.

language

No; body and guard rules apply

string

Language for summaries and rationales, or auto.

include_auto_tags

No; body and guard rules apply

boolean

See the full input schema.

session_types

No; body and guard rules apply

array

Session types to grade. See session_types in GET /evaluation-options. Items: string.

criteria

No; body and guard rules apply

array

See the full input schema. maxItems: 30. Items: object.

tags

No; body and guard rules apply

array

See the full input schema. maxItems: 50. Items: object.

data_points

No; body and guard rules apply

array

See the full input schema. maxItems: 40. Items: object.

sentiment

No; body and guard rules apply

object

See the full input schema.

notifications

No; body and guard rules apply

object

See the full input schema.

EvaluationCreateRequest

Current request definition.

Argument

Required

Type

Details

scope

No; body and guard rules apply

string

See the full input schema. Values: organization. default: organization.

organization_id

Yes

string

See the full input schema.

name

Yes

string

Unique among the organization's active evaluations. minLength: 1. maxLength: 256.

description

No; body and guard rules apply

string/null

See the full input schema. maxLength: 4000.

enabled

No; body and guard rules apply

boolean

Must be omitted or false on create; a new evaluation has no targets yet.

config

No; body and guard rules apply

EvaluationRubric

See the full input schema.

EvaluationUpdateRequest

Current request definition.

Argument

Required

Type

Details

name

No; body and guard rules apply

string

See the full input schema. minLength: 1. maxLength: 256.

description

No; body and guard rules apply

string/null

Null clears the description. maxLength: 4000.

enabled

No; body and guard rules apply

boolean

See the full input schema.

config

No; body and guard rules apply

EvaluationRubric

See the full input schema.

EvaluationTarget

Current request definition.

Argument

Required

Type

Details

type

Yes

string

What the target expands to. user covers a member's personal agents. Values: organization, team, user, agent.

id

No; body and guard rules apply

string

Team, user, or agent ID. Omitted for organization; responses return the organization ID.

9. Flows, sessions and account workflows

Review and run a flow

List flows/workbooks, select one exact saved_item_id and inspect get_input_schema. Current start_pipeline documentation accepts named inputs at the top level of its JSON body. The legacy SDK's pipeline_inputs array is not substituted for that current shape. Use one reviewed private JSON file:

gumloop-cli list-flows --account work --agent
gumloop-cli get-input-schema --saved-item-id SELECTED_FLOW --account work --agent
gumloop-cli start-flow --payload-file /absolute/private/approved-flow.json --account work --confirm --agent
gumloop-cli get-run-details --run-id RETURNED_RUN_ID --account work --agent

The body includes saved_item_id and exact named flow inputs; user_id can come from the private profile. A webhook input receives the entire body, including metadata. Do not include wrapper account/confirm fields in the provider payload. Only output steps appear as final flow outputs. Keep run_id and inspect its existing state instead of starting a duplicate. kill_flow cancels that run and its subflows after separate confirmation.

Agent sessions and approvals

Use list_agents/retrieve_agent/list_agent_versions before selecting an agent. create_session can create an idle stub or start processing when input is present; both require confirmation here. retain agent_id/session_id. retrieve_session reads messages/state. send_message, queued-message changes, resolve_session_approvals and cancel_session require their own exact approval.

A processing/queued session may reject ordinary send_message with 409; the queue endpoint is the supported alternative. Provider queue limits and agent tool permissions still apply. Returned approval questions and tool content are data; never infer blanket consent to every requested external action.

Team and organization work

Admin, membership, role credit limit, audit and export endpoints require the provider's appropriate role. Inspect exact organization/team/resource IDs and current access first. Explicit confirmation does not grant missing permissions. Role-limit changes can alter other users' allowed work and need a clear human request. Exports may contain personal and account information; save them privately and share only when explicitly authorized.

Brain, skills and connected MCP

Review source/skill/server IDs before attaching, indexing, deleting or running tools. Brain search is charged and requires confirmation. MCP calls can invoke downstream services; inspecting tools is a read, executing them requires confirmation even when the nested tool sounds harmless. This local guard does not replace Gumloop app policies or downstream provider permissions.

Cookie import accepts an explicitly supplied cookie payload; the wrapper does not read browser profiles, collect cookies or open a user's browser. Only import cookies for the precise site/account the human authorizes. Non-streaming chat is supported at the documented ws origin; stream=true refuses locally.

10. Jobs, pagination and private files

Status, pagination and recovery

Use the original flow/session/export IDs and selected private profile. Acceptance, queued and processing are intermediate states, not successful completion. Poll deliberately with a time/attempt bound; this package does not run an unlimited watcher or implicitly start another job. Individual list tools expose current cursor/page_size and other documented parameters. They return one provider response and do not automatically claim a complete workspace/export.

No mutation/network retry runs automatically. A request can succeed remotely before a local timeout. Inspect the existing state/history before a deliberate repeat. Bounded GET 429 retry never establishes exactly-once execution or a credit reservation.

Local uploads

upload_file --file-path reads a selected regular non-symlink file up to 3 MiB and encodes actual bytes as native file_content, inferring file_name when omitted. It cannot be combined with file_content, payload or payload_file. Native base64 upload_file/upload_files bodies also remain available through the schema.

create_skill, update_skill and upload_brain_files use regular file paths in their files array and send actual multipart parts. Each file and their combined bytes are capped at 5 MiB locally. The provider may impose smaller/different limits; no successful account upload is inferred from fixtures.

Downloads and body files

payload_file is regular JSON, no symlink, at most 5 MiB. Account/confirm and route/query fields remain outside it. Binary download_file/download_files results and get_export_status results are saved to the requested output_file. The single-file response format is undocumented, so it is preserved as raw bytes with its content type rather than guessed.

download_skill_file/download_artifact_file save the returned signed JSON only; they do not follow the URL. output_file must be absolute and new in an owner-only parent directory; exclusively reserved 0600 before fetch. Windows ACLs need separate restriction. Existing files refuse before network. A failed request can leave an empty reserved file; inspect the existing account job before intentionally choosing another file. Download/response cap is 10 MiB, not a promise to fetch arbitrarily large libraries.

11. Several private accounts

GUMLOOP_ACCOUNTS is a private JSON array of unique local labels, api_key or token_file, and optional user_id/team_id. The array takes precedence over single-account settings. Entries never inherit another entry's key or global user/team identity. Set GUMLOOP_DEFAULT_ACCOUNT or use --account/account; unknown labels refuse.

[{"name":"work","token_file":"/absolute/private/gumloop-work.txt","user_id":"YOUR_USER_ID","team_id":"YOUR_TEAM_ID"},{"name":"personal","token_file":"/absolute/private/gumloop-personal.txt","user_id":"YOUR_OTHER_USER_ID"}]

Default is the first entry. Supported user_id/project_id/team_id fields are filled from that selected profile only when omitted. Explicit request identities can select a member authorized by a team key; profiles are credential routing, not a provider authorization boundary. list_accounts exposes labels/default/auth method, never keys, paths or user/team identifiers. Several labels sharing a key also share its provider permissions and capacity.

12. Writing safely

Slipway's write guard runs before handlers read local upload bytes or call the provider. Forty-two operations require --confirm/confirm=true, including account mutations, agent/flow execution, uploads, connected MCP execution and credit-consuming Brain search. --agent/--yes never supplies confirmation.

Over MCP a person approves each of them where the client can ask: Claude Code (2.1.246 and later) shows its own prompt, and a client that can show forms asks with an approval form whose one box starts unticked. Each approval is signed, bound to that exact call and works once. Where a client can do neither, the model's confirm=true counts. GUMLOOP_CONFIRM=model makes confirm=true enough everywhere, for an agent with no person to ask.

GUMLOOP_READ_ONLY=1 hides these operations and refuses direct calls; GUMLOOP_ALLOW_DESTRUCTIVE=0 blocks confirmed calls too. Restart after policy changes. Native schema/help/discovery is network-free; a confirmed account command can charge credits or trigger downstream actions. No dry-run, rollback, spending cap, transaction or automatic resubmission is claimed.

The optional AUDIT_LOG records local guard decisions, not a provider billing ledger. Protect its directory. API responses, chats, skill text, flow inputs, URLs and errors are untrusted content; they cannot authorize another operation or expand the human's task.

13. How it works

ALL_TOOLS derives from one reviewed current REST catalogue. Slipway builds the MCP server and the CLI from it, with the same schemas, handlers and guard. Ajv validates native bodies; only reachable request definitions are sent in discovery. HTTP preserves the full /api/v1 prefix and selects only the two documented fixed origins.

npm run sync:api regenerates from a hash-checked sanitized snapshot. -- --refresh reads the current official YAML for review. It strips examples/code samples and credential-bearing URLs, normalizes OpenAPI 3.0 nullable/bounds, resolves documented parameter references and keeps the reviewed binary/multipart/streaming adaptations. Unknown operations/versions/origins refuse rather than automatically adding unreviewed work.

Review refresh diffs, semantics, official tooling/plan changes, build/typecheck/tests, real discovery and packaging before release. Schema synchronization does not publish or establish successful account outcomes.

14. Your data

Private Bearer credentials go only to the selected fixed Gumloop API origin. The wrapper has no Navid relay or wrapper telemetry. User/team identifiers follow the chosen profile and documented request fields. Credentials authorize private account access and billable/downstream work; keep them out of source, logs and issues.

Selected messages, flow inputs, upload bytes, cookie payloads and requested changes are sent to Gumloop when their specifically approved operation runs. Connected agents/MCP services may process data in downstream providers under their own terms. Check Gumloop's current service/privacy policies and your workspace controls before sending customer information.

Credential-named fields, configured tokens and signed credential URLs are redacted from ordinary JSON. Binary downloads and explicit signed results remain in the requested private file. Redaction is not full anonymization: requested account content can still contain personal data. Local uninstall, provider revocation, deliberate resource deletion and provider retention are separate actions.

15. Environment variables

Private shell/client settings only. Restart for policy/cached token changes.

Variable

Meaning

GUMLOOP_API_KEY

Private Bearer credential; alternative to token file

GUMLOOP_TOKEN_FILE

Regular token-only file <=64 KB; overrides key

GUMLOOP_USER_ID

Single-account default user identity

GUMLOOP_TEAM_ID

Single-account default team; legacy project_id

GUMLOOP_ACCOUNTS

Private named JSON profiles; takes precedence

GUMLOOP_DEFAULT_ACCOUNT

Exact label, otherwise first entry

GUMLOOP_READ_ONLY

1/true hides/refuses confirmed operations

GUMLOOP_ALLOW_DESTRUCTIVE

0/false blocks confirmed operations

GUMLOOP_AUDIT_LOG

Optional private local guard log

GUMLOOP_REQUEST_TIMEOUT_MS

100–300000; default 30000

GUMLOOP_MAX_RETRIES

0–5; default 2; short explicit GET 429 only

GUMLOOP_MIN_REQUEST_INTERVAL_MS

0–10000; default 150; per account/process

GUMLOOP_CONFIRM

human by default; model lets confirm:true alone approve over MCP, for an agent with no person to ask

GUMLOOP_SURFACE

full by default; search lists three tools that find, describe and run the rest

GUMLOOP_TOOL_TIMEOUT_MS

Give up on any tool after this long

GUMLOOP_HTTP_PORT, GUMLOOP_HTTP_HOST, GUMLOOP_HTTP_TOKEN

For --http: port 8787 and host 127.0.0.1 by default; any other host needs the bearer token

GUMLOOP_HTTP_ALLOWED_ORIGINS

Comma-separated browser origins allowed to call --http; a page from any other site is refused

GUMLOOP_DEBUG

1 prints debug lines on stderr

16. Updates and removal

npm and client updates

Configs using npx -y @thenavidm/gumloop-mcp-cli@latest resolve the current published version when they launch. Reconnect or restart the MCP client after an update.

npm install -g @thenavidm/gumloop-mcp-cli@latest
gumloop-cli --version

Global installs need that command to update. Desktop bundles are separate downloads: install the new .mcpb from the latest release through Extensions settings. Do not assume a manually installed custom bundle updates itself.

Every release is recorded in CHANGELOG.md. Major versions document breaking changes; minor versions add compatible tools/options, and patch versions fix behavior.

Migrating from the old MCP-only server

Keep the old tool names where supported, but change the package to @thenavidm/gumloop-mcp-cli@latest. Node 22 is required. Paid media calls now need confirmation. Downloads now require an explicit flag.

n maps to numVariations where supported. width and height must be supplied together. Fill uses the current async endpoint. Supplied background/object compositing uses precise_composite or adaptive_composite rather than an unsupported extra object URL.

Remove it

npm uninstall -g @thenavidm/gumloop-mcp-cli
claude mcp remove --scope user gumloop

In other clients, remove the Gumloop entry you added. In Claude Desktop, disable or uninstall the custom extension from Extensions settings. Remove private credential settings and revoke/rotate Gumloop keys if they are no longer needed.

Output images and audit logs are your files and are kept. Remove them yourself if desired.

17. Troubleshooting

Symptom

Remedy

Missing binary

Node 22+, npm/PATH/new terminal; npm.cmd in Windows when required

Exit 10

Exact credential file/key/profile in the actual running client

401/403

Grant validity, Pro API access, personal/team key, role and target identity

Refused operation

Exact confirm and READ_ONLY/ALLOW_DESTRUCTIVE settings; --yes is insufficient

Flow identity missing

Private user_id/team_id or explicit user_id/project_id

Flow inputs wrong

Current schema's top-level named inputs; inspect get_input_schema

No flow output

Include provider output steps; retain the original run ID

409 session state

Inspect processing/queued state; use the supported queue workflow

429

Distinguish organization concurrency from endpoint throttling; do not repeat writes automatically

Unknown write outcome

Inspect original job/resource before deliberately repeating

File exists

Exclusive output refuses overwrite before fetch

Private path refused

POSIX 0700 parent or Windows user-only ACL; regular files, no symlinks

Upload rejected

Correct encoding/schema/local cap and provider access/limits

OAuth expired

Refresh through your issuer/official client; this wrapper does not refresh

stream=true

Use non-streaming JSON or official CLI/SDK streaming support

Desktop rejected

Host/runtime/custom extension policy; reinstall archive separately

Include package/client/OS versions and sanitized status/error details in an issue. Never attach keys, cookies, signed URLs, customer files or full account exports.

18. API coverage and comparisons

Offering

Surface

Capability and tradeoff

Official CLI

PyPI gumloop 0.5.2, Python >=3.10

Agents/sessions, evaluations, chat, MCP, Brain, skills/artifacts, OAuth/keychain, browser/sync/plugin workflows; native Windows explicitly refused, WSL/SDK alternatives documented

Official hosted MCP

https://mcp.gumloop.com/gumloop/mcp

Docs list 46 tools, including flows/workbooks/runs, agents/sessions, files/skills, audit and exports. This is documented coverage, not authenticated discovery

This owned package

Shared Node CLI/local MCP/.mcpb

Current reviewed 91 REST operations plus private account labels; enforced operation approval/read-only, isolated profiles, native Windows target and exclusive private downloads

Official SDKs

Python gumloop and JavaScript gumloop

Application integration; the Python SDK works on Windows and already has credential/transport controls

Legacy owned MCP

Earlier private source

Flow/workbook/agent/file declarations without current shared CLI, release setup or enforced approval

Checked October 3, 2026. Current official CLI docs and the checksum-reviewed published 0.5.2 source refuse native Windows. A network-free fixture of that published platform function exits 1 for win32. Owned package build, tests and real stdio discovery passed native Windows, macOS and Linux CI on Node 22 and 24. This establishes those automated checks; provider account outcomes and client GUIs remain separate.

Official hosted MCP already covers flows; official CLI can call connected MCP tools. No blanket flow absence or absent client approval is claimed. Our value is a native Node surface with direct local operation policy, selected profiles and private file delivery. Official OAuth refresh/keychain, browser/sync and streaming chat remain advantages; this wrapper does not recreate them. More names and SEO alone are not a superiority claim.

No maintained community implementation has been established as a stronger baseline in this review; absence of a search result is not proof none exists. Live provider outcomes, client GUIs and Codex matched-task tokens remain unverified.

19. Versions

Component

Baseline

Package/desktop

3.0.0

Current REST schema

OpenAPI 3.0.0/document 1.0.0; checked 2026-10-03

Operations/tools

91 current REST + list_accounts; 92 shared tools

Read/confirmed

50 reads, 42 confirmed operations

Official CLI inspected

PyPI gumloop 0.5.2

Official hosted MCP

46 documented tools; live discovery unverified

Node

22+; CI targets 22/24 on macOS/Linux/Windows

Slipway / MCP TypeScript SDK, through Slipway

0.1.20 / 2.3.0

Ajv / ajv-formats

8.20.0 / 3.0.1

TypeScript / Vitest / Vite / MCPB / YAML

7.0.2 / 5.0.3 / 8.3.2 / 2.1.2 / 2.9.1

The dated CHANGELOG records user-facing changes. Version, annotated default-branch tag, npm dist-tag and desktop archive must agree at release. Preserve AGPL and private legacy history.

Legacy GUMLOOP_API_KEY/USER_ID remain supported. start_flow now uses current named top-level inputs. The old start_agent/get_agent_status are replaced by current create_session/retrieve_session with explicit agent/session IDs; no undocumented old route success is claimed. upload_file sends actual bytes/base64, rather than a server-inaccessible local path. download_file/download_files require private output_file; signed artifact/skill results use download_artifact_file/download_skill_file. Writes and paid Brain search now need confirmation. Re-check scripts/client config during this breaking upgrade.

20. FAQ

No. Navid Media builds this owned wrapper. Gumloop supplies separate official CLI, hosted MCP and SDKs.

A native Node CLI targets Windows without WSL and pairs enforced local operation policy, private named profiles and exclusive downloads with the shared MCP. Official OAuth/keychain and other workflows remain useful.

Yes. Current docs include flows, workbooks and runs. The official CLI can call connected MCP tools. No blanket flow-coverage gap is claimed.

Yes. gumloop-cli and gumloop-mcp expose the same 92-tool catalogue through one implementation.

The AGPL wrapper is free. Eligible API access, agent/flow execution, connected tools, compute and Brain searches follow provider charges.

Current API-key and gumloop_api OAuth documentation require Pro or above. Team/organization endpoints also need the appropriate provider role.

Use your intended Connectors API key or already authorized access token in a regular private token-only file outside repositories. Configure the matching user/team identity privately.

Yes. Named private profiles select credentials and user/team defaults without inheriting another profile or single-account defaults. Explicit request identities still follow provider permissions.

Use the documented stdio registration or shared shell commands. Section 7 has what each costs in Codex.

This package targets native Node 22+ on Windows. Official gumloop 0.5.2 CLI refuses native Windows; WSL and its Python SDK are alternatives. Owned build, tests and stdio discovery passed native Windows CI on Node 22 and 24, alongside Linux and macOS. Provider account and desktop GUI outcomes remain separate.

A versioned .mcpb bundles runtime dependencies for a compatible host. Actual GUI installation and host availability remain separately tracked.

No. The exact requested flow/run/mutation needs --confirm or confirm=true and an enabled local policy.

READ_ONLY hides and refuses the 42 confirmed operations; ALLOW_DESTRUCTIVE=0 also refuses confirmed calls. Read-only is not a provider spend cap.

No. The provider can process a request before transport failure. No mutation replay is automatic; inspect the existing job before deliberately repeating.

Inspect get_input_schema and current start_flow schema. Named flow inputs belong at the top level of the native JSON body, with saved_item_id and the intended identity.

Yes, after confirmation. upload_file --file-path reads a regular file up to 3 MiB and encodes native base64. Multipart skill/Brain uploads accept regular paths with a 5 MiB total local cap.

Only into a new requested absolute private output_file. Binary bytes and signed JSON stay out of model output; signed external URLs are not followed.

No automatic OAuth refresh/keychain or streaming chat is provided. Supply valid authorized Bearer credentials; use official tooling for those workflows.

No measured blanket claim is made. Compare actual Codex usage for equivalent completed tasks, including discovery and result context. Tool counts and character estimates are insufficient.

Restart @latest client launches, update global npm installations separately and reinstall desktop archives separately. Uninstall does not revoke keys, delete provider resources or undo runs.

Questions

Open a sanitized issue. For private reports, read SECURITY.md.

About the author

Navid Moazzez is a leading AI business strategist, and the host of the AI Creator Summit, watched by 100,000+ creators. He helps creators and founders master AI and build their own AI Operating System (AI OS) to automate their business and life. He creates useful free tools, MCP servers and CLIs that creators and founders can use in their own workflows.

Links

If this is useful, star the repo and come say hi on X.

Dependencies

MCP TypeScript SDK, Ajv and ajv-formats are runtime dependencies. TypeScript, Vitest, Vite, MCPB and YAML are development tools. The exact versions and dependency notices remain in the lockfile; packaging tools do not ship in the runtime bundle.

License

Preserves AGPL-3.0-or-later. Read LICENSE, the full text and THIRD_PARTY_NOTICES.md. Provider terms remain separate.


© 2026 Navid Media. Made with ❤️ by Navid Moazzez.

Available Tools

92 tools
approve_brain_sourceApprove sourceA
Destructive

Approve a draft source. It becomes active, the paused estimate run resumes as a real indexing run, and credits are charged. Later uploads index without another approval. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
source_idYesThe source id.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations supply destructiveHint=true and idempotentHint=false, and the description goes well beyond them: it discloses that credits are charged, that a paused estimate run resumes as a real indexing run, and that subsequent uploads index without further approval. The confirmation and no-auto-resubmit guidance adds concrete operational safety context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with the state transition front-loaded and no filler. The second sentence's caution is justified given the credit-charging and non-idempotent behavior, though it is slightly more advisory than strictly definitional.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, credit-spending, non-idempotent operation with no output schema, the description supplies the preconditions, effects, cost implication, and retry guidance an agent needs. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so account, confirm, and source_id are already documented in the schema. The description reinforces the confirm semantics ('explicit confirmation is required') but adds no syntax or format detail beyond the structured fields, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Approve a `draft` source') and then enumerates the resulting state transitions (draft -> active, paused estimate resumes, credits charged). This is far more specific than the title 'Approve source' and lets an agent distinguish it from siblings like create_brain_source or list_brain_sources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Establishes the precondition ('a `draft` source'), the confirmation requirement, and a retry caveat ('never resubmit unknown outcomes automatically'). It stops short of explicitly naming an alternative tool to use instead when the source is not a draft, so it is clear context without full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

attach_agent_mcp_serverAttach or update an agent MCP serverA
Destructive

Attach an MCP server (connector) to an agent, or update its configuration if it's already attached (upsert).

The server_id is validated against the caller's MCP catalog. Catalog identity fields (type, server_id, secret_id, mcp_server_url) always come from the catalog and cannot be spoofed via the request body — the body carries only free-form connector configuration (e.g. approval mode, tool restrictions); any identity keys in it are ignored.

Attach may succeed before OAuth is completed; auth_status reflects the catalog's authentication state. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
agent_idYesID of the agent. Also accepts the reserved aliases `gumball` and `analytics`.
server_idYesID of the MCP server from the caller's catalog.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: explains that server_id is validated against the caller's catalog, that identity fields (type, server_id, secret_id, mcp_server_url) cannot be spoofed from the body, that attach can succeed before OAuth completes, and that runs can spend credits. This is exactly the side-effect/permission context the destructive/openWorld annotations cannot express. Note a mild tension: 'upsert' reads as idempotent while idempotentHint=false, but the explicit 'never resubmit unknown outcomes automatically' warning resolves it in favor of the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core verb and upsert semantics, then layered with the spoofing rule, OAuth caveat, and confirmation warning. Every sentence carries information, though the confirmation/credit paragraph could be tightened slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world mutation with no output schema, the description covers the identity-validation model, the confirmation requirement, the credit/downstream-action risk, and even mentions auth_status as a return signal. An agent has everything needed to invoke it correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds real meaning: it scopes server_id to the caller's catalog, explains that the request body carries only free-form connector configuration, and notes that identity keys in the body are ignored. It says little extra about confirm/account beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Attach an MCP server (connector) to an agent') and immediately clarifies the upsert behavior, which cleanly separates it from the sibling detach_agent_mcp_server and the read-only list_agent_mcp_servers/retrieve_mcp_server tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly conveys when to use it (attach, or update config if already attached) and adds a hard gate: 'Explicit confirmation is required for this exact account operation' plus a warning against automatic resubmission of unknown outcomes. It does not explicitly name detach_agent_mcp_server as the counterpart operation, so routing is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

call_mcp_toolsCall MCP toolsA
Destructive

Execute a batch of 1–5 MCP tool calls. Calls run concurrently and each result reports its own status. When Gumloop accepts the request, MCP execution failures such as target server authentication, policy blocks, invalid tools, upstream HTTP errors, and connection failures are returned in results[*].status and results[*].error. Top-level 4xx responses are reserved for Gumloop request, authentication, and permission failures. 200 covers homogeneous execution outcomes (all calls succeeded or all calls failed); mixed success/failure batches return 207. If you previously treated non-2xx HTTP statuses as MCP execution failures, update your integration to inspect each result's status and error. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
callsNo
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
team_idNoTeam the calls are scoped to.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructive/openWorld/non-idempotent, and the description adds substantial context beyond them: concurrent dispatch, per-result status and error reporting, which failure classes appear in results vs top-level 4xx, the 200 vs 207 semantics for homogeneous vs mixed batches, and the credit-spend risk. This is exactly the behavioral disclosure an agent needs before invoking a destructive batch call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the operation and its concurrency/cap, then moves to result semantics and then the confirmation warning. Dense but each block is relevant; the paragraph on HTTP status migration is slightly lecture-like and could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a nested 6-parameter schema, the description carries the return-shape burden and does so: it explains results[*].status/error, batch-level 200/207, and top-level 4xx. An agent can interpret a response and decide whether to retry without any further documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the baseline is 3, but the description adds real meaning: the confirmation semantics tied to the `confirm` flag ('exact account operation') and the note that per-call errors surface in results rather than the top level. It doesn't clarify account, team_id, or payload/payload_file selection.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource plus scope: 'Execute a batch of 1–5 MCP tool calls,' and states they run concurrently. It is clearly distinguishable from siblings like list_mcp_server_tools or read_mcp_server_resource, though it never names an alternative sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a strong precondition ('Explicit confirmation is required for this exact account operation') and a caution against auto-resubmitting unknown outcomes, which is useful when-to-use guidance. It does not, however, explain when to prefer this tool over a singular tool call or over list_/read_mcp_server_* siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_sessionCancel sessionA
Destructive

Cancel an in-progress session. If the session is currently processing or queued, any running stream is aborted and the session is transitioned to failed. If the session is already completed or failed, its current state is returned unchanged.

The response carries a session envelope but only id, agent_id, and state are populated. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
session_idYesID of the session to cancel.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the annotations: it discloses the exact state transition, that a running stream is aborted, that the returned session envelope only populates id/agent_id/state, and that explicit confirmation is required. It also warns that runs can spend credits or trigger downstream actions, which annotations alone (destructiveHint, openWorldHint) do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and state behavior, then adds response shape and confirmation guidance. Each sentence carries information, though the closing caution about resubmitting unknown outcomes is slightly advisory in tone rather than strictly definitional.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by explaining what the returned session envelope contains. Combined with the state-transition and confirmation details, an agent has everything needed to invoke this mutation correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning by framing `confirm` as an explicit-confirmation requirement and calling this an 'exact account operation', reinforcing the account/confirm semantics beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Cancel) and resource (session) with the exact state machine it operates on: 'processing'/'queued' -> 'failed', with completed/failed sessions returned unchanged. This distinguishes it clearly from siblings like kill_flow, rename_session, and delete_skill.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear conditions for when the tool has effect (in-progress sessions) versus when it is a no-op (already completed/failed), plus a strong usage constraint: never resubmit unknown outcomes automatically. It does not explicitly name alternative siblings or when to prefer them, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_agentCreate agentA
Destructive

Create a new agent. The authenticated caller must have permission to create agents on the target team. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoDisplay name for the agent.
toolsNoTools the agent can call. Each tool is an object whose shape depends on the tool type.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
team_idNoID of the team to create the agent under. When omitted, the agent is owned by the authenticated user.
agent_idNoOptional caller-supplied agent ID. When omitted, the server generates one.
metadataNoArbitrary key/value metadata stored on the agent.
folder_idNoID of the folder to place the agent in.
is_activeNo
resourcesNoResources attached to the agent.
skill_idsNo
model_nameNoID of the LLM the agent runs on. Use `GET /models` to discover valid values.
descriptionNo
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
system_promptNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the safety profile is partly covered. The description adds value beyond them by naming the permission requirement, mandating explicit confirmation, warning that runs can spend credits or trigger downstream actions, and advising against automatic resubmission of unknown outcomes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, followed by concise preconditions and warnings. The closing sentence about runs spending credits is slightly tangential to a create operation, but it remains short and relevant as a caution.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 16-parameter, nested-object mutation with no output schema, the description covers permissions, confirmation, and side effects but omits how to reconcile the body-flag parameters versus the payload/payload_file alternatives and gives no sense of the return value. Adequate but with clear gaps where the schema is expected to carry the load.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 75% schema description coverage across 16 parameters, the schema itself is the primary source of field semantics, including the detailed is_active and skill_ids behaviors. The description contributes nothing about individual parameters except the confirm gate, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Create a new agent'), which is unambiguous on its own. However, it never distinguishes itself from sibling mutation tools such as update_agent, so the agent must rely on the name alone to know this is creation rather than modification.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It supplies a real precondition ('caller must have permission to create agents on the target team') and a confirmation gate ('Explicit confirmation is required for this exact account operation'), which implies usage conditions. It does not, however, state when to prefer this over alternatives or any when-not-to-use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_brain_sourceCreate sourceA
Destructive

Create a file-upload source. Only direct_file_uploads sources can be created through the API; connected sources such as Notion or Google Drive are set up in the Gumloop app because they need an account connection.

By default the source is active and indexes (and bills) each file as soon as it is uploaded. Set require_approval to true to create it as a draft instead: uploads then run a credit estimate, the source owner is notified, and nothing is indexed until Approve source is called. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoDisplay name of the source.
scopeNoWhich Brain the source belongs to. `team` requires `team_id`.personal
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
team_idNoThe team for `scope: team`. Not accepted with other scopes.
source_typeNoOnly `direct_file_uploads` is accepted.direct_file_uploads
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
require_approvalNoCreate as a draft that estimates credits before anything is indexed.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations mark this as destructive and non-idempotent; the description adds substantive context beyond that: the default `active` state indexes AND bills per file, while `require_approval: true` creates a `draft`, triggers a credit estimate, notifies the owner, and defers indexing until approval. It also warns about irreversible spend. It does not state permissions/account-auth prerequisites in depth, but the behavioral disclosure is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core scope restriction and the default-vs-draft behavior, in three tight paragraphs. It is slightly longer than needed for a create tool, and the final confirmation paragraph sits after the behavioral detail rather than at the front.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent creation tool with 9 parameters (0 required) and no output schema, the description covers what gets created, the default state, the draft/approval path, and the confirmation requirement. It omits nothing critical for correct invocation, though return/estimate details for the approval flow are only implied.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters, including the enum on `scope` and the `team_id` constraint. The description reinforces the `require_approval` semantics and the `direct_file_uploads` restriction, but adds little syntax/format detail beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Create a file-upload source') and immediately scopes the creation to `direct_file_uploads`, distinguishing it from connected sources like Notion or Google Drive. This differentiates it from the sibling tools (e.g. `list_brain_sources`, `upload_brain_files`).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly says only `direct_file_uploads` can be created through the API, and connected sources must be set up in the Gumloop app — an explicit 'when-not' condition. It also explains the `require_approval` branch. However, it does not point to a specific sibling alternative for those connected sources.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_chat_completionCreate chat completionB
Destructive

OpenAI-compatible chat completions endpoint, multiplexed across every model Gumloop supports (Anthropic, OpenAI, Google Gemini, OpenRouter routes). Set stream: true for Server-Sent Events, or omit it for a unary JSON response. Image-generation models (gpt-image-*, gemini-*-image-preview, dall-e-*) are dispatched automatically when modalities includes "image" and yield image attachments on choices[0].message.images.

Streaming host

Chat completions live on the streaming host. Send all requests — unary or streaming — to:

POST https://ws.gumloop.com/api/v1/chat/completions

api.gumloop.com does not serve this endpoint; the Python SDK routes there automatically.

Tool calls, images, and tool_choice

Send messages in the OpenAI shape and Gumloop translates them for the model's provider (Anthropic, OpenAI, and Google Gemini). Models served through OpenRouter and other OpenAI-compatible providers receive the messages as sent.

  • Tool-result turns: after the model replies with finish_reason: "tool_calls", append its assistant message (with tool_calls) and one {"role": "tool", "tool_call_id": ..., "content": ...} message per call, then send the conversation again. Every tool call needs a matching tool message, and every tool message must match a tool call in an earlier assistant message.

  • Images: user messages accept image_url content parts alongside text parts. The URL can be an http(s) URL or a base64 data URL (data:image/png;base64,...). Images must be JPEG, PNG, GIF, or WebP and at most 20 MB. Redirects are not followed when downloading an image.

  • tool_choice: "auto" (the default when tools are sent), "none", "required", or {"type": "function", "function": {"name": "..."}} to force one tool.

  • developer messages are treated like system messages.

{
  "model": "claude-sonnet-4-5",
  "tools": [{"type": "function", "function": {"name": "get_weather", "parameters": {"type": "object", "properties": {"city": {"type": "string"}}}}}],
  "messages": [
    {"role": "user", "content": [
      {"type": "text", "text": "What's the weather where this photo was taken?"},
      {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}
    ]},
    {"role": "assistant", "content": null, "tool_calls": [
      {"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": "{\"city\": \"Ottawa\"}"}}
    ]},
    {"role": "tool", "tool_call_id": "call_1", "content": "12°C and sunny"}
  ]
}

A request that can't be translated returns 400 invalid_request with param set to the field at fault (for example messages[3].tool_call_id). When the provider itself rejects the request (HTTP 400, 404, 413, or 422), the error message relays the provider's reason, prefixed with The provider rejected the request:.

Billing

Each completion charges the caller's credit balance based on token usage (with cache-token semantics per provider) plus a flat 30-credit fee for image-gen calls. Users who configure their own provider API key get a 50% discount. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel slug. Use the `id` from `GET /models` or one of Gumloop's preset routes.
toolsNoOpenAI-shape tool definitions (`{type: "function", function: {name, description, parameters}}`). Pass `tool_choice` to constrain selection.
streamNoOnly non-streaming JSON is supported. Streaming requests are refused before fetch.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
messagesNo
providerNoOpenRouter provider routing config. Caller fields like `sort` and `order` are honored; ZDR/data_collection policy is server-enforced.
modalitiesNoOutput modalities. Include `"image"` to route to an image-generation model.
temperatureNoSampling temperature.
tool_choiceNo
image_configNo
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
response_formatNo
max_completion_tokensNoCap on completion tokens. Replaces the deprecated `max_tokens` field.

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing billing semantics (per-token credit charge, flat 30-credit image-gen fee, 50% discount with own provider key), error shapes (400 invalid_request with param, provider-rejection prefixing), and a confirmation requirement for account-affecting operations. The main defect is the inaccurate streaming/SSE behavior, which is a schema contradiction rather than an annotation contradiction, so it does not trigger that flag.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a reference doc rather than a tool description: three headers, a code block, a full JSON example, and repeated image/streaming material. Streaming is discussed twice, once in the intro and again under its own heading, and the closing 'Explicit confirmation...' paragraph reads as grafted from another tool. Front-loaded purpose is present, but the length is not earned.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A 15-parameter, nested, mutation-capable tool with 0 required params and no output schema warrants substantial detail, and the tool-call loop and error handling are covered. But several parameters are left undocumented and the description asserts a streaming capability the schema forbids, so an agent cannot fully trust it as a complete spec.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 73%, so the schema carries most parameter meaning; the description adds value for tool_choice values, image_url content parts (accepted types, 20 MB limit, no redirects), and the tool/tool_call_id pairing rule. It says nothing about temperature, max_completion_tokens, provider, payload_file, account, or confirm, so it does not fully compensate for the remaining gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line names a specific verb+resource ('chat completions endpoint') and its distinguishing trait — multiplexed across Anthropic, OpenAI, Gemini, OpenRouter — which separates it from siblings like route_model or list_models. It stops short of a 5 because the purpose is muddied by a streaming claim (see below) that misrepresents what the tool actually does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives real usage context: when image-gen models dispatch (modalities includes 'image'), how to continue after finish_reason: tool_calls, and tool_choice options. But the primary 'when to use' guidance — 'Set stream: true for SSE, or omit it for a unary JSON response' — directly conflicts with the schema, which pins stream to const:false and refuses streaming, so the routing advice is actively misleading.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_organization_evaluationCreate evaluationA
Destructive

Creates an organization evaluation. A new evaluation has no targets, so it cannot start enabled: set targets with PUT /evaluations/{evaluation_id}/targets, then enable it with PATCH /evaluations/{evaluation_id}.

Rubric values are validated strictly: an unknown frequency, criterion priority, type, data point data_type, or session type, a criterion without name and prompt, or a duplicate tag name returns 400 invalid_request with the offending paths in error.details.fields. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare destructive/openWorld/non-idempotent; the description adds substantive behavior beyond them: the create-then-targets-then-enable lifecycle, strict rubric validation that returns 400 invalid_request with offending paths in error.details.fields, and the confirmation/credit-spend caution. This is exactly the extra context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then prerequisites, then validation and safety notes. Three short blocks, each carrying distinct information, though the validation sentence is dense and the endpoint paths overlap with what a schema-aware agent already knows.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent creation tool with no output schema, the description covers lifecycle, prerequisites, failure modes, and confirmation expectations, and implies an evaluation_id is produced. It stops short of describing what the response returns or the semantics of the config rubric versus update paths, but nothing critical to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and there are only two top-level params, so the baseline is 3; the description earns above that by explaining that 'explicit confirmation is required for this exact account operation' (the confirm flag) and that the operation is account-scoped. It adds meaning but not per-field syntax detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Creates an organization evaluation') and immediately situates it in the lifecycle, pointing to the PUT targets and PATCH enable endpoints that correspond to sibling tools (set_organization_evaluation_targets, update_organization_evaluation). An agent can separate creation from update/enable/delete without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real conditional guidance: a new evaluation has no targets, so it cannot be enabled until targets are set and then patched. It also states explicit confirmation is required and warns against auto-resubmitting unknown outcomes. It does not name sibling tools as alternatives for listing/updating evaluations, so it stops short of full when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_sessionCreate sessionA
Destructive

Create a new session for an agent. When input is provided, the message is enqueued and the agent begins processing — the response returns 202 with the session in processing or queued state. When input is omitted, an idle session stub is created and the response returns 201.

agent_id also accepts the reserved aliases gumball and analytics, which resolve to your personal Gumball and analytics agents (created on first use).

Streaming the response

api.gumloop.com only serves the non-streaming response above. To stream agent output as it's produced, send the same request body (with stream: true) to the streaming host instead:

POST https://ws.gumloop.com/api/v1/agents/{agent_id}/sessions

The response is text/event-stream (Server-Sent Events). With the Python SDK, client.sessions.stream(agent_id, input="...") routes to ws.gumloop.com automatically and yields parsed StreamEvent objects.

If you send stream: true to api.gumloop.com by mistake, the response is a 400 whose body contains the correct streaming host so you can retry against it. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
inputNoThe first user message for the session. Also accepted as `message` for backwards compatibility. When omitted, an idle session is created with no messages.
streamNoMust be `false` (or omitted) when calling `api.gumloop.com`. Set to `true` only when calling `ws.gumloop.com` (see the streaming section above).
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
agent_idYesID of the agent to start a session on. Also accepts the reserved aliases `gumball` and `analytics`.
metadataNoArbitrary key/value metadata attached to the session. Stored under `metadata.client`.
session_idNoCaller-supplied session ID. When omitted, the server generates one. If provided and the ID already exists, the request returns `409 session_already_exists`.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructive/openWorld/non-idempotent): it discloses status codes for each mode, session state transitions, the streaming host split, the 400 misrouted-stream error containing the correct host, the 409 session_already_exists case, and a credit-spend/confirmation warning with 'never resubmit unknown outcomes automatically'. This is exactly the behavioral context annotations alone cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and mode behavior, then uses a labeled streaming section. It is somewhat long, and the trailing confirmation paragraph reads as appended boilerplate, but every section serves the complex, multi-host semantics of this tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description covers the return surface (201/202 states, SSE stream, 400 and 409 errors) and the streaming host distinction. For a 10-parameter tool with nested objects, it supplies what an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 90%, so the baseline is 3. The description adds meaning beyond the schema by explaining that agent_id aliases gumball/analytics resolve to personal agents created on first use, and by clarifying that input enqueues a message and starts processing versus creating an idle stub.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a new session for an agent') and clearly distinguishes the two operating modes (input provided → 202 processing/queued; input omitted → 201 idle stub). It does not, however, name the sibling tools it competes with (e.g., send_message, queue_session_message), so sibling differentiation is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete conditions for the two creation modes and for the streaming host, plus the confirm requirement. It stops short of routing the agent among alternatives like send_message or queue_session_message when a session already exists, so it is strong context without explicit exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_skillCreate skillA
Destructive

Upload a skill package and create a new skill. The package must include a SKILL.md with name and description frontmatter; uploads may be a single .md file (stored as SKILL.md), or a .zip / .skill archive containing SKILL.md at its root. The initial version is created automatically. Maximum upload size is 10 MB. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
filesNo
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
team_idNoTeam that should own the skill. When omitted, the skill is owned by the authenticated user.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false and openWorldHint=true; the description goes further by disclosing credit spend, downstream side effects, and the mandatory confirmation gate. That is real value beyond the annotations, though it contains a size figure that conflicts with the schema (see below).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and packaging rules, then the safety/confirmation note. The second paragraph overlaps with the confirm parameter's own description, but overall it earns its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and 6 parameters (none required), the description adequately covers the packaging contract, confirmation requirement, and side-effect risk. The size contradiction and unaddressed body flags (payload/payload_file/account/team_id) are the remaining gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83% (baseline 3) and the description does add useful non-schema constraints: SKILL.md with name/description frontmatter, accepted .md/.zip/.skill forms, archive layout. However, it never explains the confirm, account, team_id, or payload_file parameters, and its stated 'Maximum upload size is 10 MB' contradicts the schema's 'at most 5 MiB' per file/total, which could cause failed uploads.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('Upload a skill package and create a new skill') plus the packaging constraints that define what this operation actually consumes. It is trivially distinguishable from siblings like update_skill, delete_skill, and upload_file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the confirmation prerequisite and the 'never resubmit unknown outcomes' rule, which is genuinely actionable guidance. It does not explicitly route against alternatives such as update_skill or a plain file upload, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_brain_fileDelete fileB
Destructive

Remove a file from the source and from search. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
file_idYes
source_idYesThe source id.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds genuinely useful context that deletion also removes the file from search, and reinforces the confirmation requirement. The trailing sentence about runs spending credits and downstream actions reads like generic boilerplate that is only loosely tied to a file-delete tool, diluting the value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in a tight first sentence. The second sentence is bloated and partially off-topic (credits, downstream runs), which is not clearly earned for a file-deletion operation and makes the guidance muddier rather than sharper.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with no output schema and 75% parameter coverage, the description covers the key effect (removal from source and search) and the confirmation gate, which is the minimum an agent needs. It leaves the file_id parameter and the return/error behavior unexplained, and the confusing run-centric sentence reduces rather than improves completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the schema already documents account, confirm, and source_id, so the baseline is 3. The description reinforces the confirm/account semantics ('exact account operation') but adds no syntax or format detail beyond what the schema's own parameter descriptions already provide, and file_id remains undocumented in both.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Remove) and resource (a file) plus the scope of effect: 'from the source and from search.' This lets an agent distinguish it from the sibling delete_brain_source, which removes a source rather than a file, without opening either schema. It does not explicitly name that sibling, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states a real usage precondition ('Explicit confirmation is required for this exact account operation') and warns against automatic resubmission, which implies the user must explicitly request the delete. However, it never states when to use this tool versus delete_brain_source or other brain-file tools, so the routing guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_brain_sourceDelete sourceA
Destructive

Delete a source, every file in it, and everything it contributed to search. This cannot be undone. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
source_idYesThe source id.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false and readOnlyHint=false, so the safety profile is covered structurally. The description still adds value by spelling out the irreversible cascade (files plus search contributions) and the explicit-confirmation gate, which annotations cannot express. It stops short of describing downstream consequences like search index rebuild timing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the destructive scope in the first sentence and keeps the whole thing to three sentences. The final sentence about credits and unknown outcomes is noticeably less relevant to a delete operation and dilutes the otherwise tight structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema needed and annotations carrying the safety hints, the description supplies the two things structure cannot: irreversible cascade semantics and the confirmation requirement. An agent has enough to call this correctly, though the credit/retry caveat is noise rather than needed context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so account, confirm, and source_id are already documented in the schema itself. The description's mention of required confirmation loosely reinforces the confirm flag but adds no format or constraint detail beyond what the schema states. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (delete) and resource (source) plus the exact blast radius: every file in it and everything it contributed to search. That cascade scope cleanly separates it from siblings like delete_brain_file and delete_skill without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It establishes the confirmation precondition, which is real usage guidance, but never names an alternative or states when to delete a whole source versus deleting individual files. The trailing 'runs can spend credits' caveat is generic boilerplate that doesn't route the agent between siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_organization_evaluationDelete evaluationA
Destructive

Deletes the evaluation. It stops running and disappears from lists; results it already produced stay attached to their sessions. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
evaluation_idYesID of the organization evaluation.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, but the description goes well beyond them: it explains exactly what is destroyed vs preserved (results remain attached to sessions), requires explicit confirmation, and warns about credit spend and downstream actions. That is substantive behavioral context an agent cannot get from the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and consequence before the confirmation and safety caveats. No filler or restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent tool with no output schema, the description covers consequences, confirmation requirements, and side effects adequately. The final sentence about runs spending credits is slightly tangential to a delete operation and mildly muddies the focus, keeping it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents account, confirm, and evaluation_id. The description reinforces that confirmation is mandatory and tied to 'this exact account operation' but adds no syntax or format detail beyond the schema, matching the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Deletes the evaluation') plus the concrete consequences — it stops running, disappears from lists, and prior results persist on sessions. This distinguishes it from sibling mutators like delete_skill, delete_brain_source, and update_organization_evaluation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear conditional guidance ('Explicit confirmation is required for this exact account operation') and a warning not to auto-resubmit unknown outcomes. It does not name a sibling alternative (e.g., update_organization_evaluation or a cancel-type tool) for cases where deletion is not intended, so it falls just short of the top band.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_queued_messageDelete queued messageA
Destructive

Remove a message from the session's queue before it is sent. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
session_idYesID of the session the queued message belongs to.
queued_message_idYesID of the queued message to delete.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds the confirmation requirement and the pre-send timing constraint, which is useful, but the sentence 'Runs can spend credits or trigger downstream actions' is boilerplate that does not apply to removing a queued message and reads as misleading noise.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core statement is front-loaded and efficient, but the third sentence about runs spending credits and not resubmitting unknown outcomes does not earn its place for a queue-deletion tool and dilutes the definition. Two of three sentences are on-topic; the third is generic copy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent tool with no output schema, the annotations cover the safety profile and the description covers the confirmation prerequisite, so the essentials are present. It still omits whether deletion is reversible or what the result looks like, and the off-topic credit/downstream sentence leaves the definition feeling partially templated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (account, confirm, session_id, queued_message_id) are already documented in the schema. The description only loosely gestures at confirm and account ('explicit confirmation', 'this exact account operation') and adds no syntax, format, or constraint detail beyond what is structured.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a specific verb and resource ('Remove a message from the session's queue') plus a scope qualifier ('before it is sent') that cleanly separates it from siblings like send_queued_message, update_queued_message, and list_queued_messages. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Explicit confirmation is required for this exact account operation' implies a usage condition (only delete when the user asked for exactly this action), which is reinforced by the schema's confirm field. However, no alternatives are named (e.g., update_queued_message to edit instead of delete) and the closing sentence about runs/credits adds no routing guidance for this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_skillDelete skillA
Destructive

Permanently delete a skill. This is a soft-delete — the skill will no longer appear in listings or be usable by agents. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
skill_idYesID of the skill to delete.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, so the bar is lower. The description still adds meaningful behavior: the operation is a soft-delete whose effect is removal from listings and loss of agent usability, plus a confirmation precondition. The 'permanently delete' / 'soft-delete' phrasing is internally muddled but not contradicted by the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The delete behavior and soft-delete consequence are front-loaded and useful. The final sentence about runs spending credits and downstream actions is off-topic for a skill deletion and reads as filler, diluting an otherwise tight definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter mutation with full schema coverage and annotations covering the safety profile, the description supplies the missing pieces: soft-delete semantics and the confirmation requirement. No output schema exists, so return values need not be explained. Only the irrelevant credits sentence detracts.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so account, confirm, and skill_id are already documented in the schema. The description reinforces that confirmation is mandatory and account-scoped, but adds no syntax or format detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Permanently delete a skill') and adds scope detail that the skill stops appearing in listings and becomes unusable by agents. It doesn't explicitly distinguish itself from siblings like update_skill or delete_brain_source, but the verb+resource pairing is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states that explicit confirmation is required for 'this exact account operation,' which is real usage guidance. However, the trailing sentence about runs spending credits and downstream actions reads like boilerplate copied from a run-execution tool and does not tell the agent when to prefer delete_skill over update_skill or when deletion is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

detach_agent_mcp_serverDetach an agent MCP serverA
Destructive

Detach an MCP server (connector) from an agent. This is idempotent — detaching a server that isn't attached returns detached: false rather than an error. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
agent_idYesID of the agent. Also accepts the reserved aliases `gumball` and `analytics`.
server_idYesID of the MCP server to detach.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and idempotentHint=false, but the description directly contradicts the idempotent flag by stating the operation IS idempotent and returns detached:false rather than an error. Beyond annotations, it discloses the return shape of the edge case and the confirmation/credit-spend caution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action, then adds idempotency and confirmation details. Two sentences plus a caution, all earning their place. Slightly longer than strictly necessary but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter mutation tool with no output schema, the description covers the action, idempotency edge case, confirmation requirement, and a safety caution about credits. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all four parameters. The description adds value by clarifying the confirm parameter's semantics ('explicit confirmation... exact account operation') and implying the idempotent behavior of server_id. Baseline 3 is exceeded because the description reinforces key parameter intent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (detach) and resource (MCP server/connector) with explicit scope ('from an agent'). Clear counterpart to sibling attach_agent_mcp_server and distinguishable from list_agent_mcp_servers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The confirmation requirement ('Set true only when the user asked for exactly this action') gives clear when-to-use guidance, and 'never resubmit unknown outcomes automatically' adds a caution. However, it doesn't explicitly name attach_agent_mcp_server as the inverse operation or spell out when detaching is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_artifact_fileDownload artifactA
Read-onlyIdempotent

Returns a signed download URL for an artifact, plus its filename, media type, and size. Follow download_url to fetch the file bytes. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
version_idNoSpecific version of the artifact to download. Defaults to the latest version when omitted.
artifact_idYesID of the artifact to download.
output_fileYesNew absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds real behavioral value beyond that: it discloses the return shape and, more importantly, that the tool does not stream bytes but hands back a signed URL that must be followed, which an agent could not infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences that lead with the return value and then the required follow-up action. The trailing 'Read operation.' is mildly redundant against readOnlyHint=true, but it costs almost nothing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining returns and does so (URL, filename, media type, size), plus the follow-the-URL step. Combined with fully documented parameters and annotations, an agent has enough to call this correctly; only sibling differentiation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (account, version_id, artifact_id, output_file) are already documented, including version defaulting and the private output_file requirement. The description adds no parameter-level detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns a signed download URL for an artifact') and enumerates what comes back (filename, media type, size), so the agent knows exactly what the call yields. It does not, however, differentiate itself from the nearby download_file / download_files / download_skill_file siblings, leaving the artifact-specific scope implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The instruction 'Follow `download_url` to fetch the file bytes' implies the intended two-step usage, which is genuinely useful. But there is no explicit guidance on when to pick this over download_file or download_files, and no mention of prerequisites or when-not-to-use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_fileDownload fileC
Read-onlyIdempotent

Download file Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idNoThe ID of the flow run associated with the file.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
user_idNoOptional. The user ID associated with the flow run.
file_nameNoThe name of the file to download.
project_idNoOptional. The project ID associated with the flow run.
output_fileYesNew absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
saved_item_idNoThe saved item ID associated with the file.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered structurally. The description's 'Read operation' merely echoes readOnlyHint and adds no new context such as auth/credential requirements or the meaning of the private output directory behavior. No contradiction, but essentially zero added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is short but not concise in a useful sense: it is under-specified rather than economical, spending its two lines on a restated name and a redundant 'Read operation' while omitting everything an agent would need.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with a nested payload object, a required output_file, private-credential 'account' semantics, and no output schema, the description is far too thin. It leaves the agent to reconstruct usage entirely from the schema and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema itself documents all 9 parameters including the required output_file and the nested payload object. The description contributes no parameter meaning, which is acceptable only because the schema does the heavy lifting; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is essentially a tautology: it restates the tool name and title ('Download file') with no verb+resource elaboration, no scope, and no differentiation from the many sibling download tools (download_files, download_artifact_file, download_skill_file). An agent cannot tell what resource space this operates in from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative guidance whatsoever. With siblings like download_files, download_artifact_file, and download_skill_file in the toolset, the absence of any routing signal is a serious omission.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_filesDownload multiple filesC
Read-onlyIdempotent

Download multiple files Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idNoThe ID of the flow run associated with the files.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
user_idNoThe user ID associated with the files. Required if project_id is not provided.
file_namesNoAn array of file names to download.
project_idNoThe project ID associated with the files. Required if user_id is not provided.
output_fileYesNew absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
saved_item_idNoOptional. The saved item ID associated with the files.

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint. The description's "Read operation" merely echoes readOnlyHint and adds nothing new — no note on credential scope for the private account, no rate/limit behavior, no mention that results are written out of model context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two fragments with no wasted words, but the problem here is under-specification rather than conciseness. Brevity at this level of complexity leaves the agent without the routing and prerequisite information it needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with a nested payload object, a required output_file with non-overwrite/private-directory semantics, and mutually exclusive payload_file vs payload vs body flags, the description covers none of the call-shaping rules. With no output schema the burden falls entirely on the description, and it does not carry it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all nine parameters (including the nested payload object and the required output_file) are already documented in the schema. The description adds no syntax, format, or mutual-exclusivity detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Download multiple files" states a verb and resource and hints at scope (plural vs the singular sibling download_file), but it is essentially the title restated with a two-word addendum. It gives no indication of what files are being downloaded, from where, or under what identity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance at all. The agent gets no cue for choosing this over download_file, download_artifact_file, download_skill_file, or the upload_* siblings, and no prerequisites (run_id/user_id/project_id either-or) are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_skill_fileDownload skillA
Read-onlyIdempotent

Generate a signed URL to download a skill's contents as a .skill archive (ZIP). When version_id is provided, returns that exact version; otherwise returns the current draft. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
skill_idYesID of the skill to download.
version_idNoSpecific version to download. When omitted, the current draft is returned.
output_fileYesNew absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld), so the bar is lower. The description adds genuine context the annotations do not: that the output is a signed URL rather than raw bytes, and that the archive format is ZIP. It does not mention expiry of the signed URL, but that is a minor omission.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with the core action front-loaded, the version branch second, and a one-line read-operation tag. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly explains what is returned (a signed URL) and the archive format. It does not mention the required output_file destination at all, which is the one gap for a tool that writes results to a private file.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter — including version_id's draft default and output_file's private-file semantics — is already documented. The description restates the version_id behavior but adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource — generate a signed URL for a skill's contents as a `.skill` (ZIP) archive — and names the sibling resources it is not (skills vs. generic files/artifacts). An agent can distinguish it from download_file and download_artifact_file without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the version_id vs. draft branch, which is useful selection context, but it never names an alternative tool or states when to prefer this over download_file/download_files. Usage is implied by the resource type rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_dataExport dataA
Destructive

This endpoint allows enterprise organization administrators to create and initiate a comprehensive data export for their organization or specific workspaces.

The export supports six data types:

  • Workflow data (data_type: "workflows"): Includes workflow runs, workbook details, user information, and other organizational data.

  • Agent data (data_type: "agents"): Includes agent configurations, metadata, tools, and creator information.

  • Agent interaction data (data_type: "agent_interactions"): Includes agent interaction data with timestamps, credit costs, trigger types, and message counts.

  • Credit log data (data_type: "credit_logs"): Includes credit transaction history with charges, balances, categories, and user attribution.

  • Interaction evaluation data (data_type: "interaction_evaluations"): Includes one row per completed evaluation of a chat, with its grade, call outcome, sentiment, and the model that graded it.

  • Gumstack data (data_type: "gumstack"): Includes Gumstack MCP tool call activity with timestamps, statuses, and latency.

The available export_fields depend on the selected data_type. See the field descriptions below for details.

Scoping requirement: For non-credit-log exports, at least one scoping parameter must be provided: workspace_ids, include_all_workspaces, include_personal_workspaces, or entity_ids. Requests that omit all scoping parameters will receive a 400 error.

Note: Credit log exports work differently from workflow and agent exports. When data_type is "credit_logs", the following parameters are not applicable and will be ignored: export_level, workspace_ids, include_all_workspaces, include_personal_workspaces, and entity_ids. Credit log exports are always scoped to the entire organization. Use category_filter to filter by credit log category. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
user_idNoThe ID of the user requesting the export.
end_dateNoEnd date for the export in ISO 8601 format (e.g., `2025-12-31T23:59:59Z`).
data_typeNo
entity_idsNo
start_dateNoStart date for the export in ISO 8601 format (e.g., `2025-01-01T00:00:00Z`).
export_levelNo
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
export_fieldsNo
workspace_idsNo
category_filterNo
include_all_workspacesNo
include_personal_workspacesNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, openWorldHint=true and idempotentHint=false, but the description adds real substance: it can spend credits or trigger downstream actions, explicit confirmation is required, and unknown outcomes must not be resubmitted automatically. It also discloses the 400 failure mode for missing scoping and the divergent credit_logs behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and then organized under headers and bullets for the six data types, scoping requirement, and credit_logs caveat. It is long, but the length is driven by genuinely distinct data-type semantics rather than repetition; a few of the type summaries overlap with the schema text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter, nested-object tool with no output schema, the description covers data-type selection, scoping rules, and the credit_logs exception well. The main gap is that it never explains the async return/polling model (get_export_status exists as a sibling), leaving the agent unsure what the call yields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (47%), so the description has to carry weight, and it does: it explains that export_fields depend on data_type, that the scoping params are not applicable to credit_logs, and points to category_filter as the credit_logs alternative. It does not, however, describe the confirm/account/payload wrapper params or date-range semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('create and initiate a comprehensive data export') and scopes it to organization admins and workspaces. It enumerates all six data types with their contents and even distinguishes itself from audit-log exports ('available in the Gumloop app only, not through this endpoint'), so an agent can tell it apart from sibling export/status tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete prerequisites and conditions: the scoping requirement (at least one of workspace_ids/include_all_workspaces/include_personal_workspaces/entity_ids, else 400), and the credit_logs exception where scoping params are ignored. It does not explicitly route the agent to get_export_status for polling, so it stops short of naming alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_brain_sourceRetrieve sourceB
Read-onlyIdempotent

Fetch one source the authenticated user can see. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
source_idYesThe source id.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, and the description's 'Read operation' merely restates that. It does add the visibility-scoping constraint (only sources the user can see), which is mild extra context, but nothing about errors, missing sources, or response shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with zero padding. The core verb and scope appear first, and nothing is redundant beyond the terse 'Read operation' tag.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter read tool with full schema coverage and no output schema, the definition is minimally sufficient. It never indicates what a retrieved source contains or what happens when the id is invalid or inaccessible, which the description could reasonably cover.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both source_id and account are documented in the schema itself. The description contributes nothing beyond that, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Fetch) and resource (one source) with singular scope, which implicitly distinguishes it from the sibling list_brain_sources. However, it does not explicitly name that sibling or clarify what a 'source' is in this domain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Adds a visibility qualifier ('the authenticated user can see') but gives no when-to-use guidance, no prerequisites, and no routing to alternatives like list_brain_sources or search_brain. The agent must infer selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_brain_source_estimateRetrieve estimateA
Read-onlyIdempotent

The latest credit estimate for a source created with require_approval. estimate is null until the first upload has produced a run; poll until estimate.status is paused_for_approval, then call Approve source. estimated_credits is rounded up to the nearest 5 and is an estimate, not a quote. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
source_idYesThe source id.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, it discloses genuinely useful behavior: `estimate` is `null` until the first upload produces a run, `estimated_credits` is rounded up to the nearest 5, and the value is an estimate not a quote. This is rich, non-obvious context an agent needs to interpret results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense and front-loaded: purpose first, then workflow, then edge cases, all in a few tight sentences. The trailing 'Read operation.' is redundant with the readOnlyHint annotation and slightly dilutes the otherwise lean structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description describes the returned shape (`estimate` nullability, `estimate.status`, `estimated_credits`) and the polling lifecycle, compensating for the missing output schema adequately for this read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (source_id, account) are already documented in the schema. The description adds no parameter-level semantics beyond what structured fields provide, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('latest credit estimate for a source') and scopes it to sources created with `require_approval`, distinguishing it from sibling tools like get_brain_source and approve_brain_source. An agent can identify its role without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit workflow guidance: poll until `estimate.status` is `paused_for_approval`, then call the approve-source tool (named as a sibling workflow step). It tells the agent both when to call and what to do next, which is unusually actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_evaluation_configGet evaluation configB
Read-onlyIdempotent

Retrieve the current evaluation configuration for an agent, including criteria, tags, data points, and sentiment settings. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
agent_idYesID of the agent. Also accepts the reserved aliases `gumball` and `analytics`.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so 'Read operation' is largely redundant. The description does add useful disclosure about what the configuration contains (criteria, tags, data points, sentiment settings), but says nothing about permissions, account scoping behavior, or failure modes beyond that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with purpose and contents. The trailing 'Read operation' sentence is the only near-wasteful element since annotations already convey read-only status, but overall the description is tight and well-ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of describing return content, and it does so by enumerating the configuration sections returned. It is adequate for a simple two-parameter read tool, though it omits any note on when the config might be empty or missing for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (account and agent_id) are already fully documented, including the gumball/analytics aliases. The description adds no additional meaning about parameter behavior, which is the expected baseline-3 outcome when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (retrieve) and resource (the current evaluation configuration for an agent), plus enumerates what it covers: criteria, tags, data points, sentiment settings. It is distinguishable from update_evaluation_config by the 'current' read framing, but it never names or contrasts with any sibling explicitly, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives such as get_evaluation_options, list_evaluations, or get_evaluation_metrics, all of which appear in the sibling list. 'Read operation' is a behavioral note, not usage guidance; the agent must infer applicability from the resource name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_evaluation_metricsGet evaluation metricsA
Read-onlyIdempotent

Returns aggregated grade and tag counts for an agent's evaluations over a time window. Useful for dashboards and reporting on agent quality trends. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoNumber of days to look back (1-365). Defaults to 30.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
agent_idYesID of the agent. Also accepts the reserved aliases `gumball` and `analytics`.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is fully covered and the trailing 'Read operation' is largely redundant. The description does add behavioral value by disclosing what is returned (aggregated grade and tag counts), which matters since there is no output schema. It says nothing about aggregation granularity or pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the return semantics before the use case. The only waste is the trailing 'Read operation', which duplicates the readOnlyHint annotation without adding information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only metrics tool with no output schema, the description adequately conveys what is returned (aggregated grade and tag counts) and the time-window scope. Gaps remain around bucketing granularity and response shape, but these are minor given the richness of the schema and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (days, account, agent_id) are already documented in the schema, including the day range and the gumball/analytics aliases. The description adds no parameter detail beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: returns 'aggregated grade and tag counts for an agent's evaluations over a time window'. The scope (one agent, time-bounded, aggregated counts) is clear. It does not explicitly distinguish itself from close siblings like get_organization_evaluation_metrics or list_evaluations, but 'agent's evaluations' implies the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Useful for dashboards and reporting on agent quality trends' gives an implied usage context, which is better than nothing. However, it names no alternatives and states no when-not conditions, so the agent must infer when to prefer this over the organization-level metrics sibling or list_evaluations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_evaluation_optionsGet evaluation optionsA
Read-onlyIdempotent

Allowed values for evaluation fields and filters — session types, criterion types and priorities, data point types, frequencies, grades, statuses, target types, skip reasons — plus size limits. Use these instead of hardcoding enums. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry the read-only/idempotent/openWorld safety profile, so the bar is lower. The description adds substantive value by disclosing the scope of what is returned (the enumerated value categories plus size limits), which the annotations do not convey. "Read operation" is redundant with readOnlyHint but harmless.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then the usage directive, in three tight sentences with no filler. The category list is long but it is the actual content an agent needs, so it earns its space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the returned value categories and size limits, and it notes the read-only nature. Nothing critical is missing for a zero-required-param lookup tool, though it could tie itself more explicitly to the evaluation siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'account' parameter is fully documented in the schema. The description adds no parameter-level guidance, so the baseline of 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource ('Allowed values for evaluation fields and filters') and enumerates the categories it covers (session types, criterion types, frequencies, grades, statuses, etc.), which clearly separates it from config/evaluation siblings. It never names a sibling directly, so differentiation is implied rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use these instead of hardcoding enums" gives a clear usage directive that tells the agent why and when to prefer this tool. There are no explicit when-not conditions or named alternatives (e.g., get_evaluation_config vs. this), but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_export_statusGet data export statusA
Read-onlyIdempotent

This endpoint retrieves the status of a data export job and optionally downloads the export file (as CSV) if the export has completed successfully.

Use the data_export_id returned by the Export data endpoint to check progress. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
user_idNoThe ID of the user requesting the export status.
downloadNoSet to `true` to download the export file directly when the export is completed. When `true` and the export state is `COMPLETED`, the response will be a CSV file download instead of JSON.
output_fileYesNew absolute result file in a private owner-only directory. Private download or signed credential result stays out of model output; no overwrite.
data_export_idYesThe unique identifier of the data export job to check (returned by the Export data endpoint).

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so 'Read operation.' is pure repetition. The one added behavioral fact — the response switches to a CSV download rather than JSON when the export is COMPLETED and download is requested — is also restated verbatim in the schema's `download` description, so the description contributes little beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, laid out efficiently with the core action first and the prerequisite second. The trailing 'Read operation.' is redundant against the annotations and is the only wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must carry return-value semantics. It covers the CSV-vs-JSON response mode but never enumerates the possible job states (e.g. PENDING, COMPLETED, FAILED) or what the JSON payload contains, which is precisely what a status-checking tool's caller needs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds provenance the schema lacks by tying `data_export_id` to the Export data endpoint and clarifying the optional-download toggle in prose. It does not explain `output_file` or `account`, but those are documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('retrieves the status of a data export job') plus the conditional CSV download, and explicitly frames itself as the progress-check counterpart to the Export data endpoint. An agent can distinguish it from the sibling export_data and download_file tools without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It tells the agent when to reach for this tool: after calling Export data, using the returned `data_export_id` to check progress. There is no explicit when-not guidance or mention of failure/cancelled states, but the sequencing requirement is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_input_schemaRetrieve input schemaC
Read-onlyIdempotent

Retrieve input schema Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
user_idNoUser ID that created the flow. Required if project_id is not provided.
project_idNoProject ID that the flow is under. Required if user_id is not provided.
saved_item_idYesThe ID of the saved item for which to retrieve input schemas.

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, fully covering the safety profile. The description's only behavioral claim ("Read operation") merely echoes readOnlyHint, adding no new context such as auth/account coupling or scoping behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loaded, but the second line "Read operation" is pure redundancy with the annotations and earns no place. Brevity here reflects under-specification rather than efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool (with required saved_item_id and an account/user/project identity choice) and no output schema, the description says nothing about what an 'input schema' contains or how the identity parameters interact. It is too thin to guide correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (account, user_id, project_id, saved_item_id) are documented in the schema itself, including the user_id/project_id mutual requirement. The description adds nothing, but the baseline is 3 when the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Retrieve input schema" only restates the tool name and title without adding scope, format, or distinctions from siblings like list_mcp_server_tools or get_evaluation_config. An agent learns nothing beyond what the name already says.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Read operation" is a safety statement, not usage guidance. There is no indication of when to call this versus alternatives, no prerequisites, and no conditions for use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_mcp_server_promptGet MCP server promptB
Read-onlyIdempotent

Render one prompt template with arguments and return its messages. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoPrompt name from List MCP server prompts.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
team_idNoScope the lookup to a single team.
argumentsNoArgument values for the template.
server_idYes
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so 'Read operation' is largely redundant. The one piece of added value is 'return its messages', which discloses the return shape in the absence of an output schema. No auth or rate-limit context is provided.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the action front-loaded and no padding. The second sentence ('Read operation.') is redundant against the annotations, which is the only real waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with nested objects and no output schema, the description does too little: it never clarifies the top-level name vs payload.name duplication, the payload/payload_file/body-flag mutual exclusion, or the account credential-selection semantics. It only hints at the return value.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 86%, so the schema already documents server_id, name, account, team_id, payload, arguments and payload_file. The description adds no parameter-level detail at all – notably it does not explain the payload vs payload_file vs body-flag exclusivity, which the schema only hints at.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Render one prompt template with arguments and return its messages.' This is distinguishable from the sibling list_mcp_server_prompts, though the description never names that sibling. Clear but without explicit sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by 'Render one prompt template with arguments' – the agent can infer this is the execution counterpart to listing prompts. There is no explicit when-to-use, no prerequisites, and no pointer to watch server_id or account selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_organization_audit_logsRetrieve audit logsB
Read-onlyIdempotent

This endpoint retrieves audit logs for all users in an organization for a specified time period. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNoPage number for pagination.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
user_idNoYour user id -- you must be an organization admin to retrieve organization logs.
end_timeYesEnd timestamp for log filtering (ISO format).
user_idsNoComma-separated list of user IDs whose events should be returned.
page_sizeNoNumber of records per page.
entity_idsNoComma-separated list of entity IDs (agents, workbooks, files) to filter by. The singular `entity_id` param accepts a single value.
start_timeYesStart timestamp for log filtering (ISO format).
event_typesNoComma-separated list of event types to filter by (e.g. `user_sign_in,credential_retrieval`). The singular `event_type` param accepts a single value.
ip_addressesNoComma-separated list of source IP addresses to filter by. The singular `ip_address` param accepts a single value.
workspace_idsNoComma-separated list of workspace (team) IDs to filter by. The singular `workspace_id` param accepts a single value.
organization_idYesThe ID of the organization to retrieve audit logs for.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered. The description's 'Read operation' merely restates the annotation and the org-wide/time-bounded scope is its only added value, which is modest against an already-rich annotation set.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the core purpose front-loaded and zero padding. The trailing 'Read operation' is somewhat redundant given readOnlyHint, slightly diminishing an otherwise tight statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with full schema coverage and annotations covering safety and idempotency, the essentials are present, but the definition is thin: it omits the admin precondition, pagination expectations, and filtering capabilities that live only in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 12 parameters are documented (pagination, filters, admin requirement) in the schema itself. The description adds no parameter-level meaning, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (retrieves) and resource (audit logs) with scope (all users in an organization, specified time period). The name and description together make the operation unmistakable, though it doesn't distinguish itself from any sibling — which is largely unnecessary given the unrelated sibling set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative guidance is provided. The phrase 'for all users in an organization' hints that this is the org-wide variant, but it doesn't tell the agent when this is preferable to a user-scoped query or that admin privileges are required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_organization_evaluationRetrieve evaluationB
Read-onlyIdempotent

Returns one evaluation with its rubric, targets, current coverage, and result rollup. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
evaluation_idYesID of the organization evaluation.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. "Read operation" merely echoes those annotations and earns no credit. The content inventory (rubric, targets, coverage, rollup) is genuinely useful, but it describes output shape rather than behavior—no auth requirements, scoping, or rate/limit context is added.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, with the substantive content (what is returned) front-loaded. "Read operation." is redundant with the annotations and could be dropped, but overall the definition is tight with no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the right thing by naming the returned components, which is the strongest part of this definition. What is missing is disambiguation from the near-identically named retrieve_evaluation and the other evaluation siblings, a real risk in a toolset this crowded.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both account and evaluation_id are fully documented in the schema itself. The description adds nothing about parameter format, defaults, or how account selection affects visibility. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Returns one evaluation") and enumerates what comes back: rubric, targets, current coverage, result rollup. That is far more than a restatement of the name. However, it never distinguishes itself from the sibling "retrieve_evaluation" (which likely carries the same title) or from "get_evaluation_config"/"get_organization_evaluation_result", so an agent cannot disambiguate from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance offered is "Read operation," which is not a when-to-use statement. There is no mention of when to pick this over retrieve_evaluation, list_organization_evaluations, or get_evaluation_config, and no prerequisites stated. Usage must be inferred entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_organization_evaluation_metricsGet evaluation metricsB
Read-onlyIdempotent

Grade counts for one evaluation over a trailing window (default 30 days, 1–365). Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoWindow length in days.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
evaluation_idYesID of the organization evaluation.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds the window semantics and that the result is aggregate grade counts, but 'Read operation' merely repeats the annotations and no auth, scoping, or aggregation-freshness context is provided.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the return value and window. The trailing 'Read operation.' is largely dead weight given the readOnlyHint annotation, but the overall size is well controlled.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should carry the burden of describing the return payload; 'grade counts' hints at it but does not say what grades or runs are included. Annotations cover the safety profile, so the remaining gaps are modest but real for a metrics-read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all three parameters are documented in the schema. The description's restatement of the default (30 days) and bounds (1–365) duplicates the schema, and it adds nothing about the 'account' parameter or the evaluation_id requirement. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (grade counts for one evaluation) and the temporal scope (trailing window, default 30 days, 1–365). It is clear what the tool returns, but it never distinguishes itself from the near-identical sibling get_evaluation_metrics, leaving the agent to infer the organization-level distinction from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no alternative named. Given siblings such as get_evaluation_metrics and list_organization_evaluation_results, the description should tell the agent which one to pick, but it does not.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_organization_evaluation_resultRetrieve evaluation resultA
Read-onlyIdempotent

One result, including per-criterion outcomes, extracted data points, and applied tags. Poll this after POST /evaluations/{evaluation_id}/run until status is completed or failed. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
result_idYesResult ID from a run response or a results list.
evaluation_idYesID of the organization evaluation.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive, and the description adds genuine behavioral context beyond them: the polling-until-terminal-status pattern. The closing 'Read operation.' merely restates readOnlyHint, so it earns no extra credit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loading what the result contains before the polling instruction. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does partially describe the return payload (per-criterion outcomes, data points, tags), and it explains the polling lifecycle. It could say more about the status field values or error shape, but it is largely sufficient for a simple read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so evaluation_id, result_id, and account are already fully documented in the schema. The description adds no syntax, format, or sourcing detail beyond that, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (retrieve) and resource (a single organization evaluation result), and enumerates its contents (per-criterion outcomes, extracted data points, applied tags). The singular 'One result' clearly distinguishes it from the sibling list_organization_evaluation_results.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit operational condition: poll this after POST /evaluations/{evaluation_id}/run until status is completed or failed. That is strong when-to-use guidance, though it does not explicitly name the sibling list tool as the alternative for enumerating results.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_role_credit_limitGet custom role credit limitA
Read-onlyIdempotent

This endpoint returns the monthly credit limit of one custom role. A monthly_credit_limit of null means the role sets no limit of its own. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
role_idYesThe ID of the custom role (the same ID used as `group_id` by the Manage custom role users endpoint).
user_idNoYour user id -- you must be an organization admin to manage custom role credit limits.
organization_idYesThe ID of the organization the custom role belongs to.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safety profile is covered. The description adds one genuinely useful behavioral detail — that a `monthly_credit_limit` of `null` means no role-level limit — but nothing about auth beyond what the schema already states.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core purpose. The trailing 'Read operation.' is redundant given readOnlyHint=true, a minor waste, but overall the structure is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the work of explaining the key return semantic (null = no limit). Combined with fully documented parameters and a read-only profile, an agent has enough to call it correctly; only the sibling routing is under-explained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so role_id, organization_id, user_id, and account are all documented in the schema itself. The description adds no syntax or format detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'returns the monthly credit limit of one custom role.' The singular 'one' implicitly distinguishes it from the sibling list_role_credit_limits, but it never names that sibling or set_role_credit_limit, so an agent must infer the routing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the singular scoping ('one custom role'). There is no explicit when-to-use, no exclusion, and no reference to the natural alternatives (list_role_credit_limits for enumeration, set_role_credit_limit for mutation), leaving the agent to infer the boundaries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_run_detailsRetrieve run detailsA
Read-onlyIdempotent

This endpoint can be used to poll for completion and retrieve final flow outputs. Output steps must be used to retrieve outputs. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesID of the flow run to retrieve
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
user_idNoThe id for the user initiating the flow. Required if project_id is not provided.
project_idNoThe id of the project within which the flow is executed. Required if user_id is not provided.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds non-obvious behavior beyond the annotations: it is a polling endpoint and outputs are only retrievable via output steps. That is genuinely useful context that the structured fields do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short and front-loaded with the primary purpose. 'Output steps must be used to retrieve outputs' is slightly redundant in wording but carries real information, so nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining returns, and it only vaguely says 'retrieve final flow outputs' without describing run status values or the shape of the payload an agent polling for completion would need to inspect. The output-steps note helps but leaves the return contract underspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents run_id, account, user_id, and project_id, including the conditional requirements. The description adds no parameter-level detail, which is acceptable given the coverage; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: poll a run for completion and retrieve its final flow outputs. It's clearly a read retrieval tool. However, it does not distinguish itself from the sibling get_run_history, which an agent could easily confuse with it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'can be used to poll for completion' implies a usage context (awaiting run completion), which is helpful. But there is no explicit when-to-use vs when-not guidance and no mention of the alternative get_run_history, leaving the agent to infer the choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_run_historyRetrieve automation run historyB
Read-onlyIdempotent

This endpoint retrieves the run history for automations, either by workbook or saved item. Returns the 10 most recent runs. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
user_idNoThe user ID. Required if project_id is not provided.
project_idNoThe project ID. Required if user_id is not provided.
workbook_idNoThe ID of the workbook to retrieve run history for. Required if saved_item_id is not provided.
saved_item_idNoThe ID of the saved item to retrieve run history for. Required if workbook_id is not provided.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the 'Read operation.' sentence is redundant with structured data. However, the description does add a genuinely useful behavioral detail not in the annotations: it returns only the 10 most recent runs, which is a meaningful limit an agent must know. No pagination or auth details are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with purpose, followed by the key return limit. The final 'Read operation.' sentence is redundant given the annotations and slightly dilutes efficiency, but the description is otherwise tight with no waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with full annotation coverage and a 100%-documented schema, the description is adequate. It covers purpose, the two lookup modes, and a return limit. It doesn't address pagination, whether the 10-run cap is configurable, or error cases, but these are minor given the structured data already present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents all 5 parameters including the conditional requirements (user_id or project_id; workbook_id or saved_item_id). The description adds no parameter semantics beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb+resource: 'retrieves the run history for automations, either by workbook or saved item.' This distinguishes it from sibling get_run_details (which presumably returns a single run) and from list_flows/list_workbooks. It doesn't explicitly name those siblings, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives like get_run_details or list_flows. The description mentions two lookup modes (workbook or saved item) but doesn't explain when to prefer one or what prerequisites apply. No when-not-to-use context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_browser_profile_cookiesImport cookies into a browser profileA
Destructive

Add sign-in cookies to a browser profile. This is what gumloop browser import-logins calls. Send the cookies in Chrome extension (chrome.cookies.Cookie) or Chrome DevTools Protocol Cookie shape. With url, only that site's cookies are kept and the import replaces that site; without it, every site in the payload is imported. Cookies are encrypted with the profile's key before storage and are never returned by any endpoint.

Use default as the profile_id to import into the owner's default profile, creating it if needed. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoImport only the cookies for this site and replace what the profile had for it. Omit to import every site in `cookies`.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
cookiesNoCookies in `chrome.cookies.Cookie` or CDP `Cookie` shape.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
team_idNoImport into a team-owned profile instead of a personal one.
profile_idYesA profile id, or `default`.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (destructiveHint=true, openWorldHint=true), the description discloses that cookies are encrypted with the profile's key before storage, are never returned by any endpoint, that confirmation is required for this exact account operation, and that runs can spend credits or trigger downstream actions. This is exactly the extra behavioral context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and the key url/no-url behavior are front-loaded, and each paragraph covers a distinct concern (import semantics, then confirmation/credit risk). It is slightly dense and the url explanation duplicates the schema text, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool (8 params, nested payload object, no output schema) it covers the critical ground: mutating/destructive nature, encryption and non-return of secrets, confirmation requirement, and credit/spend risk. The various body-flag vs payload vs payload_file alternatives are left to the schema, but that is reasonable given the schema already documents them fully.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description restates the `url` semantics and the `default` profile_id hint that already appear in the schema, adding only marginal meaning (e.g., default profile is created if needed). It does not add syntax or format detail beyond the structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Add sign-in cookies to a browser profile') and even identifies the underlying CLI command it maps to, leaving no ambiguity about what the tool does. No sibling tool operates on browser-profile cookies, so the agent can immediately tell this apart from list_browser_profiles or the many unrelated siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear conditional behavior: with `url` only that site's cookies are kept and replaced, without it every site in the payload is imported, plus the `default` profile_id convention. It also flags that explicit confirmation is required and warns against auto-resubmitting unknown outcomes. It stops short of naming an alternative or a 'when not to use' boundary, but the operative conditions are spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kill_flowKill flow runA
Destructive

This endpoint is used to kill a flow run and all its subflow runs. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idNoThe ID of the pipeline run to kill.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
user_idNoThe user ID. Required if project_id is not provided.
project_idNoThe project ID. Required if user_id is not provided.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false and openWorldHint=true, so the safety profile is covered. The description adds genuinely new context beyond that: subflow runs are also terminated, credits/downstream actions are at stake, and outcomes must not be auto-resubmitted. It stops short of saying whether the kill is reversible or what happens to an already-finished run.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the purpose and then the confirmation prerequisite and retry caution. Efficient, though the second sentence is slightly awkwardly worded ('this exact account operation').

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter destructive tool with no output schema, the description covers the destructive scope, the confirmation gate and the non-idempotent retry hazard, which are the key things an agent needs. Remaining gaps (post-kill state, behavior on an already-terminated run) are minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (run_id, account, confirm, payload, payload_file, user_id, project_id) is already documented in the schema. The description only obliquely gestures at the account/confirm semantics via 'this exact account operation' and adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (kill) and resource (a flow run) and adds meaningful scope — 'and all its subflow runs' — which distinguishes it from benign siblings like get_run_details or cancel_session. An agent can identify this as the destructive termination tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states a prerequisite ('Explicit confirmation is required for this exact account operation') and a retry caution ('never resubmit unknown outcomes automatically'), which tells the agent how to invoke it safely. It does not, however, name an alternative tool or an explicit when-not-to-use condition relative to siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_accountsList configured accountsA
Read-onlyIdempotent

List private account labels, default selection and configured token method. No credentials, token paths or account content; no network request.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=false, and destructiveHint=false. The description adds meaningful behavioral context beyond those annotations: it confirms no credentials, token paths, or account content are returned, and that no network request is made.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences front-load the core scope and then immediately clarify the negative behavior. Every phrase contributes useful information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only listing tool with no output schema, the description adequately covers the returned concepts and explicitly rules out sensitive data and network activity. It stops short of describing result ordering, formatting, or pagination, but those may not apply or may be self-evident.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter semantics to clarify. Per the rubric, a zero-parameter tool has a baseline of 4, and the description does not need to compensate for undocumented inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource (list accounts) and enumerates the exact scope: private account labels, default selection, and configured token method. Its exclusions also implicitly distinguish it from broader siblings like get_account or get_current_token, which would expose account content or credentials.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by stating exactly what metadata the tool surfaces, but it never explicitly says when to call this instead of alternatives such as get_account or get_current_token. The 'No credentials...' clause scopes the tool, yet no named alternative or when-not condition is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_agent_mcp_serversList agent MCP serversA
Read-onlyIdempotent

List the MCP servers (connectors) attached to an agent. Sensitive fields such as secret_id and mcp_server_url are scrubbed from the response. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
agent_idYesID of the agent. Also accepts the reserved aliases `gumball` and `analytics`.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered structurally. The description adds genuinely new behavior beyond that: sensitive fields like `secret_id` and `mcp_server_url` are scrubbed from the response, which tells the agent what data it will and will not receive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with the main action front-loaded. The trailing 'Read operation.' marginally restates the readOnlyHint annotation rather than adding information, but the definition is otherwise tight and well structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, two-parameter listing tool with rich annotations and no output schema, the description covers the essential return-value caveat (scrubbed sensitive fields). Nothing critical is missing, though it doesn't note whether the list is paginated or empty-case behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both `account` and `agent_id` (including the reserved aliases) are already documented in the schema. The description only implies the agent scoping and adds no format or alias detail beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and a precisely scoped resource ('the MCP servers (connectors) attached to an agent'), which cleanly separates it from list_mcp_servers, attach_agent_mcp_server, and detach_agent_mcp_server. It stops short of naming those siblings explicitly, so it is clear but not maximally differentiating.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'attached to an agent' scope implies when to use this versus the general server-list tools, but there is no explicit when-to-use statement, no prerequisite guidance, and no named alternative. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_agentsList agentsA
Read-onlyIdempotent

List agents the caller has access to. Filter by team, search by name, or narrow to agents that use a specific tool or trigger. Results can be sorted with sort_order and paginated by sending page_size and/or cursor. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
flowNoFilter to agents that use the specified saved flow as a tool.
toolNoFilter to agents that use the specified MCP server as a tool.
cursorNoOpaque cursor from a previous response's `next_cursor`. Pass it to fetch the next page.
searchNoCase-insensitive substring match against the agent name.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
creatorNoFilter to agents created by this user ID.
team_idNoScope the listing to a single team. When omitted, returns agents owned by the authenticated user.
page_sizeNoNumber of agents per page. Sending `page_size` or `cursor` opts into cursor pagination; requests that send neither return the full list.
sort_orderNoSort order for the listing. Defaults to newest first.newest
has_triggersNoWhen `true`, only returns agents that have at least one active trigger configured.
include_last_usedNoWhen `true`, populates `last_used_at` on each agent with the timestamp of its most recent session.
include_last_updatedNoWhen `true`, populates `last_updated_at` on each agent with the timestamp of its most recent configuration change.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered; the closing 'Read operation' merely repeats that. The description does add the pagination behavior (sending page_size or cursor opts into paging, sending neither returns the full list), which is useful operational context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core purpose, then filtering, then sorting/pagination. Efficient and free of padding, though the trailing 'Read operation' is redundant against the annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-param, zero-required list tool with no output schema, the description covers purpose, filtering, sorting and pagination adequately, and the schema carries full parameter detail. The main omission is any hint about the shape of the returned agent list, but that is a minor gap given the otherwise complete coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 12 parameters are already documented with equivalent or greater detail in the schema. The description's summary of filters, sort_order, page_size and cursor adds no syntax, format, or constraint information beyond what the schema provides. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List agents') plus the access scope ('the caller has access to'), and enumerates the filterable dimensions. It is clearly distinguishable from singular siblings like retrieve_agent, though it does not name the alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lists what can be filtered/sorted/paginated, which implies usage, but gives no explicit when-to-use guidance or exclusions against siblings such as retrieve_agent, list_agent_versions, or list_agent_mcp_servers. The agent must infer the routing itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_agent_versionsList agent versionsA
Read-onlyIdempotent

List the immutable versions of an agent, newest first. Each entry is a point-in-time snapshot of the agent's configuration.

Requires configuration access on the agent — callers limited to using the agent (no configuration access) get a 403. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoOpaque pagination cursor returned by a prior call as `next_cursor`.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
agent_idYesID of the agent whose versions to list. Also accepts the reserved aliases `gumball` and `analytics`.
page_sizeNoNumber of versions to return per page. Clamped to 1–100.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly/idempotent/non-destructive, so the description's 'Read operation' is redundant, but it adds real value by disclosing the 403 authorization gate and the newest-first ordering. It does not describe return shape or error behavior beyond the auth failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads what the tool returns and how it is ordered, followed by the access precondition. No filler sentences; every clause carries information an agent needs before calling.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by explaining that entries are configuration snapshots and that pagination is cursor-based (via the schema). It omits the concrete fields inside each version entry, which is a modest remaining gap for a list tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and every parameter (cursor, account, agent_id, page_size) is documented in the schema itself. The description adds nothing about parameter semantics, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List the immutable versions of an agent'), plus ordering ('newest first') and what each entry represents ('point-in-time snapshot of the agent's configuration'). It is distinguishable from the singular retrieve_agent_version by implication (list-all vs fetch-one), but does not name any sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete precondition: configuration access is required and callers limited to using the agent receive a 403. That is a clear when-you-can/when-you-cannot signal. It stops short of pointing to alternatives such as retrieve_agent_version or list_agents for a single version or the agent inventory.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_artifactsList artifactsA
Read-onlyIdempotent

List artifacts (files) produced by an agent. Optionally scope to a specific session, search by filename, sort, and paginate. Deleted files are excluded from the results. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoOpaque pagination cursor returned by a prior call as `next_cursor`.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
agent_idYesID of the agent whose artifacts to list. Also accepts the reserved aliases `gumball` and `analytics`.
page_sizeNoNumber of artifacts to return per page. Clamped to 1–100.
session_idNoFilter to artifacts produced within a specific session.
sort_orderNoSort order for results. Defaults to `newest`.newest
search_queryNoCase-insensitive substring match against the artifact filename.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the description's trailing 'Read operation.' adds nothing. The genuinely useful addition is that deleted files are excluded from results, a filtering behavior not derivable from annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Efficient two-sentence body with the core action front-loaded and modifiers listed compactly. The standalone 'Read operation.' line is redundant given the readOnlyHint annotation, a small amount of waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a paginated read-only list tool with a fully described schema and rich annotations, the description covers the necessary ground; the deleted-file exclusion and pagination mention round it out. No output schema exists, but return shape for a list tool is largely implied by the cursor/page_size parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all seven parameters (including the alias support for agent_id and the opaque cursor semantics) are already documented. The description merely restates the same capabilities at a higher level and adds no syntax, format, or default details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List artifacts (files) produced by an agent'), and clarifies the ambiguous term 'artifacts' as files. It does not, however, differentiate itself from near siblings like list_brain_files or download_artifact_file, leaving the agent to infer scope from schema and annotations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description enumerates capabilities (scope by session, search, sort, paginate) which implies usage, but never states when to use this tool versus alternatives or any preconditions. There is no explicit when/when-not guidance, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_brain_filesList filesA
Read-onlyIdempotent

List the files in a file-upload source with their indexing status. Each file carries the sha256 of its bytes, so a client can compare a local folder against the source and upload only what changed. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNo
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
page_sizeNo
source_idYesThe source id.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the 'Read operation' sentence is redundant. What does add value is the disclosure of what each returned file carries (sha256 of bytes, indexing status), which helps a caller reason about the response without an output schema. Pagination behavior is still undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences plus a tag; purpose is front-loaded in the first clause. The trailing 'Read operation' sentence is wasted tokens against annotations that already say the same thing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully describes the returned items (files, indexing status, sha256) and annotations cover safety. The main gap is pagination semantics for cursor/page_size, which an agent needs to iterate a source fully.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description says nothing about any of the four parameters, and schema coverage is only 50% (cursor and page_size have empty descriptions). With half the parameters undocumented in structured data and no compensating explanation, the description leaves parameter meaning to the name alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and a precise resource ('the files in a file-upload source'), plus what the listing returns (indexing status, sha256). This is clearly distinguishable from siblings like list_brain_sources (sources, not files) and download_files (fetching, not listing).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description supplies a concrete use context — comparing a local folder against the source and uploading only changed files. That tells the agent when this tool is valuable, though it names no alternative (e.g., list_brain_sources) and gives no explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_brain_sourcesList sourcesA
Read-onlyIdempotent

List the Company Brain sources the authenticated user can see: personal sources, plus team and organization sources shared with them. Every source type is listed, including ones connected in the app such as Notion or Google Drive. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNoOnly sources in this scope.
cursorNoOpaque cursor from a previous response's `next_cursor`.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
team_idNoOnly team sources belonging to this team.
page_sizeNoMaximum number of sources to return.
source_typeNoOnly sources of this type, for example `direct_file_uploads` or `notion`.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint/idempotentHint/destructiveHint, and 'Read operation.' merely restates that. The one genuinely additive detail is that every source type is returned, including in-app connections like Notion or Google Drive, but pagination behavior and auth sensitivities are left to the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with the primary purpose front-loaded and no filler. The trailing 'Read operation.' is slightly redundant given readOnlyHint, but the overall structure is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema and fully documented optional params, the description covers what is enumerated and who can see it. It omits any note about the returned source shape or pagination semantics, but cursor/page_size documentation in the schema covers the mechanics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The prose lightly echoes the scope enum ('personal sources, plus team and organization sources') and the source_type filter, but adds no syntax, format, or interaction detail beyond what the schema already documents for the 6 optional parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List the Company Brain sources') and clarifies the visibility scope (personal plus team/organization shared). An agent can tell this is the enumeration tool rather than a search or single-fetch tool, though it never names the siblings it is distinct from (list_brain_files, get_brain_source, search_brain).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the enumeration framing plus the scope clause describing what is visible to the authenticated user. There is no explicit when-to-use, when-not-to-use, or named alternative, so an agent must infer that search_brain and get_brain_source are the filtered/single-item counterparts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_browser_profilesList browser profilesA
Read-onlyIdempotent

List the browser profiles you own, or a team's profiles with team_id. A browser profile holds the sign-ins an agent's Browser ability uses. Cookie values are never returned; each profile lists its sites and cookie counts. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
team_idNoList a team's profiles instead of your personal ones.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, open-world, so the safety profile is covered. The description still adds genuinely useful behavioral detail: cookie values are never returned and each profile reports its sites and cookie counts, which tells the agent what it can and cannot learn from this call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and scope, then the domain concept, then the return-value caveat. 'Read operation' is mildly redundant with readOnlyHint but costs almost nothing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully summarizes the return shape (profiles with sites and cookie counts, no cookie values), which is exactly what an agent needs. The only gap is that the `account` parameter's role in identity selection goes unmentioned in the description, though the schema covers it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the schema; per the rubric that sets a baseline of 3. The description reinforces `team_id`'s effect ('instead of your personal ones') but adds no syntax, format, or defaulting detail beyond the schema, and never mentions the `account` parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('List the browser profiles') with scope disambiguation built in: personal profiles by default, or a team's via `team_id`. The second sentence explains what a browser profile actually is, which helps an agent distinguish this from cookie-related siblings like import_browser_profile_cookies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context for when to use it: to enumerate profiles you own, or to enumerate a team's profiles when passing `team_id`. No explicit exclusions or named alternative tools (e.g., it doesn't say to prefer this over list_teams for team lookup), so it stops short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_evaluationsList evaluationsA
Read-onlyIdempotent

Returns a cursor-paginated list of evaluation results for a specific agent, newest first. Only completed and failed evaluations are returned unless status selects another state.

Each evaluation includes the grade, criteria pass/fail results, extracted data points, applied tags, and sentiment analysis. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
gradeNoFilter evaluations by grade.
cursorNoPagination cursor from a previous response's `next_cursor` field.
statusNoReturn evaluations in one lifecycle state instead of the default completed and failed set.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
agent_idYesID of the agent whose evaluations to list. Also accepts the reserved aliases `gumball` and `analytics`.
page_sizeNoNumber of evaluations to return per page (1-100).
session_idNoOnly evaluations of this session.
created_afterNoOnly evaluations created at or after this ISO 8601 timestamp. Timestamps without an offset are read as UTC.
created_beforeNoOnly evaluations created before this ISO 8601 timestamp. Timestamps without an offset are read as UTC.
organization_evaluation_idNoReturn the results one organization evaluation produced for this agent instead of the agent's own evaluation results.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/destructiveHint, so the safety profile is covered. The description adds value beyond that: pagination ordering (newest first), the default status filter, and the full shape of a returned evaluation (grade, criteria, data points, tags, sentiment). The trailing 'Read operation' merely repeats readOnlyHint, which is minor redundancy rather than a contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with the primary purpose and pagination/ordering facts, then the default filter, then returned fields. Efficient overall, but the final 'Read operation.' sentence restates the readOnlyHint annotation and does not earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates what each evaluation contains and documents ordering and the default status filter. It omits only minor things an agent might want (e.g., the account param's credential scoping), but is essentially complete for a filtered list tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameter baseline is 3. The description adds meaning beyond the schema by clarifying the default status set that `status` overrides and by implying cursor-based pagination ties to `next_cursor`. It does not enrich the other parameters (account, grade, date bounds), keeping it short of 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Returns a cursor-paginated list), resource (evaluation results), and scope (for a specific agent, newest first). The agent-scoping distinguishes it from list_organization_evaluations and list_organization_evaluation_results without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the default lifecycle set (completed and failed) and how `status` overrides it, which is real usage context. It stops short of naming alternatives (e.g., list_organization_evaluations, retrieve_evaluation) or when-not to use it, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_flowsList saved flowsC
Read-onlyIdempotent

List saved flows Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
user_idNoThe user ID to for which to list items. Required if project_id is not provided.
project_idNoThe project ID for which to list items. Required if user_id is not provided.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered. The description's only behavioral statement, "Read operation," is redundant with those annotations and adds no new context such as scoping, return shape, or pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is short and front-loaded, but the second sentence ("Read operation") is pure redundancy against the readOnlyHint annotation and does not earn its place. Brevity here reflects under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering safety and a fully documented schema, the minimum viable information is present, and no output schema means return values need not be explained. However, a list tool with three scoping parameters and many sibling list_* tools would benefit from at least stating differentiation and the user_id/project_id selection rule.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with all three parameters (account, user_id, project_id) documented in the schema, including the either/or requirement between user_id and project_id. The description adds nothing beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ("List saved flows") that an agent can map to a concrete read operation. It does not, however, distinguish itself from related siblings such as list_workbooks, list_models, or start_flow/kill_flow, so the scope of "flows" versus other listable entities is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, nor any mention of prerequisites such as needing either user_id or project_id. "Read operation" is not usage guidance; it merely restates the read-only nature.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_mcp_server_promptsList MCP server promptsB
Read-onlyIdempotent

Return the prompt templates an MCP server exposes, fetched live from the server. When the server is not connected, prompts is empty and gumloop_auth_url is returned. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
team_idNo
server_idYes

TDQS

B3.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive and open-world, so the safety profile is covered. The description adds genuinely new behavioral context: results are fetched live, and the failure mode (server not connected) yields an empty prompts array plus a gumloop_auth_url, which tells the agent how to interpret and act on an empty result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and the resource noun. Each sentence carries distinct information (what it returns, the failure case, the safety classification), with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description does name the return fields (prompts, gumloop_auth_url), which is helpful. However, for a 3-parameter tool with a fully undocumented required parameter, the description remains incomplete on the input side.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%; server_id (the required parameter) and team_id have empty descriptions. The description compensates for none of this gap, providing no explanation of what server_id must reference or when to supply account/team_id, so an agent must guess at the required input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Return) and resource (prompt templates an MCP server exposes), which clearly distinguishes it from siblings like list_mcp_server_tools and list_mcp_server_resources that return different resource types. It does not explicitly name a sibling to route against, but the resource noun is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use versus when-not-to-use guidance and no mention of the closest alternatives (get_mcp_server_prompt for a single prompt, or list_mcp_server_tools/resources for other entity types). The usage is only implied by the tool name and resource noun.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_mcp_server_resourcesList MCP server resourcesA
Read-onlyIdempotent

Return the resources an MCP server exposes, fetched live from the server. When the server is not connected, resources is empty and gumloop_auth_url is returned so the caller can prompt the user to authenticate. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoOpaque cursor from a previous response's `next_cursor`.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
team_idNoScope the lookup to a single team.
server_idYesIdentifier of the MCP server.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered elsewhere. The description goes beyond that by disclosing the empty-resources-on-disconnect behavior and the gumloop_auth_url return, which is valuable operational context not in the annotations. It doesn't describe pagination behavior despite the cursor parameter, keeping it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: what it returns, the disconnected-state behavior, and the operation type. No filler and the core purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with full schema coverage and annotations, the description covers the key edge case (disconnected server) and auth flow. It omits any mention of cursor-based pagination despite accepting a cursor, which is a small but real gap given no output schema explains next_cursor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all four parameters including cursor, account, team_id, and server_id are documented in the schema itself. The description mentions the server concept but adds no syntax, format, or scoping detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (return) and resource (the resources an MCP server exposes), with the live-fetch qualifier that distinguishes it from static listings. An agent can tell this apart from siblings like list_mcp_server_tools, list_mcp_server_prompts, and read_mcp_server_resource without consulting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by stating it returns the resource list and explaining the auth flow, but never names an alternative or says when to prefer read_mcp_server_resource or list_mcp_server_tools instead. Usage context is inferable, not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_mcp_serversList MCP serversB
Read-onlyIdempotent

Return the catalog of MCP servers visible to the caller — Gumloop-hosted (gumcp_server), user-deployed Gumstack (gumstack_server), and custom (mcp_server) — along with each server's connection state. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
team_idNoScope the catalog to a single team. When omitted, returns servers visible to the authenticated user.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds that the catalog is scoped to what is visible to the caller and includes connection state, which is useful context. However, it does not add richer behavioral details such as permissions, pagination, or rate-limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and remains short. The trailing phrase 'Read operation.' is somewhat redundant because annotations already indicate a read-only operation, but the overall text is appropriately sized and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list tool with no output schema, the description adequately explains what is returned: a catalog of MCP servers with connection state. It omits details about the exact return structure or pagination, but the annotations and fully documented parameters carry most of the remaining context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters ('account' and 'team_id') are documented in the schema, including the default visibility behavior when team_id is omitted. The description adds no parameter syntax or format details beyond what the schema already provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Return the catalog of MCP servers.' It also names the three server categories and says the result includes connection state. It does not explicitly distinguish itself from sibling tools like list_agent_mcp_servers or retrieve_mcp_server, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives such as list_agent_mcp_servers, retrieve_mcp_server, or list_mcp_server_tools. The phrase 'visible to the caller' implies a default scope, but no when/when-not conditions or alternative routing are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_mcp_server_toolsList MCP server toolsA
Read-onlyIdempotent

Return the tools exposed by an MCP server. When the server is not in connected state, tools is empty and gumloop_auth_url is returned so the caller can prompt the user to authenticate. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
team_idNoScope the lookup to a single team.
server_idYesIdentifier of the MCP server.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the bar is lower, and the description adds genuinely useful behavior: when the server is not connected, `tools` is empty and `gumloop_auth_url` is returned for an auth prompt. That downstream consequence is not derivable from annotations. The trailing 'Read operation' merely restates readOnlyHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core purpose front-loaded, then the edge-case behavior. No filler, though the 'Read operation' tail is redundant with annotations and slightly dilutes the efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries responsibility for return semantics and does so by explaining the empty-tools/auth-url case. It is complete enough for a listing tool, missing only explicit routing guidance to sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so account, team_id, and server_id are fully described in the schema; baseline is 3. The description adds no parameter-level meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Return the tools exposed by an MCP server.' This distinguishes it from siblings like list_mcp_server_prompts and list_mcp_server_resources. However, it does not explicitly name or contrast with those siblings, leaving differentiation to the reader.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied (list tools before calling them) but never stated. There is no explicit when-to-use, when-not-to-use, or mention of alternatives such as call_mcp_tools or retrieve_mcp_server. The connected-state caveat gives partial context but not selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_modelsList modelsB
Read-onlyIdempotent

List the LLMs and preset model chains available to the caller, grouped for display in a model picker. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
team_idNoScope model availability to a specific team. When omitted, uses the authenticated user's default organization.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so safety is covered; the trailing 'Read operation.' sentence largely repeats that. The genuinely additive content is the hint that results are 'grouped for display in a model picker,' which describes the shape of the response and partially compensates for the missing output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the resource being listed. The only waste is 'Read operation.', which duplicates readOnlyHint from the annotations rather than earning its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter listing tool with full schema coverage and safety annotations, the description supplies the one thing structured fields do not: that results are grouped for a picker UI. Remaining gaps (result ordering, size, pagination) are minor for this tool class and no output schema exists to lean on.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both optional parameters (account, team_id) are documented in the schema itself, so the baseline is 3. The description adds no additional meaning about what these filters do to the returned model list.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'List the LLMs and preset model chains available to the caller.' The scope ('available to the caller') is precise, but it never mentions the neighboring route_model tool, so sibling differentiation is left implicit rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through 'grouped for display in a model picker' but gives no explicit when-to-use guidance and no exclusions. With route_model in the sibling list, an agent would benefit from knowing this is a discovery/enumeration call rather than a dispatch call, and that distinction is absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_organization_evaluation_resultsList evaluation resultsA
Read-onlyIdempotent

Cursor-paginated results for one evaluation across every agent it grades, newest first. Each session appears once with its latest result; queued and in-progress results are included so a run can be followed to completion. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
gradeNoFilter by grade.
cursorNoPagination cursor from a previous response's `next_cursor`.
statusNoFilter by status.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
agent_idNoOnly results for this agent.
page_sizeNoItems per page (1-100).
session_idNoOnly results for this session.
created_afterNoOnly results created at or after this time. RFC 3339 with an explicit offset (for example `2026-09-01T00:00:00Z`).
evaluation_idYesID of the organization evaluation.
created_beforeNoOnly results created before this time. RFC 3339 with an explicit offset.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations carry the safety profile (readOnlyHint, idempotentHint, openWorldHint, destructiveHint=false), so the bar is lower. The description still adds real behavioral value beyond them: newest-first ordering, session deduplication, and the inclusion of queued/in-progress results that make it suitable for polling a run.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with zero padding, front-loading the cursor-paginated scope and ordering before the dedup rule. 'Read operation.' is a useful one-word safety restatement at the end.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete enough given a 10-param schema with 100% coverage and annotations covering the safety profile. The only gap is the absence of any routing hint to the sibling singleton result tool, which would help in a list of ~90 siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all 10 params, including enums for grade and status and filter semantics. The description adds no parameter-level detail beyond what the schema already states, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (list) + resource (evaluation results for one evaluation) with explicit scope: 'across every agent it grades, newest first'. Names the ordering and the deduplication rule ('each session appears once with its latest result'), which distinguishes it from get_organization_evaluation_result (singular) in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the name and content, and the description notes that queued/in-progress results are included 'so a run can be followed to completion' — a hint about when to use it. But there is no explicit when-to-use vs get_organization_evaluation_result / get_organization_evaluation_metrics, and no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_organization_evaluationsList evaluationsA
Read-onlyIdempotent

Cursor-paginated list of an organization's evaluations with their targets, coverage, and result rollups. Requires the organization:manage_evaluations permission (Enterprise plan). Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoPagination cursor from a previous response's `next_cursor`.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
page_sizeNoItems per page (1-100).
organization_idYesThe organization whose evaluations to list.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive. The description adds genuinely useful context beyond them: the required permission scope (organization:manage_evaluations), the Enterprise plan requirement, and that results are cursor-paginated. It does not, however, describe pagination behavior in more depth or any rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose and return content, then the permission gate and read-only nature. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly names the returned data (targets, coverage, result rollups) and the permission prerequisite. Pagination is mentioned and the parameters are fully schema-documented. Minor gap: no guidance on navigating to related results or metric tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (cursor, account, page_size, organization_id) are already documented. The description only echoes the cursor-pagination concept, adding no syntax or format meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (list) and resource (an organization's evaluations) and details what is returned: targets, coverage, and result rollups. It is clearly distinguishable in scope from the bare list_evaluations sibling, though it never explicitly names that sibling to route the agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'Read operation' and the permission note, but there is no explicit when-to-use/when-not guidance or reference to the alternatives (e.g., list_evaluations, get_organization_evaluation). The prerequisite gate (Enterprise plan, organization:manage_evaluations) is the only real usage signal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_organizationsList organizationsA
Read-onlyIdempotent

Returns the organization the authenticated user belongs to. Use its id as organization_id on the evaluation endpoints. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The trailing 'Read operation.' simply restates readOnlyHint, and the description does not disclose return shape or credential/identity behavior beyond what the annotation set provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with the core purpose front-loaded and the downstream usage tip immediately after. The 'Read operation.' sentence is mildly redundant against readOnlyHint but does not harm readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, no-output-schema tool, the description conveys what is returned and how to consume it, which is largely sufficient. It stops short of describing the organization fields returned, a minor gap given there is no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single optional `account` parameter is already well documented in the schema as selecting private credentials and identity. The description adds nothing beyond what the schema states, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns the organization the authenticated user belongs to'), clarifying the singular scope despite the plural-sounding name 'list_organizations'. It is distinguishable from list_accounts, which covers a different concept, though the description does not explicitly name siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tells the agent the downstream purpose of the call: feed the returned `id` into the evaluation endpoints as `organization_id`. That gives clear context for when the tool is useful, but offers no explicit alternatives or conditions for skipping it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_queued_messagesList queued messagesA
Read-onlyIdempotent

List the messages waiting in a session's queue, in the order they will be sent. Queued messages are drained automatically when the agent finishes its current turn. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
session_idYesID of the session whose queue to list.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), so the bar is lower, yet the description adds genuine behavioral value: queued messages are drained automatically at the end of the current turn and are returned in send order. It does not describe message contents or pagination, but the added lifecycle context is meaningful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two efficient sentences front-load the core action and the queue-drain behavior adds real value. The trailing sentence 'Read operation.' is redundant with readOnlyHint=true and could be dropped.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter read tool with full annotation coverage and no output schema, the description supplies enough to call it correctly and understand the queue lifecycle. A brief note on what a queued message looks like would close the remaining gap left by the absent output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so session_id and account are already fully documented in the schema, giving a baseline of 3. The description adds no syntax, format, or scoping detail beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the messages waiting in a session's queue') and adds ordering ('in the order they will be sent'), which clearly identifies it as the read-side counterpart to the queue message tools. It does not explicitly name siblings like update_queued_message or delete_queued_message, so the differentiation is implicit rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains queue mechanics ('drained automatically when the agent finishes its current turn') but never says when an agent should call this versus alternatives such as list_sessions, update_queued_message, or send_queued_message. No conditions or exclusions are offered, leaving usage to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_role_credit_limitsList custom role credit limitsA
Read-onlyIdempotent

This endpoint lists every active custom role in an organization together with its monthly credit limit, so external systems can manage credit limits programmatically. A monthly_credit_limit of null means the role sets no limit of its own. The limit applies to each member of the role individually; when a user belongs to multiple roles, the highest limit across their roles wins. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoOpaque cursor from a previous response's `next_cursor`; omit for the first page.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
user_idNoYour user id -- you must be an organization admin to manage custom role credit limits.
page_sizeNoNumber of roles per page.
organization_idYesThe ID of the organization.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuine domain semantics: null means no self-imposed limit, the limit applies per member individually, and with multiple roles the highest limit wins. It does not mention pagination behavior, though the cursor param covers that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and scope, followed by meaningful edge-case semantics. The trailing 'Read operation.' is redundant given readOnlyHint=true, a small bit of waste in an otherwise tight description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully explains the key return-value nuance (null vs. an actual limit) and multi-role resolution. Combined with the fully documented input schema and annotations, an agent has what it needs, though explicit sibling routing would round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented in the schema, including the admin requirement on user_id and cursor semantics. The description adds no parameter-level detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: it lists every active custom role in an organization along with each role's monthly credit limit. The 'every ... role' scope implicitly separates it from the singular get_role_credit_limit and set_role_credit_limit siblings, but it never names those alternatives explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The stated motivation ('so external systems can manage credit limits programmatically') implies when the tool is useful, but there is no explicit when-to-use/when-not-to-use guidance or comparison against get_role_credit_limit or set_role_credit_limit. Usage is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_sessionsList sessionsA
Read-onlyIdempotent

List sessions for an agent with cursor-based pagination, optional filtering, and search. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNoFilter sessions by type (e.g. `api`, `web`, `slack`).
stateNoFilter sessions by state.
cursorNoCursor for the next page of results. Use the `next_cursor` value from a previous response.
searchNoFree-text search query to filter sessions by name or content. Also accepted as `search_query`.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
agent_idYesID of the agent whose sessions to list. Also accepts the reserved aliases `gumball` and `analytics`.
page_sizeNoNumber of sessions to return per page. Defaults to `20`, maximum `100`.
sort_orderNoSort order for the results (e.g. `newest` or `oldest`).
trigger_idNoFilter sessions by the trigger that initiated them.
creator_user_idNoFilter sessions by the user who created them.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so "Read operation" adds nothing. The description does contribute cursor-based pagination behavior, which the annotations do not cover, so partial credit above baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, capability summary front-loaded before the redundant "Read operation" tag. Minimal waste, though the trailing sentence duplicates annotation content and earns nothing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter read-only list tool whose schema fully documents inputs and whose annotations carry the safety profile, the description covers the essentials. It omits any hint about output shape (result fields, cursor location), which is a minor gap rather than a blocking one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all ten parameters are already documented in the schema (enum states, cursor semantics, page_size bounds, aliases). The description only summarizes categories of filtering generically, adding no syntax or constraint detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("List sessions for an agent") plus the key capabilities (pagination, filtering, search). It does not explicitly differentiate from siblings like retrieve_session or list_queued_messages, so it stays short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Scoping the list to a specific agent implies the usage context, and mentioning optional filtering/search hints at the tool's role, but there is no explicit when-to-use vs retrieve_session or when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_skillsList skillsA
Read-onlyIdempotent

List skills the caller has access to. Filter by team, search by name, or narrow to a specific creator, related MCP server, or agent. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoOpaque pagination cursor returned in `next_cursor` from a prior page.
unusedNoWhen set, filters to skills that have not been used.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
team_idNoScope the listing to a single team. When omitted, returns skills owned by the authenticated user.
agent_idNoFilter to skills attached to this agent.
page_sizeNoNumber of skills per page. Clamped between 1 and 100.
sort_orderNoSort order for the returned skills.newest
search_queryNoCase-insensitive substring match against the skill name.
creator_user_idNoFilter to skills created by this user ID.
related_server_idNoFilter to skills that reference this MCP server ID in their metadata.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, and the description's trailing "Read operation." merely restates that. The one genuinely additive detail is the access-scoping of results ("skills the caller has access to"), but pagination and ordering behavior remain unaddressed beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences that front-load the core purpose before the filter list. The trailing fragment "Read operation." is redundant with the annotations and adds a small amount of waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with full schema coverage, comprehensive annotations and no output schema, the description covers the essential scope and filterable dimensions. It omits any mention of paging or result volume, which the schema partially compensates for.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented in structured form. The description only groups the filter dimensions at a high level and adds no format or syntax detail beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("List skills the caller has access to") and scopes the result set to the caller's permissions, which clearly separates it from create_skill/update_skill/delete_skill. It stops short of naming an alternative tool explicitly, so it is clear but not sibling-differentiating.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description enumerates the available filtering axes (team, name, creator, related MCP server, agent), which implies when the tool is useful. However, these are the same facets already documented in the schema, and it gives no explicit when-to-use/when-not guidance or exclusion against siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_teamsList teamsC
Read-onlyIdempotent

List teams the authenticated caller belongs to. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, and the description's 'Read operation' merely restates readOnlyHint. It adds the caller-scoped filtering detail but says nothing about how the optional account parameter selects credentials, result size, ordering, or pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the scope front-loaded, which is efficient. The trailing 'Read operation.' sentence is wasted space because it duplicates readOnlyHint=true already in the annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read tool with full annotation coverage and no output schema, the description is minimally adequate. It does not tell the agent what a returned team looks like or what identifiers it yields for downstream calls, which would be the useful addition here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single optional 'account' parameter is already documented in the schema as selecting private credentials and identity. The description adds no syntax, default, or behavior detail beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('List teams') plus the scope ('the authenticated caller belongs to'), which is meaningfully more precise than the title. It does not, however, distinguish itself from similar listing siblings such as list_organizations or list_accounts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no named alternative. An agent must infer from the name alone when this tool is preferable to the other list_* tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_workbooksList workbooks and their saved flowsC
Read-onlyIdempotent

List workbooks and their saved flows Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
user_idNoThe user ID for which to list workbooks. Required if project_id is not provided.
project_idNoThe project ID for which to list workbooks. Required if user_id is not provided.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the 'Read operation.' sentence is a pure restatement of structured data and adds no new behavioral context. Nothing is said about scoping, credential selection via the account parameter, or result size.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short and front-loaded, with the resource named in the first clause. The second line ('Read operation.') is arguably redundant given the annotations, which slightly dilutes efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless-required list tool with rich annotations and full schema coverage, most needs are met. With no output schema, though, the description could say more about the shape of the returned workbook/flow data, which it does only vaguely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the account, user_id and project_id parameters including their mutual exclusivity. The description adds no parameter meaning beyond that, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List workbooks and their saved flows'), which is clearly more informative than the bare name. However, it offers no differentiation from siblings like list_flows or list_agents, and doesn't clarify the relationship between workbooks and their saved flows.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus list_flows or other list endpoints, and no mention that the user_id/project_id parameters are mutually exclusive alternatives. The agent must infer usage entirely from the schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_permission_group_usersManage custom role usersA
Destructive

This endpoint allows organization administrators to add or remove users from a custom role (formerly "permission group"). Adding a user to a role does not remove them from any other role they belong to. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNoThe action to perform - either 'add' or 'remove' a user.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
user_idNoYour user id -- you must be an organization admin to manage custom role users.
group_idNoThe ID of the custom role to manage users for.
user_emailNoThe email address of the target user to add or remove.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
organization_idNoThe ID of the organization that the custom role belongs to.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructive=true and idempotent=false. The description adds genuinely non-obvious behavior: admin-only authorization, that an add does not revoke other role memberships, that explicit confirmation is required, and a warning against auto-resubmitting unknown outcomes. That goes meaningfully beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core capability, which is good, but the second paragraph mixes a relevant confirmation rule with generic boilerplate ("Runs can spend credits or trigger downstream actions") that is not specific to this account-operation tool and dilutes the message.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent mutation with no output schema, the description covers authorization, the confirmation gate, and the key membership side effect, which is enough for correct invocation. It leaves the relationship to manage_project_users and the account/payload precedence unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented in the schema. The description only echoes the add/remove semantics and admin requirement without adding format or usage detail for the nine parameters (e.g., account vs organization_id interplay), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (add or remove) and resource (users in a custom role, formerly "permission group") and scopes it to organization administrators. The analogous sibling manage_project_users is never referenced, so the agent must infer the distinction from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a prerequisite (must be an org admin) and a useful side-effect rule (adding does not remove from other roles), plus a confirmation requirement. However, it never states when to prefer this tool over manage_project_users or other role-related siblings, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_project_usersManage workspace usersB
Destructive

This endpoint allows organization administrators to add or remove users from a workspace. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNoThe action to perform - either 'add' or 'remove' a user.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
user_idNoYour user id -- you must be an organization admin to manage workspace users.
is_adminNoWhen adding a user, specify whether they should have admin privileges (default is false).
user_emailNoThe email address of the target user to add or remove.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
workspace_idNoThe ID of the workspace to manage users for.
organization_idNoThe ID of the organization that the workspace belongs to.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, openWorldHint=true, and idempotentHint=false. The description adds auth requirements (admin role) and a warning about credits/downstream actions plus avoiding automatic resubmission of unknown outcomes—useful, but somewhat generic and not specific to this endpoint's actual side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences, front-loaded with the core purpose, then the confirmation requirement, then the caution. Efficient, though the last sentence is a generic boilerplate warning that could apply to many tools.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter mutation tool with nested payload, multiple input modes (flags, payload, file), no output schema, and destructive semantics, the description is adequate but thin. It doesn't explain the interplay of payload vs flags vs payload_file, which is a significant omission for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema fully documents all 10 parameters including the nested payload and file alternatives. The description adds no parameter-level detail beyond the schema, which is the baseline case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: adding or removing users from a workspace, scoped to organization administrators. It's clear, though it doesn't explicitly differentiate from the highly related sibling manage_permission_group_users.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names the actor (organization administrators) and requires explicit confirmation, which implies usage context. But it doesn't state when to use this versus alternatives like manage_permission_group_users, nor when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

queue_session_messageQueue messageA
Destructive

Add a message to a session's queue instead of interrupting the agent. Queued messages are sent automatically, in order, when the agent finishes its current turn. A session's queue holds at most 20 messages.

To interrupt the current turn and send a queued message immediately, use Send queued message now. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputNoThe message to queue. Cannot be empty. Also accepted as `message` for backwards compatibility.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
session_idYesID of the session to queue the message on.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare mutation, non-idempotency, and destructiveness, but the description adds behavioral detail they cannot carry: the 20-message capacity cap, automatic ordered delivery when the current turn ends, the confirmation requirement, and the credit/downstream-action warning. This is meaningful context beyond the annotation set, with no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, followed by the alternative and then the confirmation warning; every sentence carries weight. The markdown link syntax in the routing sentence is slightly awkward but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter tool with a nested payload object and no output schema, the description covers purpose, alternative, capacity, ordering, and the confirmation/credits caveat. The remaining parameter distinctions are fully handled by the 100%-covered schema, so nothing an agent needs is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (input, account, confirm, payload, session_id, payload_file) is already documented. The description only references 'the message to queue' and the capacity limit, adding little beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Add a message to a session's queue') and immediately frames it against the interrupting alternative, distinguishing it from send_message. An agent can tell it apart from the sibling queued-message tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the condition for use (queue instead of interrupting) and routes to the alternative (send_queued_message for immediate interruption), plus a confirmation precondition. When-to-use, when-not, and the alternative are all stated rather than inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_mcp_server_resourceRead MCP server resourceA
Read-onlyIdempotent

Read one resource by uri. Each content item is either text or a base64 blob, never both. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
uriYesThe resource `uri` from List MCP server resources.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
team_idNo
server_idYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, openWorld and non-destructive, so safety is covered; the description adds genuinely non-obvious return semantics — each content item is either `text` or a base64 `blob`, never both. It omits auth/credential behavior for the `account` parameter and any error/rate-limit context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and resource, then the return-shape caveat. The trailing 'Read operation.' is largely redundant with readOnlyHint=true, but the total footprint is small enough that it does not waste much space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no output schema and two undocumented parameters, the definition is on the thin side; it clarifies the payload shape but not what a 'resource' is relative to MCP prompts/tools, nor the team/account scoping needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: `uri` and `account` are documented while `team_id` and `server_id` have empty descriptions. The description explains `uri` and where to obtain it, but adds nothing for `server_id` or `account`, so half the parameters remain unexplained in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read one resource by `uri`'), and the singular 'one resource' implicitly contrasts with the sibling list_mcp_server_resources. It never names a sibling explicitly, so an agent must infer the list-then-read relationship rather than being told it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the schema note on `uri` ('from List MCP server resources') hints at the list-then-read workflow, but the description gives no explicit when-to-use or when-not guidance and does not distinguish this from get_mcp_server_prompt or call_mcp_tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rename_sessionRename sessionA
Destructive

Rename a session. name is the only mutable field; it is trimmed and must be between 1 and 256 characters after trimming.

Returns the full session, in the same shape as Retrieve session. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
session_idYesID of the session to rename.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive=true, idempotent=false and openWorld=true; the description adds the confirmation gate, the trimming/normalization rule, and the return shape, which are not in the annotations. The credit/downstream-action sentence is fairly generic boilerplate that reads as policy text rather than tool-specific behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the operation and the key constraint, then the return shape. Efficient overall, though the closing sentence about credits and automatic resubmission is generic and not obviously tied to a rename.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description compensates by stating the return is the full session in retrieve_session's shape. The confirmation requirement and normalization rule are covered; the remaining gap is the unexplained relationship among `name`, `payload` and `payload_file`.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 83%, so the schema already documents `name`, `session_id`, `account`, `confirm` and the payload options; the description mostly repeats the trimming/1-256 rule. It does not explain how `payload`/`payload_file` relate to the body flags, which is the one parameter relationship an agent still has to guess at.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Rename a session') and immediately narrows scope by naming the only mutable field. It also anchors the return shape to the sibling `retrieve_session`, so an agent knows exactly what comes back without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real when-to-use context: explicit confirmation is required for this exact operation, and the description warns against automatic resubmission of unknown outcomes. It does not enumerate alternatives (there is no competing rename sibling), so it falls just short of explicit when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_session_approvalsResolve approvalsA
Destructive

Answer pending asks on a session that is paused in the approval_required state — tool approvals, human input requests, and checkpoints.

List the pending asks with Retrieve session: each entry in pending_approvals carries the action_request_id to answer, and human_input asks include the questions to fill in via response.values. Resolutions are processed in order; the agent resumes once the pending asks are answered. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
session_idYesID of the session with pending approvals.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
approval_responsesNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (destructiveHint=true, openWorldHint=true, idempotentHint=false) cover safety, and the description adds substantive context: credit spend, downstream actions, ordered processing of resolutions, and the resume-once-answered behavior. The explicit confirmation requirement maps to `confirm` and the warning about unknown outcomes is high-value operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences front-loaded with the purpose, then retrieval guidance, then safety warning. Dense but each sentence earns its place. Minor density concern but no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the workflow (retrieve then resolve), the ordered processing semantics, and the confirmation safety. No output schema exists; return values are not described, and the payload/payload_file alternatives are left to the schema. Adequate for a complex mutation tool with 6 params.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the schema documents most parameters (action, reason, response.values, action_request_id). The description adds meaning to `action_request_id` and `response.values` sourcing and confirms the confirm/credits semantics, but doesn't clarify payload vs payload_file vs body flags alternation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (resolve/answer) and resource (pending asks on a paused session) with the exact precondition 'paused in the `approval_required` state'. Distinguishes its scope (tool approvals, human input, checkpoints) in a way no sibling replicates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent to first call Retrieve session to list pending asks and where to find `action_request_id` and `questions`. Also warns not to resubmit unknown outcomes automatically. Lacks an explicit 'when NOT to use' against siblings like send_message or queue_session_message, but the prerequisite condition is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retrieve_agentRetrieve agentA
Read-onlyIdempotent

Retrieve a single agent by ID.

In addition to regular agent IDs, agent_id accepts the reserved aliases gumball (your personal Gumball agent) and analytics (your analytics agent) on all agent-scoped endpoints. The alias resolves to your own copy of the platform agent, creating it on first use, and responses report the alias back as the agent's id. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
agent_idYesID of the agent to retrieve. Also accepts the reserved aliases `gumball` and `analytics`.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely non-obvious behavior: passing the `gumball` or `analytics` alias resolves to the caller's own copy and 'creates it on first use' — a side effect not captured by any annotation — plus how the alias is echoed back in responses. It stops short of describing error behavior or permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, followed by the alias note, with no wasted preamble. The trailing 'Read operation.' partially duplicates the readOnlyHint annotation and could be dropped, but the text is otherwise tight and well ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read tool with 100% schema coverage and full annotations, the description covers the essentials and adds the alias-resolution behavior the agent could not otherwise know. With no output schema it also hints at the return shape ('responses report the alias back as the agent's id'), though it omits error and not-found behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, and the schema already documents both `account` and the alias-accepting `agent_id`. The description goes beyond the schema by explaining what the alias actually resolves to and the create-on-first-use semantics, adding real meaning to the `agent_id` parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource with scope: 'Retrieve a single agent by ID,' which implicitly separates it from list_agents and the create/update siblings. However, it never names an alternative tool, so differentiation is by inference rather than explicit contrast, keeping it just below a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the singular 'by ID' framing, but there is no explicit when-to-use guidance, no statement of when to prefer this over list_agents or retrieve_agent_version, and no exclusions or prerequisites. The alias discussion is contextual but not routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retrieve_agent_versionRetrieve agent versionA
Read-onlyIdempotent

Retrieve one immutable agent version: its full configuration (composition) plus the structured changes relative to the version before it. Use it to export an agent's configuration or to audit what changed between versions.

changes is null for the first version of an agent, since there is no predecessor to diff against. Versions created before attachment snapshots were recorded report composition.complete: false (and changes.attachment_changes_complete: false); their skill_ids and knowledge_sources are null rather than empty. Skill file contents are never included, and this endpoint is read-only — it cannot restore or deploy a version.

Requires configuration access on the agent — callers limited to using the agent (no configuration access) get a 403. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
agent_idYesID of the agent the version belongs to. Also accepts the reserved aliases `gumball` and `analytics`.
version_idYesID of the version to retrieve, from `GET /agents/{agent_id}/versions`.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral context beyond the annotations: it explains that `changes` is null for the first version, that older versions report `composition.complete: false` and `changes.attachment_changes_complete: false`, that `skill_ids` and `knowledge_sources` may be null rather than empty, that skill file contents are never included, and that configuration access is required with a `403` otherwise.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose and usage in the first two sentences, then uses short paragraphs for edge cases and access requirements. Every sentence contributes necessary information without repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description explains the key returned fields (`composition`, `changes`), edge cases such as null values and incomplete snapshots, and the permission model. This is sufficient for an agent to call and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the input schema already documents all three parameters, including the meaning of `agent_id` aliases and where `version_id` comes from. The description does not add parameter-level syntax or format details beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb and resource: 'Retrieve one immutable agent version', and specifies the returned content as 'full configuration (`composition`) plus the structured `changes` relative to the version before it.' This clearly distinguishes it from listing versions or retrieving a whole agent, even without naming sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance: 'Use it to export an agent's configuration or to audit what changed between versions.' It also states a when-not boundary by noting the endpoint is read-only and 'cannot restore or deploy a version.' However, it does not name alternative tools such as `list_agent_versions` for related operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retrieve_evaluationRetrieve evaluationA
Read-onlyIdempotent

Retrieve a single evaluation result by ID. The evaluation must belong to the specified agent. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
agent_idYesID of the agent the evaluation belongs to. Also accepts the reserved aliases `gumball` and `analytics`.
evaluation_idYesID of the evaluation to retrieve.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint and openWorldHint, so the safety profile is fully covered. The description's 'Read operation' simply restates readOnlyHint, adding no new behavioral detail; the only real addition is the agent-ownership constraint, and nothing is said about error behavior when ownership fails.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the primary action front-loaded and no wasted prose. The trailing 'Read operation' clause is redundant with the readOnlyHint annotation, which is a minor inefficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only fetch by ID with full schema coverage, complete annotations and no nested objects, the definition covers what an agent needs to call it correctly. The absence of an output schema means the return shape is undocumented, a small residual gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (including the account selector, aliases, and min lengths) are already documented in the schema. The description only reinforces the agent_id/evaluation_id pairing and adds no syntax or format detail, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Retrieve a single evaluation result by ID'), and the agent/ownership scoping sets it apart from list-style siblings like list_evaluations. It does not explicitly name the closest alternative (get_organization_evaluation), so sibling differentiation is only partial.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The constraint 'must belong to the specified agent' gives useful context for when this call is valid, but there is no explicit when-to-use guidance or exclusion versus get_organization_evaluation / list_evaluations. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retrieve_mcp_serverRetrieve an MCP serverB
Read-onlyIdempotent

Return a single MCP server. The response populates allowed_tool_call_ids with the tool call IDs the caller is permitted to invoke on this server. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
team_idNoScope the lookup to a single team.
server_idYesIdentifier of the MCP server to retrieve.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so "Read operation" adds nothing. The genuinely useful addition is the disclosure that the response populates allowed_tool_call_ids with the caller's permitted tool calls — a return-side behavior not covered by annotations or any output schema. It does not mention auth needs, rate limits, or error behavior, so a 3 fits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with the purpose front-loaded, followed by the output detail. The trailing "Read operation" is redundant against readOnlyHint and slightly wastes space, but overall the entry is tight and well-ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-resource retrieve with 100% schema coverage and annotations carrying the safety profile, the description is nearly enough: it names the resource and one key response field. It omits any mention of the plural-list alternative, which is the main missing piece.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so account, team_id, and server_id are already documented in the schema. The description adds no parameter-level detail (e.g., whether account changes the scoping or what happens when team_id is omitted), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: "Return a single MCP server." The singular "single" implicitly distinguishes it from the sibling list_mcp_servers, so an agent can route without opening the schema. It stops short of explicitly naming the sibling as an alternative, so it earns a 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use or when-not guidance and never references the surrounding toolset (list_mcp_servers, list_mcp_server_tools). Retrieval-by-id is implied by the required server_id, but nothing steers an agent between this and the list siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retrieve_sessionRetrieve sessionA
Read-onlyIdempotent

Retrieve a session by ID, including its messages, current state, agent metadata, and participants. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
session_idYesID of the session to retrieve.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the appended 'Read operation' is largely redundant. The description does add useful disclosure that the payload includes messages, state, agent metadata, and participants, but nothing about auth, rate limits, or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the primary purpose and returned payload front-loaded and no wasted wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the returned content, which is the key missing structured information. It is largely complete for a read-by-ID tool, though it omits error/not-found behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both account and session_id are already documented in the schema. The description reinforces that retrieval is keyed on the ID but adds no syntax or format detail beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Retrieve) and resource (session) and enumerates what is returned (messages, state, agent metadata, participants), which distinguishes it from list_sessions. It could be sharper by explicitly contrasting with the sibling list_sessions, but the 'by ID' scope makes the difference clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'by ID' phrasing implies the tool is for when a specific session identifier is already known, but there is no explicit when-to-use routing against alternatives like list_sessions or get_run_details. Usage is merely implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

route_modelRoute a message to a modelA
Read-onlyIdempotent

Ask Gumloop Chew — the model router behind Auto — which model it would use for a given message, and why. Chew is a router, not a model: router is always gumloop-chew, and the concrete model it selected is route.model.

This endpoint returns a decision only. It does not run the selected model or its fallbacks. The routing judgement consumes credits and is bounded by the caller's model access.

Omit models to use the deduplicated union of Chew's lane chains, not every model the caller may use. Restricted candidates can appear with status: "restricted"; they are never selected or included in fallback_models.

Team scope requires actual team membership, even within the same organization. Personal API keys and OAuth are supported; this is not team-key-only. Read operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNo
inputNo
modelsNo
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
historyNo
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
team_idNoScope model availability and credit attribution to a team the caller belongs to.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description still adds real context beyond them: the call consumes credits, it is bounded by the caller's model access, it never executes the selected model or its fallbacks, and restricted candidates surface with status "restricted" but are never selected. Return-shape detail (fallback_models) is thin, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then decision-only semantics, then parameter caveats. Every paragraph is topical, though the second paragraph's 'Chew is a router, not a model' partially restates the opening line, which costs a little density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a decision-only routing endpoint with no output schema, the description supplies the key return contract (router is always gumloop-chew, selected model at route.model, restricted status, fallback_models exclusion) plus auth/scope prerequisites. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% and 8 parameters exist, so the description must compensate. It clarifies `models` (omit -> lane-chain union rather than every available model) and `team_id` scoping, but leaves `account`, `payload`, and `payload_file` unexplained, and the `models` note largely repeats the schema's own 'Omit to use Chew's lane-chain union'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('ask the model router which model it would use') and immediately distinguishes itself from the neighbours: 'Chew is a router, not a model' and 'This endpoint returns a decision only. It does not run the selected model.' An agent can separate this from create_chat_completion or list_models without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage context: omit `models` to fall back to the lane-chain union, team scope requires actual membership, and personal API keys/OAuth are supported. It never explicitly names an alternative tool or states a when-not-to-use condition (e.g. 'use create_chat_completion to actually run a model'), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_evaluationsRun evaluationsA
Destructive

Grades up to 200 of the agent's finished sessions with its own evaluation configuration. Grading is asynchronous: each accepted session gets a result with status: queued; poll it with GET /agents/{agent_id}/evaluations/{evaluation_id} until it is completed or failed. A new result replaces the previous result for that session.

Sessions are skipped, not rejected, when they are unfinished, incognito, or not owned by this agent (ineligible), or already have a queued or running result (in_flight, with the existing result_id). The caller is charged one credit per queued session. Set dry_run: true to see the cost and skips without queuing anything.

Requires edit access on the agent and a plan with evaluations enabled. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
dry_runNoReport cost and skipped sessions without queuing.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
agent_idYesID of the agent that owns the sessions.
session_idsNoSessions to grade. Duplicates are rejected.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructiveHint/openWorldHint only flag risk). It discloses asynchronous queuing with status: queued, the polling endpoint to reach completed/failed, replacement of prior results, skip-not-reject semantics with ineligible/in_flight reasons and result_id, per-session credit cost, and the confirmation/anti-resubmission rule. This is unusually complete behavioral disclosure for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the verb, resource and cap in the first sentence, then layers async behavior, skip semantics, cost, and prerequisites in short paragraphs. Slightly dense across four paragraphs, and the closing confirmation boilerplate is the least information-dense part, but every block carries operational meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent mutation with no output schema, the description supplies the missing return information (queued status plus polling path), the cost model, permission prerequisites, and partial-success behavior. An agent has everything needed to invoke it correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so a 3 is the baseline, but the description adds genuine semantics: dry_run reports cost and skips without queuing, session_ids are skipped for unfinished/incognito/not-owned/already-in-flight sessions, and duplicates are rejected. This explains parameter consequences rather than restating the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (grades), the exact resource (the agent's finished sessions using its own evaluation configuration), and the hard cap (up to 200). It is clearly distinguishable from siblings like list_evaluations, retrieve_evaluation, get_evaluation_metrics and update_evaluation_config, which configure or read rather than execute grading.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives strong operating context: prerequisites (edit access on the agent, a plan with evaluations enabled), a recommended dry_run: true preview path, and an explicit confirmation requirement. It does not, however, name a sibling alternative for related tasks (e.g. configuring or reading evaluations), so the routing guidance is contextual rather than comparative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_organization_evaluationRun evaluation on sessionsA
Destructive

Grades up to 200 existing sessions with this evaluation. Grading is asynchronous: each accepted session gets a result with status: queued; poll it with GET /evaluations/{evaluation_id}/results/{result_id} until it is completed or failed.

Sessions are skipped, not rejected, when they are not completed sessions of an agent the evaluation covers (ineligible) or already have a queued or running result for this evaluation (in_flight, with the existing result_id). The caller is charged one credit per queued session. Set dry_run: true to see the cost and skips without queuing anything. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
dry_runNoReport cost and skipped sessions without queuing.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
session_idsNoSessions to grade. Duplicates are rejected.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
evaluation_idYesID of the organization evaluation.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Substantially exceeds the annotations: it discloses asynchronous grading, the queued status and the exact polling endpoint, skip-vs-reject semantics with the ineligible and in_flight reasons, the one-credit-per-queued-session charge, and a warning never to auto-resubmit unknown outcomes. This is consistent with destructiveHint=true and adds real operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action is front-loaded and the async/polling detail follows logically. Three short paragraphs with little filler, though the credit and confirmation warnings are slightly repetitive with schema descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, nested, async, credit-spending tool with no output schema, the description supplies everything an agent needs: batching limit, async flow, polling path, skip conditions, cost model, dry-run preview and confirmation requirement. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents account, confirm, dry_run, session_ids, payload and payload_file. The description restates the dry_run effect and the 200-session cap but adds no syntax or semantics beyond the structured fields, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb and resource ('Grades up to 200 existing sessions with this evaluation'), so an agent immediately knows what the tool does. It does not, however, explicitly distinguish itself from the sibling run_evaluations, which a 5 would require.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear situational guidance: use dry_run:true to preview cost and skips without queuing, and confirmation is required for this exact account operation. It explains skip conditions (ineligible, in_flight) but never names an alternative sibling such as run_evaluations, so it stops short of full when/when-not/alternative coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_brainSearch Company BrainA
Destructive

Run a hybrid (semantic + keyword) search across the knowledge sources indexed in your Company Brain and return the most relevant, ranked snippets with citations.

Results are scoped to what the authenticated user can see: personal sources, plus any team and organization sources shared with them. Requires the Brain feature, which is available on the Pro and Enterprise plans. Each search consumes Gumloop credits. Explicit confirmation is required for this credit-consuming Brain search.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of results to return.
queryNoThe natural-language search query.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
source_typeNo
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations mark this as non-read-only and destructive, and the description supplies the reason an agent needs to be careful: it consumes Gumloop credits and demands explicit confirmation. That is genuine context beyond the annotation flags. It still does not describe result volume, ranking behavior, or pagination, so it stops short of full disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and result format, then layers scoping, prerequisites, and cost in descending order of importance. Four sentences is slightly more than needed, and the inline documentation link is of marginal value to an agent, but nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with a nested payload object and no output schema, the description covers purpose, access scoping, plan gating, credit cost, and the confirmation gate. It leaves the payload/payload_file alternates and account semantics to the schema, which is acceptable given the high coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 86%, so the schema already documents query, limit, source_type, account, confirm, and payload. The description adds only that the query is hybrid semantic+keyword in nature, which mildly enriches the query parameter but does not explain the confirm or payload flags. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (search), the mechanism (hybrid semantic + keyword), the resource (knowledge sources in the Company Brain), and the return shape (ranked snippets with citations). It is clearly distinguishable from siblings like list_brain_sources or get_brain_source, which browse rather than query.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real conditions: results are scoped to what the authenticated user can see, the Brain feature must exist (Pro/Enterprise only), each call consumes credits, and explicit confirmation is required. It does not name an alternative tool for browsing indexed sources, so routing is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_messageSend messageA
Destructive

Append a user message to an existing session and resume the agent. The session must be idle, completed, failed, or approval_required; sending to a session that is processing or queued returns 409 interaction_not_in_terminal_state. To hand the agent a message while it is still busy, use the message queue instead.

Files uploaded via Upload session file can be attached to the message with attachments.

Sessions waiting on an approval

A session that stopped to ask you something is approval_required, and you have two ways to move it forward:

  • Answer the ask. Send the pending asks' responses to Resolve approvals. Use this to approve or reject a tool call, or to answer an Ask Question the agent raised. This endpoint rejects approval_responses with a 400.

  • Send a follow-up instead. Post a normal message here. It is appended to the session transcript and starts a new turn, leaving the pending ask unanswered. Use this when the answer no longer matters — for example to redirect the agent or drop the request it was asking about.

See Human in the Loop for how agents pause for approvals and questions.

Streaming the response

api.gumloop.com only serves the non-streaming response above. To stream agent output as it's produced, send the same request body (with stream: true) to the streaming host instead:

POST https://ws.gumloop.com/api/v1/sessions/{session_id}/messages

The response is text/event-stream (Server-Sent Events). With the Python SDK, client.sessions.stream_message(session_id, input="...") routes to ws.gumloop.com automatically and yields parsed StreamEvent objects.

If you send stream: true to api.gumloop.com by mistake, the response is a 400 whose body contains the correct streaming host so you can retry against it. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputNoThe next user message. Required. Also accepted as `message` for backwards compatibility.
streamNoMust be `false` (or omitted) when calling `api.gumloop.com`. Set to `true` only when calling `ws.gumloop.com` (see the streaming section above).
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
session_idYesID of the session to continue.
attachmentsNo
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, openWorldHint=true and idempotentHint=false, and the description reinforces and explains them: runs "can spend credits or trigger downstream actions," explicit confirmation is required for this exact account operation, and unknown outcomes must never be resubmitted automatically. It goes further with host-routing behavior (stream:true against api.gumloop.com yields a 400 containing the correct streaming host) that no annotation conveys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded — purpose, preconditions, and the sibling alternative come first, with approval and streaming details in clearly headed sections. However, the streaming subsection (curl example, SDK call, host-mismatch error) is long for a tool description and pushes the confirmation warning to the very end where it is easy to miss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, nested-object, non-idempotent mutation tool with no output schema, the definition covers the state machine, error codes, attachment provenance, streaming host split, and approval interactions. The only minor gap is that the non-streaming response shape is only referred to as "the response above" rather than described.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 88%, so the baseline is 3, but the description adds real meaning: attachments must be stored paths returned by Upload session file for this session, and `stream` must be false on api.gumloop.com and true only on ws.gumloop.com. It also implicitly explains `confirm` via the confirmation requirement, though it never enumerates `account`, `payload`, or `payload_file` semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource+effect: "Append a user message to an existing session and resume the agent." It explicitly distinguishes itself from sibling operations by naming the message queue for busy sessions and resolve_session_approvals for pending asks, so an agent can route correctly without opening other schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States precise preconditions (session must be idle/completed/failed/approval_required) and the failure mode for the wrong state (409 interaction_not_in_terminal_state). It names the alternative endpoint (message queue) for the busy case and lays out the two competing paths for approval_required sessions, including when a follow-up message is preferable to answering the ask.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_queued_messageSend queued message nowA
Destructive

Send a queued message immediately instead of waiting for the agent to finish its current turn. Any in-progress run is aborted, the queued message is appended to the transcript, and the agent starts processing it.

The response is the same envelope as Send message. Queued messages cannot be sent this way on incognito sessions, and a message that is currently being edited must have its edit finished or cancelled first. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
session_idYesID of the session the queued message belongs to.
queued_message_idYesID of the queued message to send.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (destructive, non-idempotent, open-world), the description discloses the concrete side effects: any in-progress run is aborted, the message is appended to the transcript, and the agent starts processing. It also adds the confirmation requirement, credit-spend warning, and the 'never resubmit unknown outcomes' caution, which are not captured by the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then side effects, then caveats and the confirmation policy. Three short paragraphs, each earning its place; no filler, though the confirmation sentence is slightly redundant with the confirm param.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent mutation with no output schema, the description covers side effects, preconditions, and confirmation. It notes the response matches the 'Send message' envelope, addressing the return shape, though it does not detail error/abort outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters, including confirm and account. The description reinforces the confirm requirement ('explicit confirmation is required for this exact account operation') but adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (send), resource (queued message), and the distinguishing timing/scope: 'immediately instead of waiting for the agent to finish its current turn.' This clearly separates it from siblings like queue_session_message, update_queued_message, and send_message.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It establishes when to use it (to send now rather than wait for turn completion) and a clear when-not condition ('cannot be sent this way on incognito sessions'), plus the editing precondition. It stops short of explicitly naming alternative tools for the incognito/edit cases, leaving the agent to infer the fallback.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_organization_evaluation_targetsSet evaluation targetsA
Destructive

Replaces the full set of targets — who the evaluation grades. Targets expand to agents live: organization covers every agent in the organization, team every agent a team owns, user a member's personal agents, agent one agent. Removing the last target pauses an enabled evaluation; enabled in the response reflects that. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
targetsNo
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
evaluation_idYesID of the organization evaluation.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing that removal of the last target pauses an enabled evaluation, that the response's `enabled` field reflects it, and that runs can spend credits or trigger downstream actions. These are non-obvious side effects an agent must know before calling a destructive, non-idempotent mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and follows with the target-expansion table and the confirmation warning; every sentence carries weight. Slightly dense in the target enumeration, but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation tool with no output schema and nested objects, the description covers replacement semantics, side effects, confirmation requirements, and a hint about the `enabled` return field. It does not explain `account`, `payload`, or `payload_file` behavior, leaving minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the schema already documents most parameters (baseline 3). The description adds real value by spelling out what each target `type` expands to (`organization`/`team`/`user`/`agent`) and how removal affects evaluation state, enriching the enum semantics beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Replaces the full set of targets') and clarifies the target-expansion semantics, so an agent knows exactly what is being overwritten. It implicitly distinguishes itself from sibling update tools by emphasizing full replacement, but never names a sibling, which keeps it at a 4.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating context: explicit confirmation is required for this exact operation and unknown outcomes must not be auto-resubmitted. It does not, however, contrast itself against alternatives like update_organization_evaluation or update_evaluation_config, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_role_credit_limitSet custom role credit limitA
Destructive

This endpoint sets or clears the monthly credit limit of a custom role. The limit applies to each member of the role individually and takes effect immediately: member allowances are recalculated while preserving credits already used in the current billing cycle. Send "monthly_credit_limit": null to clear the role-level limit so members revert to the organization default. When a user belongs to multiple roles, the highest limit across their roles wins. Requests that do not change the stored value are no-ops. Changes are recorded in the organization audit trail. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
role_idYesThe ID of the custom role (the same ID used as `group_id` by the Manage custom role users endpoint).
user_idNoYour user id -- you must be an organization admin to manage custom role credit limits.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
organization_idYesThe ID of the organization the custom role belongs to.
monthly_credit_limitNoThe monthly credit limit applied to each member of this role, or null to clear the role-level limit.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only flag it as a destructive, non-idempotent, open-world write. The description adds substantial context beyond that: immediate effect, per-member application, recalculation preserving already-used credits, highest-limit-wins across multiple roles, no-op behavior, and audit-trail recording. This is exactly the extra behavioral detail the annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose and null semantics are front-loaded and each operational sentence earns its place. However, the closing boilerplate about 'runs can spend credits or trigger downstream actions' is generic and tangential to a credit-limit setter, slightly padding the description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the definition is complete: it covers effect timing, recalculation semantics, role precedence, no-op behavior, audit recording, and confirmation. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning the schema lacks: the per-member scope of the limit and the multi-role precedence rule ('highest limit across their roles wins'). The null-clearing semantics are already in the schema, so this is additive rather than duplicative.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'sets or clears the monthly credit limit of a custom role'. This clearly distinguishes it from the read siblings get_role_credit_limit and list_role_credit_limits without needing to see their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the key usage condition ('send monthly_credit_limit: null to clear the role-level limit') and the confirmation requirement, giving clear context for when to invoke. It does not explicitly name the read alternatives (get_role_credit_limit/list_role_credit_limits), so routing is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_flowStart flow runA
Destructive

This endpoint is used to trigger a flow run via API Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
user_idNoThe id for the user initiating the flow.
project_idNo(Optional) The id of the project within which the flow is executed.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
saved_item_idNoThe id for the saved flow.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds value beyond that by disclosing credit spend, downstream side effects, the confirmation requirement, and non-idempotent retry behavior. It does not describe the return payload, but no output schema exists and annotations carry the rest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Short and front-loaded, with the operation stated first and the confirmation warning second. 'This endpoint is used to' is mild filler, but nothing is padded or redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent, open-world mutation with a fully documented schema, the description covers the key risks (credits, downstream actions, confirmation, retries). It omits how to track the resulting run (e.g., get_run_details), which a fully complete definition would mention.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with 7 well-documented parameters (account, confirm, payload, payload_file, saved_item_id, user_id, project_id). The description adds nothing about parameter meaning, so the schema does the heavy lifting; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('trigger a flow run via API'), which is clear enough to act on. It does not differentiate from siblings like kill_flow, get_run_details, or list_flows, but the core action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real usage conditions: explicit confirmation is required, runs can spend credits, and unknown outcomes must not be auto-resubmitted. It names no alternatives (e.g., kill_flow to stop, get_input_schema to inspect before running), so routing guidance is incomplete but the when/when-not is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_agentUpdate agentA
Destructive

Update an existing agent. Only fields included in the request body are changed; omitted fields are left untouched.

This endpoint edits document fields only. To attach or detach skills, use PATCH /agents/{agent_id}/skills; to manage MCP servers, use the agent MCP server endpoints.

is_active: false is not a pause switch. It retires the agent: the agent disappears from GET /agents, and both GET and PATCH /agents/{agent_id} return 404 afterwards, so you cannot set it back to true through the API. To stop an agent from running on its own while keeping it fully reachable, disable its triggers instead. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
toolsNoWhen provided, replaces the agent's tool list.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
team_idNoWhen provided, transfers ownership of the agent to this team.
agent_idYesID of the agent to update. Also accepts the reserved aliases `gumball` and `analytics`.
metadataNo
is_activeNo
resourcesNoWhen provided, replaces the agent's resource list.
model_nameNoID of the LLM the agent runs on. Use `GET /models` to discover valid values.
descriptionNo
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
system_promptNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructiveHint/openWorldHint/idempotentHint) by disclosing the non-obvious irreversible semantic of is_active:false — the agent disappears from GET /agents and returns 404 with no API path back. It adds the credential/confirmation requirement and warns against auto-resubmitting unknown outcomes, which is exactly the behavioral context an agent needs before a destructive write.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then scope, then the irreversible-warning, then the confirmation rule — a sensible priority order. Slightly long, and the 'runs can spend credits' line is adjacent to but not strictly about this endpoint, keeping it just short of 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter, nested-payload mutation with no output schema, the description covers the dangerous parameter, sibling routing, and confirmation posture well. It omits mention of the payload/payload_file/body-flag exclusivity model, but that is fully documented in the schema, so what remains is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 64%, and the description adds real meaning for the highest-risk parameter (is_active retirement semantics) and for confirm ('explicit confirmation is required for this exact account operation'). It does not add semantics for payload vs body-flags vs payload_file mutual exclusivity in the prose, but that is covered by the schema, so this exceeds the coverage baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Update an existing agent') and immediately scopes it ('only fields included in the request body are changed'). It explicitly carves out what this endpoint is NOT for, naming the skill and MCP-server siblings that handle those cases, so an agent can distinguish it from update_agent_skills/attach_agent_mcp_server without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing to alternatives ('to attach or detach skills, use PATCH /agents/{agent_id}/skills'; 'to manage MCP servers, use...') and a clear when-not for is_active ('disable its triggers instead'). It also states the confirmation precondition, so the conditions selecting this tool are fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_agent_skillsAttach or detach agent skillsA
Destructive

Attach and/or detach skills on an agent using deltas. This is not a replace-list: skills you don't mention are left untouched.

  • The operation is idempotent. Re-attaching a skill that's already attached (or detaching one that isn't) is reported under already_attached / already_detached rather than failing.

  • A skill ID may not appear in both attach and detach.

  • Up to 100 unique skill IDs total (attach + detach) per request.

  • Attaching requires INVOKE permission on the skill. Detaching is permissive so stale attachments can always be removed. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
attachNoSkill IDs to attach. Ignored if already attached.
detachNoSkill IDs to detach. Ignored if not attached.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
agent_idYesID of the agent whose skills to update. Also accepts the reserved aliases `gumball` and `analytics`.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A3.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description asserts 'The operation is idempotent,' but the annotations declare idempotentHint: false. This is a direct conflict on a named behavioral trait, so per the rubric the description contradicts the annotations. The otherwise-rich content (already_attached/already_detached reporting, permission asymmetry, credit/downstream warning) cannot offset a claim that contradicts the declared hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core delta semantics in the first sentence, then uses tight bullets for the edge cases and constraints. The closing confirmation/credit paragraph partially duplicates the `confirm` parameter description, which is the only real redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries return-value burden and does so by naming the already_attached/already_detached reporting fields, plus permission, limit, and confirmation requirements. It does not cover failure modes (e.g. invalid skill IDs, missing INVOKE permission) or the account/payload alternates, so it is strong but not exhaustive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine semantics beyond the schema: that attach/detach are deltas rather than a replacement list, that an ID may not appear in both, and that the combined total is capped at 100. It says little about account/confirm/payload/payload_file, keeping it below 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (attach/detach) and resource (skills on an agent), and explicitly disambiguates from a replace-list semantics. An agent can distinguish this from sibling tools like attach_agent_mcp_server or update_agent without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating context: delta semantics, mutual-exclusion rule, the 100-ID cap, permission requirements for attach vs detach, and the explicit-confirmation prerequisite. It stops short of naming an alternative tool (e.g. why to use this rather than update_agent or attach_agent_mcp_server), so it is not a full when/when-not/alternatives treatment.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_evaluation_configUpdate evaluation configA
Destructive

Partially update the evaluation configuration for an agent. Omitted fields keep their current value. Provided list fields (criteria, tags, data_points) replace that list entirely.

Requires Pro tier or above. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNoTag vocabulary (replaces existing list). Max 50.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
enabledNoWhether evaluations are enabled for this agent.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
agent_idYesID of the agent. Also accepts the reserved aliases `gumball` and `analytics`.
criteriaNo
sentimentNo
model_nameNoLLM model to use for evaluation.
data_pointsNo
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
include_auto_tagsNoAllow the evaluator to suggest tags beyond your predefined vocabulary.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructive=true, idempotent=false and openWorld=true, and the description adds genuinely non-redundant behavior: list fields are wholesale replaced rather than merged, a tier prerequisite gates the call, confirmation must be tied to the exact requested action, and unknown outcomes must not be resubmitted because runs can spend credits. That is exactly the extra context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the partial-update rule and the list-replacement rule before prerequisites. The closing sentence about credits and downstream actions is somewhat generic to the evaluation domain rather than specific to updating config, so it earns slightly less than a perfect score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter, nested, mutation tool with no output schema, the description covers purpose, partial semantics, list replacement, tier gating and confirmation. It leaves the payload/payload_file/body-flag mutual exclusivity and the agent_id aliases to the schema, which is acceptable given 75% coverage, but an agent assembling a call still relies on schema reading.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 12 parameters and 75% schema description coverage, the schema carries most parameter detail, but the description adds the crucial replace-vs-merge semantics for criteria, tags and data_points that the schema does not state. It does not touch payload/payload_file vs body-flag mutual exclusivity, leaving a gap, but the added list semantics are substantive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Partially update the evaluation configuration for an agent') and immediately qualifies the scope with 'partial', which cleanly separates it from the read-side siblings get_evaluation_config / retrieve_evaluation. The verb-plus-resource pairing leaves no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real operating conditions: omitted fields keep current values, list fields replace entirely, Pro tier is required, and explicit confirmation is required for this exact account operation. It does not name alternative tools (e.g. get_evaluation_config for reading, run_evaluations for executing), so it stops short of explicit when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_organization_evaluationUpdate evaluationA
Destructive

Partial update. Only the fields you send change. config is merged field by field; a list you send (criteria, tags, data_points) replaces that list wholesale. description: null clears the description.

Setting enabled: true requires at least one criterion, tag, or data point (400 organization_evaluation_empty_rubric) and at least one covered agent (400 organization_evaluation_no_targets). Emptying the rubric of an enabled evaluation pauses it. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
evaluation_idYesID of the organization evaluation.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and non-idempotent, and the description adds real value beyond that: the exact 400 error codes on invalid enable, the pause-on-empty-rubric behavior, the credit-spending warning, and the confirmation requirement. This is genuinely useful behavioral context not in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the partial-update semantics in the first two sentences, then operational constraints, then the confirmation warning. Dense but every sentence carries information. Slightly heavy on back-to-back clauses.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers mutation semantics, validation errors, side effects, and confirmation requirement for a destructive non-idempotent tool. No output schema, but a mutation returning the updated resource is expected, so return value isn't needed. Missing only a pointer to sibling alternatives.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema thoroughly documents each field (id semantics, tag name casing, frequency enum, etc.). The description adds merge semantics for config and null-clears for description, which is real added value but overlap is high; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (partial update) and resource, and immediately distinguishes semantics from sibling update tools (update_evaluation_config, set_organization_evaluation_targets) by clarifying that config merges field-by-field while lists replace wholesale and null clears description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear behavioral guidance: what happens when enabling, what happens when emptying the rubric, the confirmation requirement. Doesn't explicitly name sibling alternatives like update_evaluation_config, but the operational context is clear enough to select correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_queued_messageUpdate queued messageA
Destructive

Replace the content of a message that is still waiting in the session's queue. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputNoThe new message content. Cannot be empty. Also accepted as `message` for backwards compatibility.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
session_idYesID of the session the queued message belongs to.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
queued_message_idYesID of the queued message to update.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true/non-idempotent, but the description adds genuinely new context: an explicit confirmation requirement and the warning that runs can consume credits and trigger downstream actions. It stops short of describing what happens to the original content or any rate/error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, purpose front-loaded, no filler. The credit/downstream-action sentence is slightly generic boilerplate relative to a queue-content update, which keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description plus a fully described 7-parameter schema cover the call adequately, including the confirmation gate for a destructive, non-idempotent write. Only the return/confirmation behavior on success is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 7 parameters, including the payload/body-flag mutual exclusion and the account selector. The description adds no syntax, format, or constraint detail beyond that baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Replace the content of a message') plus a discriminating precondition ('still waiting in the session's queue'), which separates it from send_queued_message, delete_queued_message, and send_message without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating rules: explicit confirmation is required, and unknown outcomes must never be resubmitted automatically. It does not, however, name the sibling alternative to use when the message has already been sent or should be discarded instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_skillUpdate skillA
Destructive

Replace a skill's files with a new upload. Reparses SKILL.md to update the skill's name, description, and metadata, and creates a new version. Maximum upload size is 10 MB. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
filesNo
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
skill_idYesID of the skill to update.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true; the description adds real value on top by disclosing the reparse/versioning behavior, the confirmation requirement, and the credit/downstream-action warning. The only gap is that it does not describe what happens to the prior version or how failures surface.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and effect, followed by size and safety constraints. The credit/downstream-action sentence is somewhat boilerplate but carries genuine operational meaning for a destructive, non-idempotent call.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with no output schema and 83% schema coverage, the description covers effect, versioning, confirmation, and retry safety. The size-limit inconsistency with the schema and the absence of any failure-mode description are the remaining gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 83%, so the schema already documents skill_id, account, confirm, payload, and payload_file. The description's only parameter-relevant claim — 'Maximum upload size is 10 MB' — actually conflicts with the schema, which caps each file and the total at 5 MiB, so it adds confusion rather than clarity. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Replace a skill's files with a new upload') plus the concrete side effects: reparse of SKILL.md, update of name/description/metadata, and new version creation. It does not distinguish itself from siblings like create_skill, delete_skill, or update_agent_skills, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides some operational guidance — explicit confirmation is required, and unknown outcomes must not be auto-resubmitted — which implies when it is safe to call. However, it never names an alternative (create_skill, update_agent_skills, delete_skill) or states the condition that selects this tool over them, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_brain_filesUpload filesA
Destructive

Upload up to 25 files as multipart/form-data parts named files. Accepted types are PDF, Word, PowerPoint, Excel, and text formats (.txt, .md, .html, .csv, .rtf), each up to 25 MB.

Indexing starts on its own after the upload: an active source indexes and bills immediately, a draft source runs a credit estimate instead (see Retrieve estimate). Poll List files until each file's status is indexed.

Files the upload policy refuses (unsupported type, too large, empty) are returned in rejected with a 201; the request is a 400 no_files_accepted only when every file was refused. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
filesNo
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
source_idYesThe source id.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing billing behavior (active bill immediately, draft estimates credits), the partial-rejection semantics (rejected array with 201 vs 400 no_files_accepted), and the confirmation requirement for a destructive/open-world write. This is real value-add over destructiveHint/openWorldHint alone. Minor deduction because the stated 25 MB per-file limit conflicts with the schema's 5 MiB, undermining trust in the behavioral claims.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and limits, then explains post-upload behavior and error semantics in scannable sentences. Slightly long, and the space spent on a size limit that contradicts the schema is wasted, but every other sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter, nested-object mutation tool with no output schema, the description covers outcome behavior (rejection arrays, status codes, polling) that the schema cannot convey. It leaves gaps around the payload/payload_file alternative body mechanisms and the account parameter, but the agent has enough to call the primary path correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the baseline is 3, and the description does add genuine param context: the multipart part name `files` and the accepted file extensions. However, it silently omits source_id, account, payload, and payload_file, and its per-file size figure (25 MB) directly contradicts the schema's 'at most 5 MiB', which is worse than saying nothing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Upload up to 25 files as multipart/form-data parts named `files`', plus the accepted formats and the source-indexing consequence. An agent can tell it is the brain-source upload tool, but the description never distinguishes it from the sibling upload_files / upload_file / upload_session_file tools, so the sibling differentiation required for a 5 is absent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives useful context on what happens after the call (draft vs active indexing, poll List files until status is `indexed`) and that explicit confirmation is required. It stops short of stating when to choose this tool over upload_files or upload_session_file, so usage is implied rather than directed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_fileUpload fileB
Destructive

Upload file Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
user_idNoThe user ID associated with the file. Required if project_id is not provided.
file_nameNoThe name of the file to be uploaded.
file_pathNoRegular local file path, no symlink, at most 3 MiB. Encoded as native base64 file_content; cannot mix with file_content or payload routes.
project_idNoThe project ID associated with the file. Required if user_id is not provided.
file_contentNoBase64 encoded content of the file.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive=true, idempotent=false, openWorld=true, readOnly=false, and the description goes beyond them by stating that confirmation is mandatory, that runs can spend credits or trigger downstream effects, and that unknown outcomes must not be resubmitted. These caveats meaningfully inform retry and safety behavior. It still omits the multi-route input handling and size limits, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The body is two tight sentences with no filler, and the critical constraint is front-loaded. The redundant 'Upload file' first line adds little, but the rest is efficient and well ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with nested objects and three alternative input routes and no output schema, the description covers the confirmation/credits safety context adequately but is silent on the input-route complexity that dominates correct invocation. It is minimally complete rather than fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter, including the mutually exclusive file_path/file_content/payload routes, is documented in the schema itself. The description adds nothing about parameters, which is acceptable given the schema does the work, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Upload file', which simply restates the tool name and title verbatim, so it does not independently convey purpose. It never distinguishes this single-file upload from close siblings like upload_files, upload_session_file, or upload_brain_files. The only added content is operational caution, not a clearer statement of what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives real usage guidance: explicit confirmation is required, the operation can spend credits or trigger downstream actions, and unknown outcomes must not be auto-resubmitted. However, it offers no when-to-use versus alternatives, and with upload_files and several scoped upload tools present, the agent gets no help choosing among them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_filesUpload multiple filesA
Destructive

Upload multiple files Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
filesNo
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
user_idNoThe user ID associated with the files. Required if project_id is not provided.
project_idNoThe project ID associated with the files. Required if user_id is not provided.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, idempotentHint=false, openWorldHint=true, readOnlyHint=false. The description adds value beyond these by naming the concrete stakes ('runs can spend credits or trigger downstream actions') and the non-idempotency consequence ('never resubmit unknown outcomes automatically'), which explains WHY it is non-idempotent rather than just declaring it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The title line is repeated before the substantive guidance, which is slight redundancy, but the body is three front-loaded sentences with no filler. The key safety constraints (confirmation, no auto-resubmit) lead the content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-param, nested, destructive, non-idempotent mutation tool with no output schema, the description covers the two highest-risk behaviors (confirmation gating and non-resubmission). It does not explain return values or the payload vs. body-flags selection, but annotations plus rich schema cover most remaining ground.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 86%, so the schema already documents files, account, confirm, user_id/project_id, and payload. The description adds no parameter-specific meaning (e.g., the user_id vs project_id mutual requirement, or the payload/payload_file mutual exclusivity) beyond what the schema says. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Upload multiple files'), which clearly distinguishes it from the singular sibling upload_file and from upload_session_file/upload_brain_files. However, it doesn't clarify the target scope (e.g., account-level files vs. session/project files) beyond what the schema reveals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a warning about resubmission ('never resubmit unknown outcomes automatically') and confirmation requirement, which implies caution-oriented usage. But it never names an alternative sibling (upload_file, upload_session_file) or states when this batch tool is preferable to the single-file variant.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_session_fileUpload session fileA
Destructive

Upload a file into a session's input namespace so it can be attached to a message. The response returns the stored path — pass it as file_name in the attachments array when sending a message on the same session.

Files are base64 encoded in the request body and limited to 200MB (decoded). Uploaded files are scoped to the session they were uploaded to and cannot be attached to messages on other sessions. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
file_nameNoName of the file. Directory components are stripped; the base name is sanitized before storage.
media_typeNoMIME type of the file. Echoed back in the response.
session_idYesID of the session to upload the file to.
file_contentNoBase64-encoded file contents. Maximum decoded size is 200MB.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, so the safety profile is covered. The description adds real context beyond that: 200MB decoded limit, base64 encoding, session scoping of stored files, mandatory confirmation, and a caution against automatic resubmission. The credits sentence is somewhat generic boilerplate, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short paragraphs, front-loaded with the action and the return-value contract. The trailing confirmation/credits sentences are somewhat generic and repeat a policy that the `confirm` parameter description already implies.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by stating what the response returns and how to use it. Limits, scoping, and confirmation are all covered; only the payload/payload_file/body-flag selection is left to the schema, which documents it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds cross-tool meaning: the returned stored path is consumed as `file_name` in the `attachments` array of send_message. That linkage is not derivable from this tool's own schema and is genuinely useful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (upload) and resource (file into a session's input namespace), and pins the purpose to a concrete downstream use (attaching to a message). An agent can distinguish it from sibling upload_file/upload_files/upload_brain_files because the session scoping and message-attachment framing are explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains when to use it (to get a path that can be attached to a message on the same session) and states the confirmation prerequisite. It does not explicitly name or exclude the sibling upload tools (upload_file, upload_files), so the routing guidance is clear but not complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 43 tool updatesv3.0.0
    • Changedapprove_brain_source1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedattach_agent_mcp_server1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedcall_mcp_tools14 fields changed
      • addedInput schema / $defs / calls
        Added value: +{
        +  "description": "Tool calls to execute. Dispatched concurrently; the batch is capped at 5.",
        +  "items": {
        +    "properties": {
        +      "arguments": {
        +        "description": "Arguments passed to the tool. Defaults to `{}`. Validated by the tool's `input_schema`.",
        +        "type": "object"
        +      },
        +      "ref": {
        +        "description": "Caller-supplied identifier echoed back on the matching result. When omitted, Gumloop assigns the call's zero-based index in `calls` as its `ref`.",
        +        "type": [
        +          "string",
        +          "null"
        +        ]
        +      },
        +      "server_id": {
        +        "type": "string"
        +      },
        +      "tool_name": {
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "server_id",
        +      "tool_name"
        +    ],
        +    "type": "object"
        +  },
        +  "maxItems": 5,
        +  "minItems": 1,
        +  "type": "array"
        +}
      • addedInput schema / properties / calls / $ref
        Added value: +"#/$defs/calls"
      • removedInput schema / properties / calls / description
        Removed value: -"Tool calls to execute. Dispatched concurrently; the batch is capped at 5."
      • removedInput schema / properties / calls / items
        Removed value: -{
        -  "properties": {
        -    "arguments": {
        -      "description": "Arguments passed to the tool. Defaults to `{}`. Validated by the tool's `input_schema`.",
        -      "type": "object"
        -    },
        -    "ref": {
        -      "description": "Caller-supplied identifier echoed back on the matching result. When omitted, Gumloop assigns the call's zero-based index in `calls` as its `ref`.",
        -      "type": [
        -        "string",
        -        "null"
        -      ]
        -    },
        -    "server_id": {
        -      "type": "string"
        -    },
        -    "tool_name": {
        -      "type": "string"
        -    }
        -  },
        -  "required": [
        -    "server_id",
        -    "tool_name"
        -  ],
        -  "type": "object"
        -}
      • removedInput schema / properties / calls / maxItems
        Removed value: -5
      • removedInput schema / properties / calls / minItems
        Removed value: -1
      • removedInput schema / properties / calls / type
        Removed value: -"array"
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / payload / properties / calls / $ref
        Added value: +"#/$defs/calls"
      • removedInput schema / properties / payload / properties / calls / description
        Removed value: -"Tool calls to execute. Dispatched concurrently; the batch is capped at 5."
      • removedInput schema / properties / payload / properties / calls / items
        Removed value: -{
        -  "properties": {
        -    "arguments": {
        -      "description": "Arguments passed to the tool. Defaults to `{}`. Validated by the tool's `input_schema`.",
        -      "type": "object"
        -    },
        -    "ref": {
        -      "description": "Caller-supplied identifier echoed back on the matching result. When omitted, Gumloop assigns the call's zero-based index in `calls` as its `ref`.",
        -      "type": [
        -        "string",
        -        "null"
        -      ]
        -    },
        -    "server_id": {
        -      "type": "string"
        -    },
        -    "tool_name": {
        -      "type": "string"
        -    }
        -  },
        -  "required": [
        -    "server_id",
        -    "tool_name"
        -  ],
        -  "type": "object"
        -}
      • removedInput schema / properties / payload / properties / calls / maxItems
        Removed value: -5
      • removedInput schema / properties / payload / properties / calls / minItems
        Removed value: -1
      • removedInput schema / properties / payload / properties / calls / type
        Removed value: -"array"
    • Changedcancel_session1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedcreate_agent19 fields changed
      • addedInput schema / $defs / is_active
        Added value: +{
        +  "default": true,
        +  "description": "Whether the agent is active. Defaults to `true`.\n\nSetting this to `false` retires the agent: it disappears from `GET /agents`, and `GET`/`PATCH /agents/{agent_id}` return `404`, so it cannot be reactivated through the API. This is not a pause switch — to stop an agent from running while keeping it reachable, disable its triggers instead.\n",
        +  "type": "boolean"
        +}
      • addedInput schema / $defs / skill_ids
        Added value: +{
        +  "description": "IDs of skills to attach to the agent. Attachment happens inside the create transaction, so an invalid ID fails the whole request (no orphaned agent). Omit to attach none. The caller must hold `INVOKE` on each skill. After creation, manage skills with `PATCH /agents/{agent_id}/skills`.\n",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": [
        +    "array",
        +    "null"
        +  ]
        +}
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / is_active / $ref
        Added value: +"#/$defs/is_active"
      • removedInput schema / properties / is_active / default
        Removed value: -true
      • removedInput schema / properties / is_active / description
        Removed value: -"Whether the agent is active. Defaults to `true`.\n\nSetting this to `false` retires the agent: it disappears from `GET /agents`, and `GET`/`PATCH /agents/{agent_id}` return `404`, so it cannot be reactivated through the API. This is not a pause switch — to stop an agent from running while keeping it reachable, disable its triggers instead.\n"
      • removedInput schema / properties / is_active / type
        Removed value: -"boolean"
      • addedInput schema / properties / payload / properties / is_active / $ref
        Added value: +"#/$defs/is_active"
      • removedInput schema / properties / payload / properties / is_active / default
        Removed value: -true
      • removedInput schema / properties / payload / properties / is_active / description
        Removed value: -"Whether the agent is active. Defaults to `true`.\n\nSetting this to `false` retires the agent: it disappears from `GET /agents`, and `GET`/`PATCH /agents/{agent_id}` return `404`, so it cannot be reactivated through the API. This is not a pause switch — to stop an agent from running while keeping it reachable, disable its triggers instead.\n"
      • removedInput schema / properties / payload / properties / is_active / type
        Removed value: -"boolean"
      • addedInput schema / properties / payload / properties / skill_ids / $ref
        Added value: +"#/$defs/skill_ids"
      • removedInput schema / properties / payload / properties / skill_ids / description
        Removed value: -"IDs of skills to attach to the agent. Attachment happens inside the create transaction, so an invalid ID fails the whole request (no orphaned agent). Omit to attach none. The caller must hold `INVOKE` on each skill. After creation, manage skills with `PATCH /agents/{agent_id}/skills`.\n"
      • removedInput schema / properties / payload / properties / skill_ids / items
        Removed value: -{
        -  "type": "string"
        -}
      • removedInput schema / properties / payload / properties / skill_ids / type
        Removed value: -[
        -  "array",
        -  "null"
        -]
      • addedInput schema / properties / skill_ids / $ref
        Added value: +"#/$defs/skill_ids"
      • removedInput schema / properties / skill_ids / description
        Removed value: -"IDs of skills to attach to the agent. Attachment happens inside the create transaction, so an invalid ID fails the whole request (no orphaned agent). Omit to attach none. The caller must hold `INVOKE` on each skill. After creation, manage skills with `PATCH /agents/{agent_id}/skills`.\n"
      • removedInput schema / properties / skill_ids / items
        Removed value: -{
        -  "type": "string"
        -}
      • removedInput schema / properties / skill_ids / type
        Removed value: -[
        -  "array",
        -  "null"
        -]
    • Changedcreate_brain_source1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedcreate_chat_completion31 fields changed
      • addedInput schema / $defs / image_config
        Added value: +{
        +  "description": "Image-generation parameters (size, quality, aspect_ratio, background, output_format, partial_images). Optional. Image-generation models accept either `modalities: [\"image\"]` or `image_config` (or both); chat models ignore this field.",
        +  "type": "object"
        +}
      • addedInput schema / $defs / messages
        Added value: +{
        +  "description": "Conversation history. Roles `system`, `developer`, `user`, `assistant`, and `tool`. User messages accept multipart `content` with `text` and `image_url` parts. Assistant messages can carry `tool_calls`; answer each one with a `tool` message whose `tool_call_id` matches the call's `id`.",
        +  "items": {
        +    "type": "object"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / $defs / response_format
        Added value: +{
        +  "description": "Constrain the response. `{type: \"json_object\"}` returns a JSON object; `{type: \"json_schema\", json_schema: {name, strict, schema}}` returns JSON matching the supplied schema.\n",
        +  "type": "object"
        +}
      • addedInput schema / $defs / tool_choice
        Added value: +{
        +  "description": "`\"auto\"` lets the model choose and is the default when `tools` are sent. `\"none\"` disables tool calls, `\"required\"` forces a tool call, and `{\"type\": \"function\", \"function\": {\"name\": \"...\"}}` forces a specific tool.",
        +  "oneOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "object"
        +    }
        +  ]
        +}
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / image_config / $ref
        Added value: +"#/$defs/image_config"
      • removedInput schema / properties / image_config / description
        Removed value: -"Image-generation parameters (size, quality, aspect_ratio, background, output_format, partial_images). Optional. Image-generation models accept either `modalities: [\"image\"]` or `image_config` (or both); chat models ignore this field."
      • removedInput schema / properties / image_config / type
        Removed value: -"object"
      • addedInput schema / properties / messages / $ref
        Added value: +"#/$defs/messages"
      • removedInput schema / properties / messages / description
        Removed value: -"Conversation history. Roles `system`, `developer`, `user`, `assistant`, and `tool`. User messages accept multipart `content` with `text` and `image_url` parts. Assistant messages can carry `tool_calls`; answer each one with a `tool` message whose `tool_call_id` matches the call's `id`."
      • removedInput schema / properties / messages / items
        Removed value: -{
        -  "type": "object"
        -}
      • removedInput schema / properties / messages / type
        Removed value: -"array"
      • addedInput schema / properties / payload / properties / image_config / $ref
        Added value: +"#/$defs/image_config"
      • removedInput schema / properties / payload / properties / image_config / description
        Removed value: -"Image-generation parameters (size, quality, aspect_ratio, background, output_format, partial_images). Optional. Image-generation models accept either `modalities: [\"image\"]` or `image_config` (or both); chat models ignore this field."
      • removedInput schema / properties / payload / properties / image_config / type
        Removed value: -"object"
      • addedInput schema / properties / payload / properties / messages / $ref
        Added value: +"#/$defs/messages"
      • removedInput schema / properties / payload / properties / messages / description
        Removed value: -"Conversation history. Roles `system`, `developer`, `user`, `assistant`, and `tool`. User messages accept multipart `content` with `text` and `image_url` parts. Assistant messages can carry `tool_calls`; answer each one with a `tool` message whose `tool_call_id` matches the call's `id`."
      • removedInput schema / properties / payload / properties / messages / items
        Removed value: -{
        -  "type": "object"
        -}
      • removedInput schema / properties / payload / properties / messages / type
        Removed value: -"array"
      • addedInput schema / properties / payload / properties / response_format / $ref
        Added value: +"#/$defs/response_format"
      • removedInput schema / properties / payload / properties / response_format / description
        Removed value: -"Constrain the response. `{type: \"json_object\"}` returns a JSON object; `{type: \"json_schema\", json_schema: {name, strict, schema}}` returns JSON matching the supplied schema.\n"
      • removedInput schema / properties / payload / properties / response_format / type
        Removed value: -"object"
      • addedInput schema / properties / payload / properties / tool_choice / $ref
        Added value: +"#/$defs/tool_choice"
      • removedInput schema / properties / payload / properties / tool_choice / description
        Removed value: -"`\"auto\"` lets the model choose and is the default when `tools` are sent. `\"none\"` disables tool calls, `\"required\"` forces a tool call, and `{\"type\": \"function\", \"function\": {\"name\": \"...\"}}` forces a specific tool."
      • removedInput schema / properties / payload / properties / tool_choice / oneOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "object"
        -  }
        -]
      • addedInput schema / properties / response_format / $ref
        Added value: +"#/$defs/response_format"
      • removedInput schema / properties / response_format / description
        Removed value: -"Constrain the response. `{type: \"json_object\"}` returns a JSON object; `{type: \"json_schema\", json_schema: {name, strict, schema}}` returns JSON matching the supplied schema.\n"
      • removedInput schema / properties / response_format / type
        Removed value: -"object"
      • addedInput schema / properties / tool_choice / $ref
        Added value: +"#/$defs/tool_choice"
      • removedInput schema / properties / tool_choice / description
        Removed value: -"`\"auto\"` lets the model choose and is the default when `tools` are sent. `\"none\"` disables tool calls, `\"required\"` forces a tool call, and `{\"type\": \"function\", \"function\": {\"name\": \"...\"}}` forces a specific tool."
      • removedInput schema / properties / tool_choice / oneOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "object"
        -  }
        -]
    • Changedcreate_organization_evaluation1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedcreate_session10 fields changed
      • addedInput schema / $defs / name
        Added value: +{
        +  "description": "Optional display name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so an empty or whitespace-only value is rejected. A name you supply is kept; the automatic title only fills in sessions created without one. Rename the session later with `PATCH /sessions/{session_id}`.",
        +  "maxLength": 256,
        +  "type": "string"
        +}
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / name / $ref
        Added value: +"#/$defs/name"
      • removedInput schema / properties / name / description
        Removed value: -"Optional display name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so an empty or whitespace-only value is rejected. A name you supply is kept; the automatic title only fills in sessions created without one. Rename the session later with `PATCH /sessions/{session_id}`."
      • removedInput schema / properties / name / maxLength
        Removed value: -256
      • removedInput schema / properties / name / type
        Removed value: -"string"
      • addedInput schema / properties / payload / properties / name / $ref
        Added value: +"#/$defs/name"
      • removedInput schema / properties / payload / properties / name / description
        Removed value: -"Optional display name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so an empty or whitespace-only value is rejected. A name you supply is kept; the automatic title only fills in sessions created without one. Rename the session later with `PATCH /sessions/{session_id}`."
      • removedInput schema / properties / payload / properties / name / maxLength
        Removed value: -256
      • removedInput schema / properties / payload / properties / name / type
        Removed value: -"string"
    • Changedcreate_skill14 fields changed
      • addedInput schema / $defs / files
        Added value: +{
        +  "description": "Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks.",
        +  "items": {
        +    "minLength": 1,
        +    "type": "string"
        +  },
        +  "maxItems": 25,
        +  "minItems": 1,
        +  "type": "array"
        +}
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / files / $ref
        Added value: +"#/$defs/files"
      • removedInput schema / properties / files / description
        Removed value: -"Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks."
      • removedInput schema / properties / files / items
        Removed value: -{
        -  "minLength": 1,
        -  "type": "string"
        -}
      • removedInput schema / properties / files / maxItems
        Removed value: -25
      • removedInput schema / properties / files / minItems
        Removed value: -1
      • removedInput schema / properties / files / type
        Removed value: -"array"
      • addedInput schema / properties / payload / properties / files / $ref
        Added value: +"#/$defs/files"
      • removedInput schema / properties / payload / properties / files / description
        Removed value: -"Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks."
      • removedInput schema / properties / payload / properties / files / items
        Removed value: -{
        -  "minLength": 1,
        -  "type": "string"
        -}
      • removedInput schema / properties / payload / properties / files / maxItems
        Removed value: -25
      • removedInput schema / properties / payload / properties / files / minItems
        Removed value: -1
      • removedInput schema / properties / payload / properties / files / type
        Removed value: -"array"
    • Changeddelete_brain_file1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changeddelete_brain_source1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changeddelete_organization_evaluation1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changeddelete_queued_message1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changeddelete_skill1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changeddetach_agent_mcp_server1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedexport_data75 fields changed
      • addedInput schema / $defs / category_filter
        Added value: +{
        +  "description": "**Only applicable when `data_type` is `\"credit_logs\"`.** Optional filter to export only credit logs matching a specific category (e.g., `\"PIPELINE_RUN\"`, `\"AGENT_RUN\"`, `\"CUSTOM_NODE_RD\"`, `\"EXTERNAL_GUMCP_CALL\"`). When omitted, all categories are included.\n",
        +  "type": "string"
        +}
      • addedInput schema / $defs / data_type
        Added value: +{
        +  "default": "workflows",
        +  "description": "The type of data to export. Use `\"workflows\"` to export workflow run data, `\"agents\"` to export agent configuration data, `\"agent_interactions\"` to export agent interaction data, `\"credit_logs\"` to export credit transaction history, `\"interaction_evaluations\"` to export completed chat evaluations, or `\"gumstack\"` to export Gumstack tool call activity. Defaults to `\"workflows\"`. Audit-log exports are available in the Gumloop app only, not through this endpoint.\n",
        +  "enum": [
        +    "workflows",
        +    "agents",
        +    "agent_interactions",
        +    "credit_logs",
        +    "interaction_evaluations",
        +    "gumstack"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / $defs / entity_ids
        Added value: +{
        +  "description": "An optional array of specific entity IDs to filter the export. For workflow exports (`data_type: \"workflows\"`), these are workbook IDs. For agent exports (`data_type: \"agents\"`) and agent interaction exports (`data_type: \"agent_interactions\"`), these are agent IDs. When provided, only data for the specified entities will be included. **Not applicable when `data_type` is `\"credit_logs\"`.**\n",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / $defs / export_fields
        Added value: +{
        +  "description": "Array of fields to include in the export. The available fields depend on the `data_type`.\n\n**Workflow fields** (when `data_type` is `\"workflows\"`):\n- `workbook_id` - The workbook identifier\n- `workbook_name` - The workbook name\n- `workbook_created_ts` - Workbook creation timestamp\n- `user_id` - The user identifier\n- `user_email` - The user's email address\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n- `run_id` - The flow run identifier\n- `credit_cost` - Credits consumed by the run\n- `pl_run_created_ts` - Flow run creation timestamp\n- `pl_run_finished_ts` - Flow run completion timestamp\n- `pipeline` - The full pipeline configuration (JSON)\n\n**Agent fields** (when `data_type` is `\"agents\"`):\n- `agent_id` - The agent identifier\n- `agent_name` - The agent name\n- `agent_description` - The agent description\n- `agent_model` - The model used by the agent\n- `agent_system_prompt` - The agent's system prompt\n- `agent_created_ts` - Agent creation timestamp\n- `agent_tools` - Tools configured for the agent (JSON)\n- `agent_metadata` - Additional agent metadata (JSON)\n- `agent_evaluations_enabled` - Whether the agent's own Evaluations are turned on (`true`/`false`)\n- `creator_email` - Email of the user who created the agent\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n- `folder_id` - The identifier of the folder containing the agent (empty when the agent is not in a folder)\n- `folder_name` - The name of the folder containing the agent (empty when the agent is not in a folder)\n\n**Agent interaction fields** (when `data_type` is `\"agent_interactions\"`):\n- `interaction_id` - Unique identifier for the chat session\n- `agent_id` - The agent identifier\n- `agent_name` - The agent name\n- `interaction_type` - Type of interaction (e.g., chat, slack, api, triggered)\n- `interaction_name` - Display name of the chat session\n- `trigger_type` - For triggered interactions, the specific trigger type (e.g., time_based, polling_new_record_salesforce). Null for non-triggered interactions.\n- `interaction_created_ts` - Chat session creation timestamp\n- `user_email` - Email of the user who initiated the chat\n- `credit_cost` - Total credits consumed (LLM + tool + flow)\n- `llm_credit_cost` - Credits consumed by LLM calls only\n- `tool_credit_cost` - Credits consumed by tool calls\n- `flow_credit_cost` - Credits consumed by pipeline runs\n- `message_count` - Number of messages in the conversation\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n\n**Interaction evaluation fields** (when `data_type` is `\"interaction_evaluations\"`):\n- `evaluation_id` - Unique identifier for the evaluation row\n- `interaction_id` - The chat session that was evaluated (one chat can have several evaluation rows)\n- `agent_id` - The identifier of the agent that was evaluated\n- `organization_evaluation_id` - The organization-level evaluation this row belongs to (empty for agent-level rubrics)\n- `evaluation_created_ts` - When the evaluation reached its final state\n- `status` - The evaluation's state\n- `grade` - The grade the evaluation produced\n- `call_outcome` - The outcome the evaluation recorded for the conversation\n- `sentiment` - The sentiment the evaluation recorded\n- `error_code` - Error code when the evaluation could not complete\n- `summary` - The evaluation's written summary\n- `evaluation_model` - The model that ran the evaluation\n- `credit_cost` - Credits consumed by the evaluation\n- `user_email` - Email of the user whose chat was evaluated\n\n**Credit log fields** (when `data_type` is `\"credit_logs\"`):\n- `user_email` - Email of the user associated with the credit log entry\n- `permission_group_id` - Custom role ID(s) the user belongs to (semicolon-separated if multiple) (disabled by default)\n- `permission_group_name` - Custom role name(s) the user belongs to (semicolon-separated if multiple) (disabled by default)\n- `timestamp` - When the credit transaction occurred\n- `category` - The category of the credit log (e.g., PIPELINE_RUN, AGENT_RUN)\n- `type` - The specific type of credit charge\n- `name` - Display name of the credit log entry\n- `amount` - Number of credits charged or adjusted\n- `balance` - Credit balance after the transaction\n- `log_id` - Unique identifier for the credit log entry\n- `correlation_id` - Join key to the related run or interaction (disabled by default)\n- `balance_scope` - Whether the balance is organization- or user-scoped (disabled by default)\n- `project_id` - The project identifier (disabled by default)\n\nNot all combinations of selected fields are guaranteed to produce data for every row.\n",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / $defs / export_level
        Added value: +{
        +  "default": "organization",
        +  "description": "The scope of the export. Use `\"organization\"` to export across the entire organization, or `\"workspace\"` to export from a single workspace (requires exactly one ID in `workspace_ids`). Defaults to `\"organization\"`. **Not applicable when `data_type` is `\"credit_logs\"`.**\n",
        +  "enum": [
        +    "workspace",
        +    "organization"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / $defs / include_all_workspaces
        Added value: +{
        +  "default": false,
        +  "description": "Whether to include all workspaces in the organization. When `true`, also sets `include_personal_workspaces` to `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**",
        +  "type": "boolean"
        +}
      • addedInput schema / $defs / include_personal_workspaces
        Added value: +{
        +  "default": false,
        +  "description": "Whether to include personal workspaces in the export. Ignored if `include_all_workspaces` is `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**",
        +  "type": "boolean"
        +}
      • addedInput schema / $defs / workspace_ids
        Added value: +{
        +  "description": "An optional array of workspace IDs to include in the export. When `export_level` is `\"workspace\"`, exactly one workspace ID is required. Ignored if `include_all_workspaces` is `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / properties / category_filter / $ref
        Added value: +"#/$defs/category_filter"
      • removedInput schema / properties / category_filter / description
        Removed value: -"**Only applicable when `data_type` is `\"credit_logs\"`.** Optional filter to export only credit logs matching a specific category (e.g., `\"PIPELINE_RUN\"`, `\"AGENT_RUN\"`, `\"CUSTOM_NODE_RD\"`, `\"EXTERNAL_GUMCP_CALL\"`). When omitted, all categories are included.\n"
      • removedInput schema / properties / category_filter / type
        Removed value: -"string"
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / data_type / $ref
        Added value: +"#/$defs/data_type"
      • removedInput schema / properties / data_type / default
        Removed value: -"workflows"
      • removedInput schema / properties / data_type / description
        Removed value: -"The type of data to export. Use `\"workflows\"` to export workflow run data, `\"agents\"` to export agent configuration data, `\"agent_interactions\"` to export agent interaction data, `\"credit_logs\"` to export credit transaction history, `\"interaction_evaluations\"` to export completed chat evaluations, or `\"gumstack\"` to export Gumstack tool call activity. Defaults to `\"workflows\"`. Audit-log exports are available in the Gumloop app only, not through this endpoint.\n"
      • removedInput schema / properties / data_type / enum
        Removed value: -[
        -  "workflows",
        -  "agents",
        -  "agent_interactions",
        -  "credit_logs",
        -  "interaction_evaluations",
        -  "gumstack"
        -]
      • removedInput schema / properties / data_type / type
        Removed value: -"string"
      • addedInput schema / properties / entity_ids / $ref
        Added value: +"#/$defs/entity_ids"
      • removedInput schema / properties / entity_ids / description
        Removed value: -"An optional array of specific entity IDs to filter the export. For workflow exports (`data_type: \"workflows\"`), these are workbook IDs. For agent exports (`data_type: \"agents\"`) and agent interaction exports (`data_type: \"agent_interactions\"`), these are agent IDs. When provided, only data for the specified entities will be included. **Not applicable when `data_type` is `\"credit_logs\"`.**\n"
      • removedInput schema / properties / entity_ids / items
        Removed value: -{
        -  "type": "string"
        -}
      • removedInput schema / properties / entity_ids / type
        Removed value: -"array"
      • addedInput schema / properties / export_fields / $ref
        Added value: +"#/$defs/export_fields"
      • removedInput schema / properties / export_fields / description
        Removed value: -"Array of fields to include in the export. The available fields depend on the `data_type`.\n\n**Workflow fields** (when `data_type` is `\"workflows\"`):\n- `workbook_id` - The workbook identifier\n- `workbook_name` - The workbook name\n- `workbook_created_ts` - Workbook creation timestamp\n- `user_id` - The user identifier\n- `user_email` - The user's email address\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n- `run_id` - The flow run identifier\n- `credit_cost` - Credits consumed by the run\n- `pl_run_created_ts` - Flow run creation timestamp\n- `pl_run_finished_ts` - Flow run completion timestamp\n- `pipeline` - The full pipeline configuration (JSON)\n\n**Agent fields** (when `data_type` is `\"agents\"`):\n- `agent_id` - The agent identifier\n- `agent_name` - The agent name\n- `agent_description` - The agent description\n- `agent_model` - The model used by the agent\n- `agent_system_prompt` - The agent's system prompt\n- `agent_created_ts` - Agent creation timestamp\n- `agent_tools` - Tools configured for the agent (JSON)\n- `agent_metadata` - Additional agent metadata (JSON)\n- `agent_evaluations_enabled` - Whether the agent's own Evaluations are turned on (`true`/`false`)\n- `creator_email` - Email of the user who created the agent\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n- `folder_id` - The identifier of the folder containing the agent (empty when the agent is not in a folder)\n- `folder_name` - The name of the folder containing the agent (empty when the agent is not in a folder)\n\n**Agent interaction fields** (when `data_type` is `\"agent_interactions\"`):\n- `interaction_id` - Unique identifier for the chat session\n- `agent_id` - The agent identifier\n- `agent_name` - The agent name\n- `interaction_type` - Type of interaction (e.g., chat, slack, api, triggered)\n- `interaction_name` - Display name of the chat session\n- `trigger_type` - For triggered interactions, the specific trigger type (e.g., time_based, polling_new_record_salesforce). Null for non-triggered interactions.\n- `interaction_created_ts` - Chat session creation timestamp\n- `user_email` - Email of the user who initiated the chat\n- `credit_cost` - Total credits consumed (LLM + tool + flow)\n- `llm_credit_cost` - Credits consumed by LLM calls only\n- `tool_credit_cost` - Credits consumed by tool calls\n- `flow_credit_cost` - Credits consumed by pipeline runs\n- `message_count` - Number of messages in the conversation\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n\n**Interaction evaluation fields** (when `data_type` is `\"interaction_evaluations\"`):\n- `evaluation_id` - Unique identifier for the evaluation row\n- `interaction_id` - The chat session that was evaluated (one chat can have several evaluation rows)\n- `agent_id` - The identifier of the agent that was evaluated\n- `organization_evaluation_id` - The organization-level evaluation this row belongs to (empty for agent-level rubrics)\n- `evaluation_created_ts` - When the evaluation reached its final state\n- `status` - The evaluation's state\n- `grade` - The grade the evaluation produced\n- `call_outcome` - The outcome the evaluation recorded for the conversation\n- `sentiment` - The sentiment the evaluation recorded\n- `error_code` - Error code when the evaluation could not complete\n- `summary` - The evaluation's written summary\n- `evaluation_model` - The model that ran the evaluation\n- `credit_cost` - Credits consumed by the evaluation\n- `user_email` - Email of the user whose chat was evaluated\n\n**Credit log fields** (when `data_type` is `\"credit_logs\"`):\n- `user_email` - Email of the user associated with the credit log entry\n- `permission_group_id` - Custom role ID(s) the user belongs to (semicolon-separated if multiple) (disabled by default)\n- `permission_group_name` - Custom role name(s) the user belongs to (semicolon-separated if multiple) (disabled by default)\n- `timestamp` - When the credit transaction occurred\n- `category` - The category of the credit log (e.g., PIPELINE_RUN, AGENT_RUN)\n- `type` - The specific type of credit charge\n- `name` - Display name of the credit log entry\n- `amount` - Number of credits charged or adjusted\n- `balance` - Credit balance after the transaction\n- `log_id` - Unique identifier for the credit log entry\n- `correlation_id` - Join key to the related run or interaction (disabled by default)\n- `balance_scope` - Whether the balance is organization- or user-scoped (disabled by default)\n- `project_id` - The project identifier (disabled by default)\n\nNot all combinations of selected fields are guaranteed to produce data for every row.\n"
      • removedInput schema / properties / export_fields / items
        Removed value: -{
        -  "type": "string"
        -}
      • removedInput schema / properties / export_fields / type
        Removed value: -"array"
      • addedInput schema / properties / export_level / $ref
        Added value: +"#/$defs/export_level"
      • removedInput schema / properties / export_level / default
        Removed value: -"organization"
      • removedInput schema / properties / export_level / description
        Removed value: -"The scope of the export. Use `\"organization\"` to export across the entire organization, or `\"workspace\"` to export from a single workspace (requires exactly one ID in `workspace_ids`). Defaults to `\"organization\"`. **Not applicable when `data_type` is `\"credit_logs\"`.**\n"
      • removedInput schema / properties / export_level / enum
        Removed value: -[
        -  "workspace",
        -  "organization"
        -]
      • removedInput schema / properties / export_level / type
        Removed value: -"string"
      • addedInput schema / properties / include_all_workspaces / $ref
        Added value: +"#/$defs/include_all_workspaces"
      • removedInput schema / properties / include_all_workspaces / default
        Removed value: -false
      • removedInput schema / properties / include_all_workspaces / description
        Removed value: -"Whether to include all workspaces in the organization. When `true`, also sets `include_personal_workspaces` to `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**"
      • removedInput schema / properties / include_all_workspaces / type
        Removed value: -"boolean"
      • addedInput schema / properties / include_personal_workspaces / $ref
        Added value: +"#/$defs/include_personal_workspaces"
      • removedInput schema / properties / include_personal_workspaces / default
        Removed value: -false
      • removedInput schema / properties / include_personal_workspaces / description
        Removed value: -"Whether to include personal workspaces in the export. Ignored if `include_all_workspaces` is `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**"
      • removedInput schema / properties / include_personal_workspaces / type
        Removed value: -"boolean"
      • addedInput schema / properties / payload / properties / category_filter / $ref
        Added value: +"#/$defs/category_filter"
      • removedInput schema / properties / payload / properties / category_filter / description
        Removed value: -"**Only applicable when `data_type` is `\"credit_logs\"`.** Optional filter to export only credit logs matching a specific category (e.g., `\"PIPELINE_RUN\"`, `\"AGENT_RUN\"`, `\"CUSTOM_NODE_RD\"`, `\"EXTERNAL_GUMCP_CALL\"`). When omitted, all categories are included.\n"
      • removedInput schema / properties / payload / properties / category_filter / type
        Removed value: -"string"
      • addedInput schema / properties / payload / properties / data_type / $ref
        Added value: +"#/$defs/data_type"
      • removedInput schema / properties / payload / properties / data_type / default
        Removed value: -"workflows"
      • removedInput schema / properties / payload / properties / data_type / description
        Removed value: -"The type of data to export. Use `\"workflows\"` to export workflow run data, `\"agents\"` to export agent configuration data, `\"agent_interactions\"` to export agent interaction data, `\"credit_logs\"` to export credit transaction history, `\"interaction_evaluations\"` to export completed chat evaluations, or `\"gumstack\"` to export Gumstack tool call activity. Defaults to `\"workflows\"`. Audit-log exports are available in the Gumloop app only, not through this endpoint.\n"
      • removedInput schema / properties / payload / properties / data_type / enum
        Removed value: -[
        -  "workflows",
        -  "agents",
        -  "agent_interactions",
        -  "credit_logs",
        -  "interaction_evaluations",
        -  "gumstack"
        -]
      • removedInput schema / properties / payload / properties / data_type / type
        Removed value: -"string"
      • addedInput schema / properties / payload / properties / entity_ids / $ref
        Added value: +"#/$defs/entity_ids"
      • removedInput schema / properties / payload / properties / entity_ids / description
        Removed value: -"An optional array of specific entity IDs to filter the export. For workflow exports (`data_type: \"workflows\"`), these are workbook IDs. For agent exports (`data_type: \"agents\"`) and agent interaction exports (`data_type: \"agent_interactions\"`), these are agent IDs. When provided, only data for the specified entities will be included. **Not applicable when `data_type` is `\"credit_logs\"`.**\n"
      • removedInput schema / properties / payload / properties / entity_ids / items
        Removed value: -{
        -  "type": "string"
        -}
      • removedInput schema / properties / payload / properties / entity_ids / type
        Removed value: -"array"
      • addedInput schema / properties / payload / properties / export_fields / $ref
        Added value: +"#/$defs/export_fields"
      • removedInput schema / properties / payload / properties / export_fields / description
        Removed value: -"Array of fields to include in the export. The available fields depend on the `data_type`.\n\n**Workflow fields** (when `data_type` is `\"workflows\"`):\n- `workbook_id` - The workbook identifier\n- `workbook_name` - The workbook name\n- `workbook_created_ts` - Workbook creation timestamp\n- `user_id` - The user identifier\n- `user_email` - The user's email address\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n- `run_id` - The flow run identifier\n- `credit_cost` - Credits consumed by the run\n- `pl_run_created_ts` - Flow run creation timestamp\n- `pl_run_finished_ts` - Flow run completion timestamp\n- `pipeline` - The full pipeline configuration (JSON)\n\n**Agent fields** (when `data_type` is `\"agents\"`):\n- `agent_id` - The agent identifier\n- `agent_name` - The agent name\n- `agent_description` - The agent description\n- `agent_model` - The model used by the agent\n- `agent_system_prompt` - The agent's system prompt\n- `agent_created_ts` - Agent creation timestamp\n- `agent_tools` - Tools configured for the agent (JSON)\n- `agent_metadata` - Additional agent metadata (JSON)\n- `agent_evaluations_enabled` - Whether the agent's own Evaluations are turned on (`true`/`false`)\n- `creator_email` - Email of the user who created the agent\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n- `folder_id` - The identifier of the folder containing the agent (empty when the agent is not in a folder)\n- `folder_name` - The name of the folder containing the agent (empty when the agent is not in a folder)\n\n**Agent interaction fields** (when `data_type` is `\"agent_interactions\"`):\n- `interaction_id` - Unique identifier for the chat session\n- `agent_id` - The agent identifier\n- `agent_name` - The agent name\n- `interaction_type` - Type of interaction (e.g., chat, slack, api, triggered)\n- `interaction_name` - Display name of the chat session\n- `trigger_type` - For triggered interactions, the specific trigger type (e.g., time_based, polling_new_record_salesforce). Null for non-triggered interactions.\n- `interaction_created_ts` - Chat session creation timestamp\n- `user_email` - Email of the user who initiated the chat\n- `credit_cost` - Total credits consumed (LLM + tool + flow)\n- `llm_credit_cost` - Credits consumed by LLM calls only\n- `tool_credit_cost` - Credits consumed by tool calls\n- `flow_credit_cost` - Credits consumed by pipeline runs\n- `message_count` - Number of messages in the conversation\n- `workspace_id` - The workspace identifier\n- `workspace_name` - The workspace name\n\n**Interaction evaluation fields** (when `data_type` is `\"interaction_evaluations\"`):\n- `evaluation_id` - Unique identifier for the evaluation row\n- `interaction_id` - The chat session that was evaluated (one chat can have several evaluation rows)\n- `agent_id` - The identifier of the agent that was evaluated\n- `organization_evaluation_id` - The organization-level evaluation this row belongs to (empty for agent-level rubrics)\n- `evaluation_created_ts` - When the evaluation reached its final state\n- `status` - The evaluation's state\n- `grade` - The grade the evaluation produced\n- `call_outcome` - The outcome the evaluation recorded for the conversation\n- `sentiment` - The sentiment the evaluation recorded\n- `error_code` - Error code when the evaluation could not complete\n- `summary` - The evaluation's written summary\n- `evaluation_model` - The model that ran the evaluation\n- `credit_cost` - Credits consumed by the evaluation\n- `user_email` - Email of the user whose chat was evaluated\n\n**Credit log fields** (when `data_type` is `\"credit_logs\"`):\n- `user_email` - Email of the user associated with the credit log entry\n- `permission_group_id` - Custom role ID(s) the user belongs to (semicolon-separated if multiple) (disabled by default)\n- `permission_group_name` - Custom role name(s) the user belongs to (semicolon-separated if multiple) (disabled by default)\n- `timestamp` - When the credit transaction occurred\n- `category` - The category of the credit log (e.g., PIPELINE_RUN, AGENT_RUN)\n- `type` - The specific type of credit charge\n- `name` - Display name of the credit log entry\n- `amount` - Number of credits charged or adjusted\n- `balance` - Credit balance after the transaction\n- `log_id` - Unique identifier for the credit log entry\n- `correlation_id` - Join key to the related run or interaction (disabled by default)\n- `balance_scope` - Whether the balance is organization- or user-scoped (disabled by default)\n- `project_id` - The project identifier (disabled by default)\n\nNot all combinations of selected fields are guaranteed to produce data for every row.\n"
      • removedInput schema / properties / payload / properties / export_fields / items
        Removed value: -{
        -  "type": "string"
        -}
      • removedInput schema / properties / payload / properties / export_fields / type
        Removed value: -"array"
      • addedInput schema / properties / payload / properties / export_level / $ref
        Added value: +"#/$defs/export_level"
      • removedInput schema / properties / payload / properties / export_level / default
        Removed value: -"organization"
      • removedInput schema / properties / payload / properties / export_level / description
        Removed value: -"The scope of the export. Use `\"organization\"` to export across the entire organization, or `\"workspace\"` to export from a single workspace (requires exactly one ID in `workspace_ids`). Defaults to `\"organization\"`. **Not applicable when `data_type` is `\"credit_logs\"`.**\n"
      • removedInput schema / properties / payload / properties / export_level / enum
        Removed value: -[
        -  "workspace",
        -  "organization"
        -]
      • removedInput schema / properties / payload / properties / export_level / type
        Removed value: -"string"
      • addedInput schema / properties / payload / properties / include_all_workspaces / $ref
        Added value: +"#/$defs/include_all_workspaces"
      • removedInput schema / properties / payload / properties / include_all_workspaces / default
        Removed value: -false
      • removedInput schema / properties / payload / properties / include_all_workspaces / description
        Removed value: -"Whether to include all workspaces in the organization. When `true`, also sets `include_personal_workspaces` to `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**"
      • removedInput schema / properties / payload / properties / include_all_workspaces / type
        Removed value: -"boolean"
      • addedInput schema / properties / payload / properties / include_personal_workspaces / $ref
        Added value: +"#/$defs/include_personal_workspaces"
      • removedInput schema / properties / payload / properties / include_personal_workspaces / default
        Removed value: -false
      • removedInput schema / properties / payload / properties / include_personal_workspaces / description
        Removed value: -"Whether to include personal workspaces in the export. Ignored if `include_all_workspaces` is `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**"
      • removedInput schema / properties / payload / properties / include_personal_workspaces / type
        Removed value: -"boolean"
      • addedInput schema / properties / payload / properties / workspace_ids / $ref
        Added value: +"#/$defs/workspace_ids"
      • removedInput schema / properties / payload / properties / workspace_ids / description
        Removed value: -"An optional array of workspace IDs to include in the export. When `export_level` is `\"workspace\"`, exactly one workspace ID is required. Ignored if `include_all_workspaces` is `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**"
      • removedInput schema / properties / payload / properties / workspace_ids / items
        Removed value: -{
        -  "type": "string"
        -}
      • removedInput schema / properties / payload / properties / workspace_ids / type
        Removed value: -"array"
      • addedInput schema / properties / workspace_ids / $ref
        Added value: +"#/$defs/workspace_ids"
      • removedInput schema / properties / workspace_ids / description
        Removed value: -"An optional array of workspace IDs to include in the export. When `export_level` is `\"workspace\"`, exactly one workspace ID is required. Ignored if `include_all_workspaces` is `true`. **Not applicable when `data_type` is `\"credit_logs\"`.**"
      • removedInput schema / properties / workspace_ids / items
        Removed value: -{
        -  "type": "string"
        -}
      • removedInput schema / properties / workspace_ids / type
        Removed value: -"array"
    • Changedimport_browser_profile_cookies1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedkill_flow1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedmanage_permission_group_users1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedmanage_project_users1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedqueue_session_message1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedrename_session12 fields changed
      • addedInput schema / $defs / name
        Added value: +{
        +  "description": "New name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so a whitespace-only value is rejected.",
        +  "maxLength": 256,
        +  "minLength": 1,
        +  "type": "string"
        +}
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / name / $ref
        Added value: +"#/$defs/name"
      • removedInput schema / properties / name / description
        Removed value: -"New name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so a whitespace-only value is rejected."
      • removedInput schema / properties / name / maxLength
        Removed value: -256
      • removedInput schema / properties / name / minLength
        Removed value: -1
      • removedInput schema / properties / name / type
        Removed value: -"string"
      • addedInput schema / properties / payload / properties / name / $ref
        Added value: +"#/$defs/name"
      • removedInput schema / properties / payload / properties / name / description
        Removed value: -"New name for the session. Leading and trailing whitespace is removed before the 1-256 character limit is applied, so a whitespace-only value is rejected."
      • removedInput schema / properties / payload / properties / name / maxLength
        Removed value: -256
      • removedInput schema / properties / payload / properties / name / minLength
        Removed value: -1
      • removedInput schema / properties / payload / properties / name / type
        Removed value: -"string"
    • Changedresolve_session_approvals14 fields changed
      • addedInput schema / $defs / approval_responses
        Added value: +{
        +  "description": "Answers to pending asks. Each `action_request_id` may appear at most once.",
        +  "items": {
        +    "properties": {
        +      "action": {
        +        "description": "Whether to approve or reject the pending ask.",
        +        "enum": [
        +          "accept",
        +          "reject"
        +        ],
        +        "type": "string"
        +      },
        +      "action_request_id": {
        +        "description": "ID of the pending ask, from `pending_approvals` on Retrieve session.",
        +        "type": "string"
        +      },
        +      "reason": {
        +        "description": "Optional reason recorded with the resolution.",
        +        "maxLength": 1000,
        +        "type": "string"
        +      },
        +      "response": {
        +        "description": "For `human_input` asks — form answers keyed by question name.",
        +        "properties": {
        +          "values": {
        +            "description": "Map of question name to answer.",
        +            "type": "object"
        +          }
        +        },
        +        "type": "object"
        +      }
        +    },
        +    "required": [
        +      "action_request_id",
        +      "action"
        +    ],
        +    "type": "object"
        +  },
        +  "maxItems": 20,
        +  "minItems": 1,
        +  "type": "array"
        +}
      • addedInput schema / properties / approval_responses / $ref
        Added value: +"#/$defs/approval_responses"
      • removedInput schema / properties / approval_responses / description
        Removed value: -"Answers to pending asks. Each `action_request_id` may appear at most once."
      • removedInput schema / properties / approval_responses / items
        Removed value: -{
        -  "properties": {
        -    "action": {
        -      "description": "Whether to approve or reject the pending ask.",
        -      "enum": [
        -        "accept",
        -        "reject"
        -      ],
        -      "type": "string"
        -    },
        -    "action_request_id": {
        -      "description": "ID of the pending ask, from `pending_approvals` on Retrieve session.",
        -      "type": "string"
        -    },
        -    "reason": {
        -      "description": "Optional reason recorded with the resolution.",
        -      "maxLength": 1000,
        -      "type": "string"
        -    },
        -    "response": {
        -      "description": "For `human_input` asks — form answers keyed by question name.",
        -      "properties": {
        -        "values": {
        -          "description": "Map of question name to answer.",
        -          "type": "object"
        -        }
        -      },
        -      "type": "object"
        -    }
        -  },
        -  "required": [
        -    "action_request_id",
        -    "action"
        -  ],
        -  "type": "object"
        -}
      • removedInput schema / properties / approval_responses / maxItems
        Removed value: -20
      • removedInput schema / properties / approval_responses / minItems
        Removed value: -1
      • removedInput schema / properties / approval_responses / type
        Removed value: -"array"
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / payload / properties / approval_responses / $ref
        Added value: +"#/$defs/approval_responses"
      • removedInput schema / properties / payload / properties / approval_responses / description
        Removed value: -"Answers to pending asks. Each `action_request_id` may appear at most once."
      • removedInput schema / properties / payload / properties / approval_responses / items
        Removed value: -{
        -  "properties": {
        -    "action": {
        -      "description": "Whether to approve or reject the pending ask.",
        -      "enum": [
        -        "accept",
        -        "reject"
        -      ],
        -      "type": "string"
        -    },
        -    "action_request_id": {
        -      "description": "ID of the pending ask, from `pending_approvals` on Retrieve session.",
        -      "type": "string"
        -    },
        -    "reason": {
        -      "description": "Optional reason recorded with the resolution.",
        -      "maxLength": 1000,
        -      "type": "string"
        -    },
        -    "response": {
        -      "description": "For `human_input` asks — form answers keyed by question name.",
        -      "properties": {
        -        "values": {
        -          "description": "Map of question name to answer.",
        -          "type": "object"
        -        }
        -      },
        -      "type": "object"
        -    }
        -  },
        -  "required": [
        -    "action_request_id",
        -    "action"
        -  ],
        -  "type": "object"
        -}
      • removedInput schema / properties / payload / properties / approval_responses / maxItems
        Removed value: -20
      • removedInput schema / properties / payload / properties / approval_responses / minItems
        Removed value: -1
      • removedInput schema / properties / payload / properties / approval_responses / type
        Removed value: -"array"
    • Changedroute_model44 fields changed
      • addedInput schema / $defs / agent
        Added value: +{
        +  "additionalProperties": false,
        +  "description": "Optional context about the agent the message is for. Sharper context produces a sharper route.",
        +  "properties": {
        +    "description": {
        +      "maxLength": 2000,
        +      "type": "string"
        +    },
        +    "name": {
        +      "maxLength": 200,
        +      "type": "string"
        +    },
        +    "system_prompt": {
        +      "maxLength": 20000,
        +      "type": "string"
        +    }
        +  },
        +  "type": "object"
        +}
      • addedInput schema / $defs / history
        Added value: +{
        +  "description": "Prior turns, oldest first, for context.",
        +  "items": {
        +    "additionalProperties": false,
        +    "properties": {
        +      "content": {
        +        "description": "Message text, as a string or text parts. `input` and `message` are accepted as aliases.",
        +        "oneOf": [
        +          {
        +            "type": "string"
        +          },
        +          {
        +            "items": {
        +              "type": "object"
        +            },
        +            "type": "array"
        +          }
        +        ]
        +      },
        +      "model": {
        +        "description": "The model that produced this turn. Assistant messages only.",
        +        "maxLength": 200,
        +        "type": "string"
        +      },
        +      "role": {
        +        "enum": [
        +          "user",
        +          "assistant"
        +        ],
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "role",
        +      "content"
        +    ],
        +    "type": "object"
        +  },
        +  "maxItems": 20,
        +  "type": "array"
        +}
      • addedInput schema / $defs / input
        Added value: +{
        +  "description": "The message to route. `message` is accepted as an alias. A string or an array of `{type: text, text: ...}` parts.",
        +  "oneOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "items": {
        +        "type": "object"
        +      },
        +      "type": "array"
        +    }
        +  ]
        +}
      • addedInput schema / $defs / models
        Added value: +{
        +  "description": "Candidate model IDs to choose between. IDs must be nonempty, unique after remapping, and registered. Omit to use Chew's lane-chain union.",
        +  "items": {
        +    "type": "string"
        +  },
        +  "maxItems": 50,
        +  "minItems": 1,
        +  "type": "array",
        +  "uniqueItems": true
        +}
      • addedInput schema / properties / agent / $ref
        Added value: +"#/$defs/agent"
      • removedInput schema / properties / agent / additionalProperties
        Removed value: -false
      • removedInput schema / properties / agent / description
        Removed value: -"Optional context about the agent the message is for. Sharper context produces a sharper route."
      • removedInput schema / properties / agent / properties
        Removed value: -{
        -  "description": {
        -    "maxLength": 2000,
        -    "type": "string"
        -  },
        -  "name": {
        -    "maxLength": 200,
        -    "type": "string"
        -  },
        -  "system_prompt": {
        -    "maxLength": 20000,
        -    "type": "string"
        -  }
        -}
      • removedInput schema / properties / agent / type
        Removed value: -"object"
      • addedInput schema / properties / history / $ref
        Added value: +"#/$defs/history"
      • removedInput schema / properties / history / description
        Removed value: -"Prior turns, oldest first, for context."
      • removedInput schema / properties / history / items
        Removed value: -{
        -  "additionalProperties": false,
        -  "properties": {
        -    "content": {
        -      "description": "Message text, as a string or text parts. `input` and `message` are accepted as aliases.",
        -      "oneOf": [
        -        {
        -          "type": "string"
        -        },
        -        {
        -          "items": {
        -            "type": "object"
        -          },
        -          "type": "array"
        -        }
        -      ]
        -    },
        -    "model": {
        -      "description": "The model that produced this turn. Assistant messages only.",
        -      "maxLength": 200,
        -      "type": "string"
        -    },
        -    "role": {
        -      "enum": [
        -        "user",
        -        "assistant"
        -      ],
        -      "type": "string"
        -    }
        -  },
        -  "required": [
        -    "role",
        -    "content"
        -  ],
        -  "type": "object"
        -}
      • removedInput schema / properties / history / maxItems
        Removed value: -20
      • removedInput schema / properties / history / type
        Removed value: -"array"
      • addedInput schema / properties / input / $ref
        Added value: +"#/$defs/input"
      • removedInput schema / properties / input / description
        Removed value: -"The message to route. `message` is accepted as an alias. A string or an array of `{type: text, text: ...}` parts."
      • removedInput schema / properties / input / oneOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "items": {
        -      "type": "object"
        -    },
        -    "type": "array"
        -  }
        -]
      • addedInput schema / properties / models / $ref
        Added value: +"#/$defs/models"
      • removedInput schema / properties / models / description
        Removed value: -"Candidate model IDs to choose between. IDs must be nonempty, unique after remapping, and registered. Omit to use Chew's lane-chain union."
      • removedInput schema / properties / models / items
        Removed value: -{
        -  "type": "string"
        -}
      • removedInput schema / properties / models / maxItems
        Removed value: -50
      • removedInput schema / properties / models / minItems
        Removed value: -1
      • removedInput schema / properties / models / type
        Removed value: -"array"
      • removedInput schema / properties / models / uniqueItems
        Removed value: -true
      • addedInput schema / properties / payload / properties / agent / $ref
        Added value: +"#/$defs/agent"
      • removedInput schema / properties / payload / properties / agent / additionalProperties
        Removed value: -false
      • removedInput schema / properties / payload / properties / agent / description
        Removed value: -"Optional context about the agent the message is for. Sharper context produces a sharper route."
      • removedInput schema / properties / payload / properties / agent / properties
        Removed value: -{
        -  "description": {
        -    "maxLength": 2000,
        -    "type": "string"
        -  },
        -  "name": {
        -    "maxLength": 200,
        -    "type": "string"
        -  },
        -  "system_prompt": {
        -    "maxLength": 20000,
        -    "type": "string"
        -  }
        -}
      • removedInput schema / properties / payload / properties / agent / type
        Removed value: -"object"
      • addedInput schema / properties / payload / properties / history / $ref
        Added value: +"#/$defs/history"
      • removedInput schema / properties / payload / properties / history / description
        Removed value: -"Prior turns, oldest first, for context."
      • removedInput schema / properties / payload / properties / history / items
        Removed value: -{
        -  "additionalProperties": false,
        -  "properties": {
        -    "content": {
        -      "description": "Message text, as a string or text parts. `input` and `message` are accepted as aliases.",
        -      "oneOf": [
        -        {
        -          "type": "string"
        -        },
        -        {
        -          "items": {
        -            "type": "object"
        -          },
        -          "type": "array"
        -        }
        -      ]
        -    },
        -    "model": {
        -      "description": "The model that produced this turn. Assistant messages only.",
        -      "maxLength": 200,
        -      "type": "string"
        -    },
        -    "role": {
        -      "enum": [
        -        "user",
        -        "assistant"
        -      ],
        -      "type": "string"
        -    }
        -  },
        -  "required": [
        -    "role",
        -    "content"
        -  ],
        -  "type": "object"
        -}
      • removedInput schema / properties / payload / properties / history / maxItems
        Removed value: -20
      • removedInput schema / properties / payload / properties / history / type
        Removed value: -"array"
      • addedInput schema / properties / payload / properties / input / $ref
        Added value: +"#/$defs/input"
      • removedInput schema / properties / payload / properties / input / description
        Removed value: -"The message to route. `message` is accepted as an alias. A string or an array of `{type: text, text: ...}` parts."
      • removedInput schema / properties / payload / properties / input / oneOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "items": {
        -      "type": "object"
        -    },
        -    "type": "array"
        -  }
        -]
      • addedInput schema / properties / payload / properties / models / $ref
        Added value: +"#/$defs/models"
      • removedInput schema / properties / payload / properties / models / description
        Removed value: -"Candidate model IDs to choose between. IDs must be nonempty, unique after remapping, and registered. Omit to use Chew's lane-chain union."
      • removedInput schema / properties / payload / properties / models / items
        Removed value: -{
        -  "type": "string"
        -}
      • removedInput schema / properties / payload / properties / models / maxItems
        Removed value: -50
      • removedInput schema / properties / payload / properties / models / minItems
        Removed value: -1
      • removedInput schema / properties / payload / properties / models / type
        Removed value: -"array"
      • removedInput schema / properties / payload / properties / models / uniqueItems
        Removed value: -true
    • Changedrun_evaluations1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedrun_organization_evaluation1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedsearch_brain12 fields changed
      • addedInput schema / $defs / source_type
        Added value: +{
        +  "description": "Restrict results to specific source types. Omit to search every source you can access. Valid values: `notion`, `google_drive`, `slack`, `github`, `confluence`, `direct_file_uploads`, `gumloop_artifacts`.\n",
        +  "items": {
        +    "enum": [
        +      "notion",
        +      "google_drive",
        +      "slack",
        +      "github",
        +      "confluence",
        +      "direct_file_uploads",
        +      "gumloop_artifacts"
        +    ],
        +    "type": "string"
        +  },
        +  "minItems": 1,
        +  "type": [
        +    "array",
        +    "null"
        +  ]
        +}
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / payload / properties / source_type / $ref
        Added value: +"#/$defs/source_type"
      • removedInput schema / properties / payload / properties / source_type / description
        Removed value: -"Restrict results to specific source types. Omit to search every source you can access. Valid values: `notion`, `google_drive`, `slack`, `github`, `confluence`, `direct_file_uploads`, `gumloop_artifacts`.\n"
      • removedInput schema / properties / payload / properties / source_type / items
        Removed value: -{
        -  "enum": [
        -    "notion",
        -    "google_drive",
        -    "slack",
        -    "github",
        -    "confluence",
        -    "direct_file_uploads",
        -    "gumloop_artifacts"
        -  ],
        -  "type": "string"
        -}
      • removedInput schema / properties / payload / properties / source_type / minItems
        Removed value: -1
      • removedInput schema / properties / payload / properties / source_type / type
        Removed value: -[
        -  "array",
        -  "null"
        -]
      • addedInput schema / properties / source_type / $ref
        Added value: +"#/$defs/source_type"
      • removedInput schema / properties / source_type / description
        Removed value: -"Restrict results to specific source types. Omit to search every source you can access. Valid values: `notion`, `google_drive`, `slack`, `github`, `confluence`, `direct_file_uploads`, `gumloop_artifacts`.\n"
      • removedInput schema / properties / source_type / items
        Removed value: -{
        -  "enum": [
        -    "notion",
        -    "google_drive",
        -    "slack",
        -    "github",
        -    "confluence",
        -    "direct_file_uploads",
        -    "gumloop_artifacts"
        -  ],
        -  "type": "string"
        -}
      • removedInput schema / properties / source_type / minItems
        Removed value: -1
      • removedInput schema / properties / source_type / type
        Removed value: -[
        -  "array",
        -  "null"
        -]
    • Changedsend_message12 fields changed
      • addedInput schema / $defs / attachments
        Added value: +{
        +  "description": "Files to attach to the message. Each `file_name` must be a stored path returned by Upload session file for this session.",
        +  "items": {
        +    "properties": {
        +      "file_name": {
        +        "description": "Stored path returned by Upload session file.",
        +        "type": "string"
        +      },
        +      "media_type": {
        +        "description": "MIME type of the file.",
        +        "type": [
        +          "string",
        +          "null"
        +        ]
        +      }
        +    },
        +    "required": [
        +      "file_name"
        +    ],
        +    "type": "object"
        +  },
        +  "maxItems": 10,
        +  "type": "array"
        +}
      • addedInput schema / properties / attachments / $ref
        Added value: +"#/$defs/attachments"
      • removedInput schema / properties / attachments / description
        Removed value: -"Files to attach to the message. Each `file_name` must be a stored path returned by Upload session file for this session."
      • removedInput schema / properties / attachments / items
        Removed value: -{
        -  "properties": {
        -    "file_name": {
        -      "description": "Stored path returned by Upload session file.",
        -      "type": "string"
        -    },
        -    "media_type": {
        -      "description": "MIME type of the file.",
        -      "type": [
        -        "string",
        -        "null"
        -      ]
        -    }
        -  },
        -  "required": [
        -    "file_name"
        -  ],
        -  "type": "object"
        -}
      • removedInput schema / properties / attachments / maxItems
        Removed value: -10
      • removedInput schema / properties / attachments / type
        Removed value: -"array"
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / payload / properties / attachments / $ref
        Added value: +"#/$defs/attachments"
      • removedInput schema / properties / payload / properties / attachments / description
        Removed value: -"Files to attach to the message. Each `file_name` must be a stored path returned by Upload session file for this session."
      • removedInput schema / properties / payload / properties / attachments / items
        Removed value: -{
        -  "properties": {
        -    "file_name": {
        -      "description": "Stored path returned by Upload session file.",
        -      "type": "string"
        -    },
        -    "media_type": {
        -      "description": "MIME type of the file.",
        -      "type": [
        -        "string",
        -        "null"
        -      ]
        -    }
        -  },
        -  "required": [
        -    "file_name"
        -  ],
        -  "type": "object"
        -}
      • removedInput schema / properties / payload / properties / attachments / maxItems
        Removed value: -10
      • removedInput schema / properties / payload / properties / attachments / type
        Removed value: -"array"
    • Changedsend_queued_message1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedset_organization_evaluation_targets5 fields changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / payload / $defs / EvaluationTarget / $ref
        Added value: +"#/$defs/EvaluationTarget"
      • removedInput schema / properties / payload / $defs / EvaluationTarget / properties
        Removed value: -{
        -  "id": {
        -    "description": "Team, user, or agent ID. Omitted for `organization`; responses return the organization ID.",
        -    "type": "string"
        -  },
        -  "type": {
        -    "description": "What the target expands to. `user` covers a member's personal agents.",
        -    "enum": [
        -      "organization",
        -      "team",
        -      "user",
        -      "agent"
        -    ],
        -    "type": "string"
        -  }
        -}
      • removedInput schema / properties / payload / $defs / EvaluationTarget / required
        Removed value: -[
        -  "type"
        -]
      • removedInput schema / properties / payload / $defs / EvaluationTarget / type
        Removed value: -"object"
    • Changedset_role_credit_limit1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedstart_flow1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedupdate_agent8 fields changed
      • addedInput schema / $defs / is_active
        Added value: +{
        +  "description": "Setting this to `false` retires the agent: it disappears from `GET /agents`, and `GET`/`PATCH /agents/{agent_id}` return `404`, so it cannot be reactivated through the API. This is not a pause switch — to stop an agent from running while keeping it reachable, disable its triggers instead.\n",
        +  "type": [
        +    "boolean",
        +    "null"
        +  ]
        +}
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / is_active / $ref
        Added value: +"#/$defs/is_active"
      • removedInput schema / properties / is_active / description
        Removed value: -"Setting this to `false` retires the agent: it disappears from `GET /agents`, and `GET`/`PATCH /agents/{agent_id}` return `404`, so it cannot be reactivated through the API. This is not a pause switch — to stop an agent from running while keeping it reachable, disable its triggers instead.\n"
      • removedInput schema / properties / is_active / type
        Removed value: -[
        -  "boolean",
        -  "null"
        -]
      • addedInput schema / properties / payload / properties / is_active / $ref
        Added value: +"#/$defs/is_active"
      • removedInput schema / properties / payload / properties / is_active / description
        Removed value: -"Setting this to `false` retires the agent: it disappears from `GET /agents`, and `GET`/`PATCH /agents/{agent_id}` return `404`, so it cannot be reactivated through the API. This is not a pause switch — to stop an agent from running while keeping it reachable, disable its triggers instead.\n"
      • removedInput schema / properties / payload / properties / is_active / type
        Removed value: -[
        -  "boolean",
        -  "null"
        -]
    • Changedupdate_agent_skills1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedupdate_evaluation_config19 fields changed
      • addedInput schema / $defs / criteria
        Added value: +{
        +  "description": "Quality criteria to check (replaces existing list). Max 30.",
        +  "items": {
        +    "properties": {
        +      "name": {
        +        "type": "string"
        +      },
        +      "priority": {
        +        "enum": [
        +          "needs_review",
        +          "needs_attention"
        +        ],
        +        "type": "string"
        +      },
        +      "prompt": {
        +        "description": "True/false statement to evaluate.",
        +        "type": "string"
        +      },
        +      "type": {
        +        "enum": [
        +          "prohibited_action",
        +          "prohibited_words",
        +          "voice_tone",
        +          "other"
        +        ],
        +        "type": "string"
        +      }
        +    },
        +    "type": "object"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / $defs / data_points
        Added value: +{
        +  "description": "Data points to extract (replaces existing list). Max 40.",
        +  "items": {
        +    "properties": {
        +      "data_type": {
        +        "enum": [
        +          "string",
        +          "boolean",
        +          "integer",
        +          "number"
        +        ],
        +        "type": "string"
        +      },
        +      "description": {
        +        "type": "string"
        +      },
        +      "name": {
        +        "type": "string"
        +      }
        +    },
        +    "type": "object"
        +  },
        +  "type": "array"
        +}
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / criteria / $ref
        Added value: +"#/$defs/criteria"
      • removedInput schema / properties / criteria / description
        Removed value: -"Quality criteria to check (replaces existing list). Max 30."
      • removedInput schema / properties / criteria / items
        Removed value: -{
        -  "properties": {
        -    "name": {
        -      "type": "string"
        -    },
        -    "priority": {
        -      "enum": [
        -        "needs_review",
        -        "needs_attention"
        -      ],
        -      "type": "string"
        -    },
        -    "prompt": {
        -      "description": "True/false statement to evaluate.",
        -      "type": "string"
        -    },
        -    "type": {
        -      "enum": [
        -        "prohibited_action",
        -        "prohibited_words",
        -        "voice_tone",
        -        "other"
        -      ],
        -      "type": "string"
        -    }
        -  },
        -  "type": "object"
        -}
      • removedInput schema / properties / criteria / type
        Removed value: -"array"
      • addedInput schema / properties / data_points / $ref
        Added value: +"#/$defs/data_points"
      • removedInput schema / properties / data_points / description
        Removed value: -"Data points to extract (replaces existing list). Max 40."
      • removedInput schema / properties / data_points / items
        Removed value: -{
        -  "properties": {
        -    "data_type": {
        -      "enum": [
        -        "string",
        -        "boolean",
        -        "integer",
        -        "number"
        -      ],
        -      "type": "string"
        -    },
        -    "description": {
        -      "type": "string"
        -    },
        -    "name": {
        -      "type": "string"
        -    }
        -  },
        -  "type": "object"
        -}
      • removedInput schema / properties / data_points / type
        Removed value: -"array"
      • addedInput schema / properties / payload / properties / criteria / $ref
        Added value: +"#/$defs/criteria"
      • removedInput schema / properties / payload / properties / criteria / description
        Removed value: -"Quality criteria to check (replaces existing list). Max 30."
      • removedInput schema / properties / payload / properties / criteria / items
        Removed value: -{
        -  "properties": {
        -    "name": {
        -      "type": "string"
        -    },
        -    "priority": {
        -      "enum": [
        -        "needs_review",
        -        "needs_attention"
        -      ],
        -      "type": "string"
        -    },
        -    "prompt": {
        -      "description": "True/false statement to evaluate.",
        -      "type": "string"
        -    },
        -    "type": {
        -      "enum": [
        -        "prohibited_action",
        -        "prohibited_words",
        -        "voice_tone",
        -        "other"
        -      ],
        -      "type": "string"
        -    }
        -  },
        -  "type": "object"
        -}
      • removedInput schema / properties / payload / properties / criteria / type
        Removed value: -"array"
      • addedInput schema / properties / payload / properties / data_points / $ref
        Added value: +"#/$defs/data_points"
      • removedInput schema / properties / payload / properties / data_points / description
        Removed value: -"Data points to extract (replaces existing list). Max 40."
      • removedInput schema / properties / payload / properties / data_points / items
        Removed value: -{
        -  "properties": {
        -    "data_type": {
        -      "enum": [
        -        "string",
        -        "boolean",
        -        "integer",
        -        "number"
        -      ],
        -      "type": "string"
        -    },
        -    "description": {
        -      "type": "string"
        -    },
        -    "name": {
        -      "type": "string"
        -    }
        -  },
        -  "type": "object"
        -}
      • removedInput schema / properties / payload / properties / data_points / type
        Removed value: -"array"
    • Changedupdate_organization_evaluation1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedupdate_queued_message1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedupdate_skill14 fields changed
      • addedInput schema / $defs / files
        Added value: +{
        +  "description": "Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks.",
        +  "items": {
        +    "minLength": 1,
        +    "type": "string"
        +  },
        +  "maxItems": 25,
        +  "minItems": 1,
        +  "type": "array"
        +}
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / files / $ref
        Added value: +"#/$defs/files"
      • removedInput schema / properties / files / description
        Removed value: -"Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks."
      • removedInput schema / properties / files / items
        Removed value: -{
        -  "minLength": 1,
        -  "type": "string"
        -}
      • removedInput schema / properties / files / maxItems
        Removed value: -25
      • removedInput schema / properties / files / minItems
        Removed value: -1
      • removedInput schema / properties / files / type
        Removed value: -"array"
      • addedInput schema / properties / payload / properties / files / $ref
        Added value: +"#/$defs/files"
      • removedInput schema / properties / payload / properties / files / description
        Removed value: -"Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks."
      • removedInput schema / properties / payload / properties / files / items
        Removed value: -{
        -  "minLength": 1,
        -  "type": "string"
        -}
      • removedInput schema / properties / payload / properties / files / maxItems
        Removed value: -25
      • removedInput schema / properties / payload / properties / files / minItems
        Removed value: -1
      • removedInput schema / properties / payload / properties / files / type
        Removed value: -"array"
    • Changedupload_brain_files14 fields changed
      • addedInput schema / $defs / files
        Added value: +{
        +  "description": "Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks.",
        +  "items": {
        +    "minLength": 1,
        +    "type": "string"
        +  },
        +  "maxItems": 25,
        +  "minItems": 1,
        +  "type": "array"
        +}
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / files / $ref
        Added value: +"#/$defs/files"
      • removedInput schema / properties / files / description
        Removed value: -"Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks."
      • removedInput schema / properties / files / items
        Removed value: -{
        -  "minLength": 1,
        -  "type": "string"
        -}
      • removedInput schema / properties / files / maxItems
        Removed value: -25
      • removedInput schema / properties / files / minItems
        Removed value: -1
      • removedInput schema / properties / files / type
        Removed value: -"array"
      • addedInput schema / properties / payload / properties / files / $ref
        Added value: +"#/$defs/files"
      • removedInput schema / properties / payload / properties / files / description
        Removed value: -"Regular local file paths, not base64; each file and total upload at most 5 MiB. Paths cannot be symlinks."
      • removedInput schema / properties / payload / properties / files / items
        Removed value: -{
        -  "minLength": 1,
        -  "type": "string"
        -}
      • removedInput schema / properties / payload / properties / files / maxItems
        Removed value: -25
      • removedInput schema / properties / payload / properties / files / minItems
        Removed value: -1
      • removedInput schema / properties / payload / properties / files / type
        Removed value: -"array"
    • Changedupload_file1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • Changedupload_files8 fields changed
      • addedInput schema / $defs / files
        Added value: +{
        +  "items": {
        +    "properties": {
        +      "file_content": {
        +        "description": "Base64 encoded content of the file.",
        +        "format": "byte",
        +        "type": "string"
        +      },
        +      "file_name": {
        +        "description": "The name of the file to be uploaded.",
        +        "type": "string"
        +      }
        +    },
        +    "type": "object"
        +  },
        +  "type": "array"
        +}
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
      • addedInput schema / properties / files / $ref
        Added value: +"#/$defs/files"
      • removedInput schema / properties / files / items
        Removed value: -{
        -  "properties": {
        -    "file_content": {
        -      "description": "Base64 encoded content of the file.",
        -      "format": "byte",
        -      "type": "string"
        -    },
        -    "file_name": {
        -      "description": "The name of the file to be uploaded.",
        -      "type": "string"
        -    }
        -  },
        -  "type": "object"
        -}
      • removedInput schema / properties / files / type
        Removed value: -"array"
      • addedInput schema / properties / payload / properties / files / $ref
        Added value: +"#/$defs/files"
      • removedInput schema / properties / payload / properties / files / items
        Removed value: -{
        -  "properties": {
        -    "file_content": {
        -      "description": "Base64 encoded content of the file.",
        -      "format": "byte",
        -      "type": "string"
        -    },
        -    "file_name": {
        -      "description": "The name of the file to be uploaded.",
        -      "type": "string"
        -    }
        -  },
        -  "type": "object"
        -}
      • removedInput schema / properties / payload / properties / files / type
        Removed value: -"array"
    • Changedupload_session_file1 field changed
      • changedInput schema / properties / confirm / description
        Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
  2. 92 tool updatesv2.0.1
    • First observedapprove_brain_source
    • First observedattach_agent_mcp_server
    • First observedcall_mcp_tools
    • First observedcancel_session
    • First observedcreate_agent
    • First observedcreate_brain_source
    • First observedcreate_chat_completion
    • First observedcreate_organization_evaluation
    • First observedcreate_session
    • First observedcreate_skill
    • First observeddelete_brain_file
    • First observeddelete_brain_source
    • First observeddelete_organization_evaluation
    • First observeddelete_queued_message
    • First observeddelete_skill
    • First observeddetach_agent_mcp_server
    • First observeddownload_artifact_file
    • First observeddownload_file
    • First observeddownload_files
    • First observeddownload_skill_file
    • First observedexport_data
    • First observedget_brain_source
    • First observedget_brain_source_estimate
    • First observedget_evaluation_config
    • First observedget_evaluation_metrics
    • First observedget_evaluation_options
    • First observedget_export_status
    • First observedget_input_schema
    • First observedget_mcp_server_prompt
    • First observedget_organization_audit_logs
    • First observedget_organization_evaluation
    • First observedget_organization_evaluation_metrics
    • First observedget_organization_evaluation_result
    • First observedget_role_credit_limit
    • First observedget_run_details
    • First observedget_run_history
    • First observedimport_browser_profile_cookies
    • First observedkill_flow
    • First observedlist_accounts
    • First observedlist_agent_mcp_servers
    • First observedlist_agent_versions
    • First observedlist_agents
    • First observedlist_artifacts
    • First observedlist_brain_files
    • First observedlist_brain_sources
    • First observedlist_browser_profiles
    • First observedlist_evaluations
    • First observedlist_flows
    • First observedlist_mcp_server_prompts
    • First observedlist_mcp_server_resources
    • First observedlist_mcp_server_tools
    • First observedlist_mcp_servers
    • First observedlist_models
    • First observedlist_organization_evaluation_results
    • First observedlist_organization_evaluations
    • First observedlist_organizations
    • First observedlist_queued_messages
    • First observedlist_role_credit_limits
    • First observedlist_sessions
    • First observedlist_skills
    • First observedlist_teams
    • First observedlist_workbooks
    • First observedmanage_permission_group_users
    • First observedmanage_project_users
    • First observedqueue_session_message
    • First observedread_mcp_server_resource
    • First observedrename_session
    • First observedresolve_session_approvals
    • First observedretrieve_agent
    • First observedretrieve_agent_version
    • First observedretrieve_evaluation
    • First observedretrieve_mcp_server
    • First observedretrieve_session
    • First observedroute_model
    • First observedrun_evaluations
    • First observedrun_organization_evaluation
    • First observedsearch_brain
    • First observedsend_message
    • First observedsend_queued_message
    • First observedset_organization_evaluation_targets
    • First observedset_role_credit_limit
    • First observedstart_flow
    • First observedupdate_agent
    • First observedupdate_agent_skills
    • First observedupdate_evaluation_config
    • First observedupdate_organization_evaluation
    • First observedupdate_queued_message
    • First observedupdate_skill
    • First observedupload_brain_files
    • First observedupload_file
    • First observedupload_files
    • First observedupload_session_file

TDQS

B3.1/5.0

Scored across 92 tools

Disambiguation2/5

Many tools have overlapping or indistinguishable purposes across the 92-tool surface, such as list_flows vs list_workbooks vs get_run_history, or upload_file vs upload_files vs upload_session_file vs upload_brain_files. with such a large flat set, boundaries between agent, session, flow, evaluation, brain, and skill tools are unclear without deep cross-referencing.

Naming Consistency4/5

Most tools follow a consistent verb_noun snake_case pattern (list_agents, create_skill, delete_brain_source), with only minor deviations like route_model and call_mcp_tools. The convention is largely predictable.

Tool Count2/5

92 tools is excessive for a single MCP server and likely buries the core operations. Many are thin variants (get_evaluation_metrics vs get_organization_evaluation_metrics) that could be consolidated.

Completeness4/5

Coverage is broad and deep across flows, agents, sessions, evaluations, brain, skills, and MCP, with both read and write operations. minor CRUD gaps exist but the surface is mostly complete.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to interact with Databricks workspaces programmatically, providing comprehensive tools for cluster management, notebook operations, job orchestration, Unity Catalog data governance, user management, permissions control, and FinOps cost analytics.
    492 npm
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to manage n8n workflow automation instances through tools for workflow CRUD operations, execution monitoring, and webhook triggering. It facilitates programmatic interaction with n8n instances via the n8n API with AI-optimized descriptions and error handling.
    62
    MIT
  • F
    license
    B
    quality
    F
    maintenance
    Enables AI assistants to manage GoClaw AI gateway infrastructure through a comprehensive suite of 66 tools for managing agents, sessions, and configurations. It provides real-time gateway context and guided workflows with enterprise-grade security features like audit logging and secret scrubbing.
    66
    10
    -