Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
GUMLOOP_API_KEYNoPrivate Bearer credential; alternative to token file.
GUMLOOP_TEAM_IDNoSingle-account default team; legacy project_id.
GUMLOOP_USER_IDNoSingle-account default user identity.
GUMLOOP_ACCOUNTSNoPrivate named JSON profiles; takes precedence over single-account settings.
GUMLOOP_AUDIT_LOGNoOptional private local guard log.
GUMLOOP_READ_ONLYNo1/true hides/refuses confirmed operations.
GUMLOOP_TOKEN_FILENoRegular token-only file <=64 KB; overrides key.
GUMLOOP_MAX_RETRIESNo0-5; default 2; short explicit GET 429 only.2
GUMLOOP_DEFAULT_ACCOUNTNoExact label, otherwise first entry.
GUMLOOP_ALLOW_DESTRUCTIVENo0/false blocks confirmed operations.
GUMLOOP_REQUEST_TIMEOUT_MSNo100-300000; default 30000.30000
GUMLOOP_MIN_REQUEST_INTERVAL_MSNo0-10000; default 150; per account/process.150

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": true
}

Tools

Functions exposed to the LLM to take actions

NameDescription
start_flowA

This endpoint is used to trigger a flow run via API Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

kill_flowA

This endpoint is used to kill a flow run and all its subflow runs. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

get_run_detailsA

This endpoint can be used to poll for completion and retrieve final flow outputs. Output steps must be used to retrieve outputs. Read operation.

list_workbooksC

List workbooks and their saved flows Read operation.

list_flowsC

List saved flows Read operation.

get_input_schemaC

Retrieve input schema Read operation.

get_run_historyB

This endpoint retrieves the run history for automations, either by workbook or saved item. Returns the 10 most recent runs. Read operation.

download_fileC

Download file Read operation.

download_filesC

Download multiple files Read operation.

upload_fileB

Upload file Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

upload_filesA

Upload multiple files Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

get_organization_audit_logsB

This endpoint retrieves audit logs for all users in an organization for a specified time period. Read operation.

manage_project_usersB

This endpoint allows organization administrators to add or remove users from a workspace. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

manage_permission_group_usersA

This endpoint allows organization administrators to add or remove users from a custom role (formerly "permission group"). Adding a user to a role does not remove them from any other role they belong to. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

list_role_credit_limitsA

This endpoint lists every active custom role in an organization together with its monthly credit limit, so external systems can manage credit limits programmatically. A monthly_credit_limit of null means the role sets no limit of its own. The limit applies to each member of the role individually; when a user belongs to multiple roles, the highest limit across their roles wins. Read operation.

get_role_credit_limitA

This endpoint returns the monthly credit limit of one custom role. A monthly_credit_limit of null means the role sets no limit of its own. Read operation.

set_role_credit_limitA

This endpoint sets or clears the monthly credit limit of a custom role. The limit applies to each member of the role individually and takes effect immediately: member allowances are recalculated while preserving credits already used in the current billing cycle. Send "monthly_credit_limit": null to clear the role-level limit so members revert to the organization default. When a user belongs to multiple roles, the highest limit across their roles wins. Requests that do not change the stored value are no-ops. Changes are recorded in the organization audit trail. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

export_dataA

This endpoint allows enterprise organization administrators to create and initiate a comprehensive data export for their organization or specific workspaces.

The export supports six data types:

  • Workflow data (data_type: "workflows"): Includes workflow runs, workbook details, user information, and other organizational data.

  • Agent data (data_type: "agents"): Includes agent configurations, metadata, tools, and creator information.

  • Agent interaction data (data_type: "agent_interactions"): Includes agent interaction data with timestamps, credit costs, trigger types, and message counts.

  • Credit log data (data_type: "credit_logs"): Includes credit transaction history with charges, balances, categories, and user attribution.

  • Interaction evaluation data (data_type: "interaction_evaluations"): Includes one row per completed evaluation of a chat, with its grade, call outcome, sentiment, and the model that graded it.

  • Gumstack data (data_type: "gumstack"): Includes Gumstack MCP tool call activity with timestamps, statuses, and latency.

The available export_fields depend on the selected data_type. See the field descriptions below for details.

Scoping requirement: For non-credit-log exports, at least one scoping parameter must be provided: workspace_ids, include_all_workspaces, include_personal_workspaces, or entity_ids. Requests that omit all scoping parameters will receive a 400 error.

Note: Credit log exports work differently from workflow and agent exports. When data_type is "credit_logs", the following parameters are not applicable and will be ignored: export_level, workspace_ids, include_all_workspaces, include_personal_workspaces, and entity_ids. Credit log exports are always scoped to the entire organization. Use category_filter to filter by credit log category. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

get_export_statusA

This endpoint retrieves the status of a data export job and optionally downloads the export file (as CSV) if the export has completed successfully.

Use the data_export_id returned by the Export data endpoint to check progress. Read operation.

list_agentsA

List agents the caller has access to. Filter by team, search by name, or narrow to agents that use a specific tool or trigger. Results can be sorted with sort_order and paginated by sending page_size and/or cursor. Read operation.

create_agentA

Create a new agent. The authenticated caller must have permission to create agents on the target team. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

retrieve_agentA

Retrieve a single agent by ID.

In addition to regular agent IDs, agent_id accepts the reserved aliases gumball (your personal Gumball agent) and analytics (your analytics agent) on all agent-scoped endpoints. The alias resolves to your own copy of the platform agent, creating it on first use, and responses report the alias back as the agent's id. Read operation.

update_agentA

Update an existing agent. Only fields included in the request body are changed; omitted fields are left untouched.

This endpoint edits document fields only. To attach or detach skills, use PATCH /agents/{agent_id}/skills; to manage MCP servers, use the agent MCP server endpoints.

is_active: false is not a pause switch. It retires the agent: the agent disappears from GET /agents, and both GET and PATCH /agents/{agent_id} return 404 afterwards, so you cannot set it back to true through the API. To stop an agent from running on its own while keeping it fully reachable, disable its triggers instead. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

create_chat_completionB

OpenAI-compatible chat completions endpoint, multiplexed across every model Gumloop supports (Anthropic, OpenAI, Google Gemini, OpenRouter routes). Set stream: true for Server-Sent Events, or omit it for a unary JSON response. Image-generation models (gpt-image-*, gemini-*-image-preview, dall-e-*) are dispatched automatically when modalities includes "image" and yield image attachments on choices[0].message.images.

Streaming host

Chat completions live on the streaming host. Send all requests — unary or streaming — to:

POST https://ws.gumloop.com/api/v1/chat/completions

api.gumloop.com does not serve this endpoint; the Python SDK routes there automatically.

Tool calls, images, and tool_choice

Send messages in the OpenAI shape and Gumloop translates them for the model's provider (Anthropic, OpenAI, and Google Gemini). Models served through OpenRouter and other OpenAI-compatible providers receive the messages as sent.

  • Tool-result turns: after the model replies with finish_reason: "tool_calls", append its assistant message (with tool_calls) and one {"role": "tool", "tool_call_id": ..., "content": ...} message per call, then send the conversation again. Every tool call needs a matching tool message, and every tool message must match a tool call in an earlier assistant message.

  • Images: user messages accept image_url content parts alongside text parts. The URL can be an http(s) URL or a base64 data URL (data:image/png;base64,...). Images must be JPEG, PNG, GIF, or WebP and at most 20 MB. Redirects are not followed when downloading an image.

  • tool_choice: "auto" (the default when tools are sent), "none", "required", or {"type": "function", "function": {"name": "..."}} to force one tool.

  • developer messages are treated like system messages.

{
  "model": "claude-sonnet-4-5",
  "tools": [{"type": "function", "function": {"name": "get_weather", "parameters": {"type": "object", "properties": {"city": {"type": "string"}}}}}],
  "messages": [
    {"role": "user", "content": [
      {"type": "text", "text": "What's the weather where this photo was taken?"},
      {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}
    ]},
    {"role": "assistant", "content": null, "tool_calls": [
      {"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": "{\"city\": \"Ottawa\"}"}}
    ]},
    {"role": "tool", "tool_call_id": "call_1", "content": "12°C and sunny"}
  ]
}

A request that can't be translated returns 400 invalid_request with param set to the field at fault (for example messages[3].tool_call_id). When the provider itself rejects the request (HTTP 400, 404, 413, or 422), the error message relays the provider's reason, prefixed with The provider rejected the request:.

Billing

Each completion charges the caller's credit balance based on token usage (with cache-token semantics per provider) plus a flat 30-credit fee for image-gen calls. Users who configure their own provider API key get a 50% discount. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

list_modelsB

List the LLMs and preset model chains available to the caller, grouped for display in a model picker. Read operation.

route_modelA

Ask Gumloop Chew — the model router behind Auto — which model it would use for a given message, and why. Chew is a router, not a model: router is always gumloop-chew, and the concrete model it selected is route.model.

This endpoint returns a decision only. It does not run the selected model or its fallbacks. The routing judgement consumes credits and is bounded by the caller's model access.

Omit models to use the deduplicated union of Chew's lane chains, not every model the caller may use. Restricted candidates can appear with status: "restricted"; they are never selected or included in fallback_models.

Team scope requires actual team membership, even within the same organization. Personal API keys and OAuth are supported; this is not team-key-only. Read operation.

list_agent_versionsA

List the immutable versions of an agent, newest first. Each entry is a point-in-time snapshot of the agent's configuration.

Requires configuration access on the agent — callers limited to using the agent (no configuration access) get a 403. Read operation.

retrieve_agent_versionA

Retrieve one immutable agent version: its full configuration (composition) plus the structured changes relative to the version before it. Use it to export an agent's configuration or to audit what changed between versions.

changes is null for the first version of an agent, since there is no predecessor to diff against. Versions created before attachment snapshots were recorded report composition.complete: false (and changes.attachment_changes_complete: false); their skill_ids and knowledge_sources are null rather than empty. Skill file contents are never included, and this endpoint is read-only — it cannot restore or deploy a version.

Requires configuration access on the agent — callers limited to using the agent (no configuration access) get a 403. Read operation.

update_agent_skillsA

Attach and/or detach skills on an agent using deltas. This is not a replace-list: skills you don't mention are left untouched.

  • The operation is idempotent. Re-attaching a skill that's already attached (or detaching one that isn't) is reported under already_attached / already_detached rather than failing.

  • A skill ID may not appear in both attach and detach.

  • Up to 100 unique skill IDs total (attach + detach) per request.

  • Attaching requires INVOKE permission on the skill. Detaching is permissive so stale attachments can always be removed. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

list_agent_mcp_serversA

List the MCP servers (connectors) attached to an agent. Sensitive fields such as secret_id and mcp_server_url are scrubbed from the response. Read operation.

attach_agent_mcp_serverA

Attach an MCP server (connector) to an agent, or update its configuration if it's already attached (upsert).

The server_id is validated against the caller's MCP catalog. Catalog identity fields (type, server_id, secret_id, mcp_server_url) always come from the catalog and cannot be spoofed via the request body — the body carries only free-form connector configuration (e.g. approval mode, tool restrictions); any identity keys in it are ignored.

Attach may succeed before OAuth is completed; auth_status reflects the catalog's authentication state. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

detach_agent_mcp_serverA

Detach an MCP server (connector) from an agent. This is idempotent — detaching a server that isn't attached returns detached: false rather than an error. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

list_sessionsA

List sessions for an agent with cursor-based pagination, optional filtering, and search. Read operation.

create_sessionA

Create a new session for an agent. When input is provided, the message is enqueued and the agent begins processing — the response returns 202 with the session in processing or queued state. When input is omitted, an idle session stub is created and the response returns 201.

agent_id also accepts the reserved aliases gumball and analytics, which resolve to your personal Gumball and analytics agents (created on first use).

Streaming the response

api.gumloop.com only serves the non-streaming response above. To stream agent output as it's produced, send the same request body (with stream: true) to the streaming host instead:

POST https://ws.gumloop.com/api/v1/agents/{agent_id}/sessions

The response is text/event-stream (Server-Sent Events). With the Python SDK, client.sessions.stream(agent_id, input="...") routes to ws.gumloop.com automatically and yields parsed StreamEvent objects.

If you send stream: true to api.gumloop.com by mistake, the response is a 400 whose body contains the correct streaming host so you can retry against it. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

retrieve_sessionA

Retrieve a session by ID, including its messages, current state, agent metadata, and participants. Read operation.

rename_sessionA

Rename a session. name is the only mutable field; it is trimmed and must be between 1 and 256 characters after trimming.

Returns the full session, in the same shape as Retrieve session. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

send_messageA

Append a user message to an existing session and resume the agent. The session must be idle, completed, failed, or approval_required; sending to a session that is processing or queued returns 409 interaction_not_in_terminal_state. To hand the agent a message while it is still busy, use the message queue instead.

Files uploaded via Upload session file can be attached to the message with attachments.

Sessions waiting on an approval

A session that stopped to ask you something is approval_required, and you have two ways to move it forward:

  • Answer the ask. Send the pending asks' responses to Resolve approvals. Use this to approve or reject a tool call, or to answer an Ask Question the agent raised. This endpoint rejects approval_responses with a 400.

  • Send a follow-up instead. Post a normal message here. It is appended to the session transcript and starts a new turn, leaving the pending ask unanswered. Use this when the answer no longer matters — for example to redirect the agent or drop the request it was asking about.

See Human in the Loop for how agents pause for approvals and questions.

Streaming the response

api.gumloop.com only serves the non-streaming response above. To stream agent output as it's produced, send the same request body (with stream: true) to the streaming host instead:

POST https://ws.gumloop.com/api/v1/sessions/{session_id}/messages

The response is text/event-stream (Server-Sent Events). With the Python SDK, client.sessions.stream_message(session_id, input="...") routes to ws.gumloop.com automatically and yields parsed StreamEvent objects.

If you send stream: true to api.gumloop.com by mistake, the response is a 400 whose body contains the correct streaming host so you can retry against it. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

cancel_sessionA

Cancel an in-progress session. If the session is currently processing or queued, any running stream is aborted and the session is transitioned to failed. If the session is already completed or failed, its current state is returned unchanged.

The response carries a session envelope but only id, agent_id, and state are populated. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

upload_session_fileA

Upload a file into a session's input namespace so it can be attached to a message. The response returns the stored path — pass it as file_name in the attachments array when sending a message on the same session.

Files are base64 encoded in the request body and limited to 200MB (decoded). Uploaded files are scoped to the session they were uploaded to and cannot be attached to messages on other sessions. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

resolve_session_approvalsA

Answer pending asks on a session that is paused in the approval_required state — tool approvals, human input requests, and checkpoints.

List the pending asks with Retrieve session: each entry in pending_approvals carries the action_request_id to answer, and human_input asks include the questions to fill in via response.values. Resolutions are processed in order; the agent resumes once the pending asks are answered. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

list_queued_messagesA

List the messages waiting in a session's queue, in the order they will be sent. Queued messages are drained automatically when the agent finishes its current turn. Read operation.

queue_session_messageA

Add a message to a session's queue instead of interrupting the agent. Queued messages are sent automatically, in order, when the agent finishes its current turn. A session's queue holds at most 20 messages.

To interrupt the current turn and send a queued message immediately, use Send queued message now. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

update_queued_messageA

Replace the content of a message that is still waiting in the session's queue. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

delete_queued_messageA

Remove a message from the session's queue before it is sent. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

send_queued_messageA

Send a queued message immediately instead of waiting for the agent to finish its current turn. Any in-progress run is aborted, the queued message is appended to the transcript, and the agent starts processing it.

The response is the same envelope as Send message. Queued messages cannot be sent this way on incognito sessions, and a message that is currently being edited must have its edit finished or cancelled first. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

list_mcp_serversB

Return the catalog of MCP servers visible to the caller — Gumloop-hosted (gumcp_server), user-deployed Gumstack (gumstack_server), and custom (mcp_server) — along with each server's connection state. Read operation.

retrieve_mcp_serverB

Return a single MCP server. The response populates allowed_tool_call_ids with the tool call IDs the caller is permitted to invoke on this server. Read operation.

list_mcp_server_toolsA

Return the tools exposed by an MCP server. When the server is not in connected state, tools is empty and gumloop_auth_url is returned so the caller can prompt the user to authenticate. Read operation.

list_mcp_server_resourcesA

Return the resources an MCP server exposes, fetched live from the server. When the server is not connected, resources is empty and gumloop_auth_url is returned so the caller can prompt the user to authenticate. Read operation.

read_mcp_server_resourceA

Read one resource by uri. Each content item is either text or a base64 blob, never both. Read operation.

list_mcp_server_promptsB

Return the prompt templates an MCP server exposes, fetched live from the server. When the server is not connected, prompts is empty and gumloop_auth_url is returned. Read operation.

get_mcp_server_promptB

Render one prompt template with arguments and return its messages. Read operation.

call_mcp_toolsA

Execute a batch of 1–5 MCP tool calls. Calls run concurrently and each result reports its own status. When Gumloop accepts the request, MCP execution failures such as target server authentication, policy blocks, invalid tools, upstream HTTP errors, and connection failures are returned in results[*].status and results[*].error. Top-level 4xx responses are reserved for Gumloop request, authentication, and permission failures. 200 covers homogeneous execution outcomes (all calls succeeded or all calls failed); mixed success/failure batches return 207. If you previously treated non-2xx HTTP statuses as MCP execution failures, update your integration to inspect each result's status and error. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

search_brainA

Run a hybrid (semantic + keyword) search across the knowledge sources indexed in your Company Brain and return the most relevant, ranked snippets with citations.

Results are scoped to what the authenticated user can see: personal sources, plus any team and organization sources shared with them. Requires the Brain feature, which is available on the Pro and Enterprise plans. Each search consumes Gumloop credits. Explicit confirmation is required for this credit-consuming Brain search.

list_brain_sourcesA

List the Company Brain sources the authenticated user can see: personal sources, plus team and organization sources shared with them. Every source type is listed, including ones connected in the app such as Notion or Google Drive. Read operation.

create_brain_sourceA

Create a file-upload source. Only direct_file_uploads sources can be created through the API; connected sources such as Notion or Google Drive are set up in the Gumloop app because they need an account connection.

By default the source is active and indexes (and bills) each file as soon as it is uploaded. Set require_approval to true to create it as a draft instead: uploads then run a credit estimate, the source owner is notified, and nothing is indexed until Approve source is called. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

get_brain_sourceB

Fetch one source the authenticated user can see. Read operation.

delete_brain_sourceA

Delete a source, every file in it, and everything it contributed to search. This cannot be undone. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

list_brain_filesA

List the files in a file-upload source with their indexing status. Each file carries the sha256 of its bytes, so a client can compare a local folder against the source and upload only what changed. Read operation.

upload_brain_filesA

Upload up to 25 files as multipart/form-data parts named files. Accepted types are PDF, Word, PowerPoint, Excel, and text formats (.txt, .md, .html, .csv, .rtf), each up to 25 MB.

Indexing starts on its own after the upload: an active source indexes and bills immediately, a draft source runs a credit estimate instead (see Retrieve estimate). Poll List files until each file's status is indexed.

Files the upload policy refuses (unsupported type, too large, empty) are returned in rejected with a 201; the request is a 400 no_files_accepted only when every file was refused. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

delete_brain_fileB

Remove a file from the source and from search. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

get_brain_source_estimateA

The latest credit estimate for a source created with require_approval. estimate is null until the first upload has produced a run; poll until estimate.status is paused_for_approval, then call Approve source. estimated_credits is rounded up to the nearest 5 and is an estimate, not a quote. Read operation.

approve_brain_sourceA

Approve a draft source. It becomes active, the paused estimate run resumes as a real indexing run, and credits are charged. Later uploads index without another approval. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

list_skillsA

List skills the caller has access to. Filter by team, search by name, or narrow to a specific creator, related MCP server, or agent. Read operation.

create_skillA

Upload a skill package and create a new skill. The package must include a SKILL.md with name and description frontmatter; uploads may be a single .md file (stored as SKILL.md), or a .zip / .skill archive containing SKILL.md at its root. The initial version is created automatically. Maximum upload size is 10 MB. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

update_skillA

Replace a skill's files with a new upload. Reparses SKILL.md to update the skill's name, description, and metadata, and creates a new version. Maximum upload size is 10 MB. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

delete_skillA

Permanently delete a skill. This is a soft-delete — the skill will no longer appear in listings or be usable by agents. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

download_skill_fileA

Generate a signed URL to download a skill's contents as a .skill archive (ZIP). When version_id is provided, returns that exact version; otherwise returns the current draft. Read operation.

list_artifactsA

List artifacts (files) produced by an agent. Optionally scope to a specific session, search by filename, sort, and paginate. Deleted files are excluded from the results. Read operation.

download_artifact_fileA

Returns a signed download URL for an artifact, plus its filename, media type, and size. Follow download_url to fetch the file bytes. Read operation.

list_browser_profilesA

List the browser profiles you own, or a team's profiles with team_id. A browser profile holds the sign-ins an agent's Browser ability uses. Cookie values are never returned; each profile lists its sites and cookie counts. Read operation.

import_browser_profile_cookiesA

Add sign-in cookies to a browser profile. This is what gumloop browser import-logins calls. Send the cookies in Chrome extension (chrome.cookies.Cookie) or Chrome DevTools Protocol Cookie shape. With url, only that site's cookies are kept and the import replaces that site; without it, every site in the payload is imported. Cookies are encrypted with the profile's key before storage and are never returned by any endpoint.

Use default as the profile_id to import into the owner's default profile, creating it if needed. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

list_teamsC

List teams the authenticated caller belongs to. Read operation.

list_evaluationsA

Returns a cursor-paginated list of evaluation results for a specific agent, newest first. Only completed and failed evaluations are returned unless status selects another state.

Each evaluation includes the grade, criteria pass/fail results, extracted data points, applied tags, and sentiment analysis. Read operation.

run_evaluationsA

Grades up to 200 of the agent's finished sessions with its own evaluation configuration. Grading is asynchronous: each accepted session gets a result with status: queued; poll it with GET /agents/{agent_id}/evaluations/{evaluation_id} until it is completed or failed. A new result replaces the previous result for that session.

Sessions are skipped, not rejected, when they are unfinished, incognito, or not owned by this agent (ineligible), or already have a queued or running result (in_flight, with the existing result_id). The caller is charged one credit per queued session. Set dry_run: true to see the cost and skips without queuing anything.

Requires edit access on the agent and a plan with evaluations enabled. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

get_evaluation_metricsA

Returns aggregated grade and tag counts for an agent's evaluations over a time window. Useful for dashboards and reporting on agent quality trends. Read operation.

retrieve_evaluationA

Retrieve a single evaluation result by ID. The evaluation must belong to the specified agent. Read operation.

get_evaluation_configB

Retrieve the current evaluation configuration for an agent, including criteria, tags, data points, and sentiment settings. Read operation.

update_evaluation_configA

Partially update the evaluation configuration for an agent. Omitted fields keep their current value. Provided list fields (criteria, tags, data_points) replace that list entirely.

Requires Pro tier or above. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

list_organizationsA

Returns the organization the authenticated user belongs to. Use its id as organization_id on the evaluation endpoints. Read operation.

get_evaluation_optionsA

Allowed values for evaluation fields and filters — session types, criterion types and priorities, data point types, frequencies, grades, statuses, target types, skip reasons — plus size limits. Use these instead of hardcoding enums. Read operation.

list_organization_evaluationsA

Cursor-paginated list of an organization's evaluations with their targets, coverage, and result rollups. Requires the organization:manage_evaluations permission (Enterprise plan). Read operation.

create_organization_evaluationA

Creates an organization evaluation. A new evaluation has no targets, so it cannot start enabled: set targets with PUT /evaluations/{evaluation_id}/targets, then enable it with PATCH /evaluations/{evaluation_id}.

Rubric values are validated strictly: an unknown frequency, criterion priority, type, data point data_type, or session type, a criterion without name and prompt, or a duplicate tag name returns 400 invalid_request with the offending paths in error.details.fields. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

get_organization_evaluationB

Returns one evaluation with its rubric, targets, current coverage, and result rollup. Read operation.

update_organization_evaluationA

Partial update. Only the fields you send change. config is merged field by field; a list you send (criteria, tags, data_points) replaces that list wholesale. description: null clears the description.

Setting enabled: true requires at least one criterion, tag, or data point (400 organization_evaluation_empty_rubric) and at least one covered agent (400 organization_evaluation_no_targets). Emptying the rubric of an enabled evaluation pauses it. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

delete_organization_evaluationA

Deletes the evaluation. It stops running and disappears from lists; results it already produced stay attached to their sessions. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

set_organization_evaluation_targetsA

Replaces the full set of targets — who the evaluation grades. Targets expand to agents live: organization covers every agent in the organization, team every agent a team owns, user a member's personal agents, agent one agent. Removing the last target pauses an enabled evaluation; enabled in the response reflects that. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

run_organization_evaluationA

Grades up to 200 existing sessions with this evaluation. Grading is asynchronous: each accepted session gets a result with status: queued; poll it with GET /evaluations/{evaluation_id}/results/{result_id} until it is completed or failed.

Sessions are skipped, not rejected, when they are not completed sessions of an agent the evaluation covers (ineligible) or already have a queued or running result for this evaluation (in_flight, with the existing result_id). The caller is charged one credit per queued session. Set dry_run: true to see the cost and skips without queuing anything. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

list_organization_evaluation_resultsA

Cursor-paginated results for one evaluation across every agent it grades, newest first. Each session appears once with its latest result; queued and in-progress results are included so a run can be followed to completion. Read operation.

get_organization_evaluation_resultA

One result, including per-criterion outcomes, extracted data points, and applied tags. Poll this after POST /evaluations/{evaluation_id}/run until status is completed or failed. Read operation.

get_organization_evaluation_metricsB

Grade counts for one evaluation over a trailing window (default 30 days, 1–365). Read operation.

list_accountsA

List private account labels, default selection and configured token method. No credentials, token paths or account content; no network request.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

B3.1/5.0

Scored across 92 tools

Disambiguation2/5

Many tools have overlapping or indistinguishable purposes across the 92-tool surface, such as list_flows vs list_workbooks vs get_run_history, or upload_file vs upload_files vs upload_session_file vs upload_brain_files. with such a large flat set, boundaries between agent, session, flow, evaluation, brain, and skill tools are unclear without deep cross-referencing.

Naming Consistency4/5

Most tools follow a consistent verb_noun snake_case pattern (list_agents, create_skill, delete_brain_source), with only minor deviations like route_model and call_mcp_tools. The convention is largely predictable.

Tool Count2/5

92 tools is excessive for a single MCP server and likely buries the core operations. Many are thin variants (get_evaluation_metrics vs get_organization_evaluation_metrics) that could be consolidated.

Completeness4/5

Coverage is broad and deep across flows, agents, sessions, evaluations, brain, skills, and MCP, with both read and write operations. minor CRUD gaps exist but the surface is mostly complete.

Maintenance

ActivityMaintained
ResponsivenessNo issues