Skip to main content
Glama

Set evaluation targets

set_organization_evaluation_targets
Destructive

Replaces all targets for an organization evaluation, defining which agents, teams, users, or organization it grades. Requires explicit confirmation for this account operation.

Instructions

Replaces the full set of targets — who the evaluation grades. Targets expand to agents live: organization covers every agent in the organization, team every agent a team owns, user a member's personal agents, agent one agent. Removing the last target pauses an enabled evaluation; enabled in the response reflects that. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
targetsNo
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.
evaluation_idYesID of the organization evaluation.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed5 schema fields changedv3.0.0
    • changedInput schema / properties / confirm / description
      Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
    • addedInput schema / properties / payload / $defs / EvaluationTarget / $ref
      Added value: +"#/$defs/EvaluationTarget"
    • removedInput schema / properties / payload / $defs / EvaluationTarget / properties
      Removed value: -{
      -  "id": {
      -    "description": "Team, user, or agent ID. Omitted for `organization`; responses return the organization ID.",
      -    "type": "string"
      -  },
      -  "type": {
      -    "description": "What the target expands to. `user` covers a member's personal agents.",
      -    "enum": [
      -      "organization",
      -      "team",
      -      "user",
      -      "agent"
      -    ],
      -    "type": "string"
      -  }
      -}
    • removedInput schema / properties / payload / $defs / EvaluationTarget / required
      Removed value: -[
      -  "type"
      -]
    • removedInput schema / properties / payload / $defs / EvaluationTarget / type
      Removed value: -"object"
  2. First observedv2.0.1

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing that removal of the last target pauses an enabled evaluation, that the response's `enabled` field reflects it, and that runs can spend credits or trigger downstream actions. These are non-obvious side effects an agent must know before calling a destructive, non-idempotent mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and follows with the target-expansion table and the confirmation warning; every sentence carries weight. Slightly dense in the target enumeration, but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation tool with no output schema and nested objects, the description covers replacement semantics, side effects, confirmation requirements, and a hint about the `enabled` return field. It does not explain `account`, `payload`, or `payload_file` behavior, leaving minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the schema already documents most parameters (baseline 3). The description adds real value by spelling out what each target `type` expands to (`organization`/`team`/`user`/`agent`) and how removal affects evaluation state, enriching the enum semantics beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Replaces the full set of targets') and clarifies the target-expansion semantics, so an agent knows exactly what is being overwritten. It implicitly distinguishes itself from sibling update tools by emphasizing full replacement, but never names a sibling, which keeps it at a 4.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating context: explicit confirmation is required for this exact operation and unknown outcomes must not be auto-resubmitted. It does not, however, contrast itself against alternatives like update_organization_evaluation or update_evaluation_config, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools