Featureflip
OfficialSummary: Manage Featureflip feature flags end-to-end from an AI coding assistant — flags, targeting, variations, segments, webhooks, and cleanup workflows (but not flag evaluation).
Discover:
list_projects,list_environments, andlist_flags(filter by search, tag, type, owner, archived), plusget_flag,list_segments, andget_segment.Manage flags:
create_flag(Boolean/String/Number/Json, created disabled in every environment),update_flagmetadata,delete_flag,archive_flag/restore_flag, andtoggle_flagto enable/disable in one environment.Targeting & config: read and fully replace targeting rules with
get_targeting/update_targeting(attributes, operators, segment references, percentage rollouts), set per-environment defaults, strategy, and prerequisites viaupdate_flag_environment_config, and inspect cross-environment state withflag_status.Variations:
manage_variationto add, update, or remove a flag's variations.Governance:
set_flag_expiryandset_flag_ownerto date and assign responsibility for cleanup notices (owner assignment needs Pro or above).Cleanup workflows:
find_stale_flagsandlist_removal_candidatessurface dead/stale/expired flags withblockedBydependents so you remove them dependents-first.Wrap new features:
wrap_featurecreates a Boolean flag and returns the SDK snippet (js, node, browser, react, python, csharp, java, go, php, ruby, swift) — you apply the edit yourself.Webhooks (Admin token):
list_webhooks,list_webhook_deliveries,manage_webhook(create/update/delete, rotate or retire signing secrets), anddeliver_webhook(send a test event or redeliver).Limits: flag evaluation is not exposed (use the language SDKs), and every mutation is audit-logged and attributed to the token.
@featureflip/mcp
MCP (Model Context Protocol) server for Featureflip — manage feature flags from AI coding assistants (Claude Code, Cursor, Copilot, Cline) and autonomous agents.
Setup
Create an API token in the Featureflip dashboard:
Personal token (
ffp_...): Settings → API Tokens — acts as you, for interactive editor use.Service token (
ffs_...): Organization Settings → Service Tokens — scoped machine identity for CI/agents.
Add the server to your MCP client:
Claude Code
claude mcp add featureflip -e FEATUREFLIP_TOKEN=ffp_your_token -- npx -y @featureflip/mcpCursor / Cline / generic JSON config
{
"mcpServers": {
"featureflip": {
"command": "npx",
"args": ["-y", "@featureflip/mcp"],
"env": { "FEATUREFLIP_TOKEN": "ffp_your_token" }
}
}
}Related MCP server: Featureflow MCP Server
Configuration
Env var | Required | Default | Purpose |
| yes | — |
|
| no |
| API base URL |
| no | auto | Org slug (needed only for multi-org personal tokens) |
Tools
CRUD: list_projects, list_environments, list_flags, get_flag, create_flag, update_flag,
delete_flag, archive_flag, restore_flag, set_flag_expiry, set_flag_owner, toggle_flag,
update_flag_environment_config, get_targeting, update_targeting, manage_variation, list_segments,
get_segment
Webhooks (Admin token required): list_webhooks, list_webhook_deliveries, manage_webhook, deliver_webhook
Workflows: flag_status (cross-environment view), find_stale_flags (cleanup candidates),
list_removal_candidates (flags safe to remove, with the prerequisite dependents to remove first),
wrap_feature (create flag + get the SDK snippet for your language)
Full reference: https://featureflip.io/docs/integrations/mcp/
Notes
Flag evaluation is not exposed over MCP — use the language SDKs in application code.
Every mutation is audit-logged and attributed to the token.
License
Apache-2.0
Available Tools
26 toolsarchive_flagArchive feature flagADestructive
Archive a flag (soft-hide, evaluation stops serving it). Reversible with restore_flag. Archive dependents first: refused with FLAG_HAS_DEPENDENTS while another live flag lists this one as a prerequisite, so archive those flags (blockedBy in find_stale_flags / list_removal_candidates) first. Refused with FLAG_RECENTLY_EVALUATED while live traffic is still evaluating the flag, because archiving makes every caller fall back to its own hardcoded default — normally that means the code removal has merged but not deployed yet, and the refusal clears itself once it has.
| Name | Required | Description | Default |
|---|---|---|---|
| flag | Yes | ||
| force | No | Archive even though traffic is still evaluating the flag. Only for clients that can never be updated (old mobile app versions), where traffic will never drain on its own. | |
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the destructiveHint annotation by naming two refusal codes (FLAG_HAS_DEPENDENTS, FLAG_RECENTLY_EVALUATED), explaining the fallback-to-hardcoded-default consequence, and noting the refusal self-clears after deploy. Minor tension with destructiveHint=true, since the description stresses reversibility, but soft-hide + restore_flag makes this defensible rather than contradictory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: purpose first, then prerequisites, then refusal conditions. Every sentence carries operational information; it is long only because the failure modes genuinely require explanation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, refusal-prone mutation with no output schema, the description supplies the missing operational context: reversibility, ordering constraints, error codes, and the lifecycle reason for refusal. Nothing an agent needs to call this correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, and the description never mentions the 'force' parameter or its escape-hatch semantics, leaving that entirely to the schema. Required project/flag are self-evident, so the description neither compensates for nor worsens the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Archive a flag') and immediately disambiguates it from delete_flag by parenthetically defining it as a 'soft-hide, evaluation stops serving it.' The relationship to restore_flag is named, so the agent can place it in the sibling set without reading a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit preconditions and alternatives: archive dependents first, pointing to 'blockedBy in find_stale_flags / list_removal_candidates'. It also tells the agent when the tool will refuse and why, which is exactly the when/when-not guidance an agent needs before calling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_flagCreate feature flagA
Create a feature flag. type is one of Boolean|String|Number|Json. Boolean flags get true/false variations automatically; for other types pass initialVariations. The flag is created in every environment of the project (disabled).
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Unique flag key, e.g. checkout-v2 | |
| name | Yes | ||
| tags | No | ||
| type | Yes | ||
| project | Yes | Project key | |
| ownerEmail | No | Optional owner: the email of an active member (Pro plan and above, see set_flag_owner). Omitted, a personal access token makes its own user the owner and a service token leaves the flag unowned | |
| description | No | ||
| expiresAtUtc | No | Optional date (YYYY-MM-DD, end of that UTC day) or ISO-8601 timestamp the flag is expected to be removed by (see set_flag_expiry) | |
| idempotency_key | No | Idempotency-Key header for safe retries | |
| clientSideVisible | No | Expose to client-side SDKs (browser/mobile) | |
| initialVariations | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply destructiveHint=false, so the description carries most of the burden and delivers real value: the flag is created in every environment and starts disabled, and Boolean flags auto-generate true/false variations. It omits return shape, key-conflict behavior, and permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the verb+resource front-loaded, then type semantics, then environment behavior. The only minor redundancy is restating the enum values already in the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter create tool with no output schema and thin annotations, the description covers the key creation semantics but leaves gaps around conflict/duplicate handling, default idempotency behavior, and the created resource's identity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
At 55% schema coverage, the description adds genuine cross-parameter meaning: type drives whether initialVariations is needed, and it restates the type enum. However the remaining parameters (name, tags, description, and the rest) get no added semantic detail, so it only partly compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create a feature flag') that is immediately distinguishable from sibling mutations like update_flag, toggle_flag, and delete_flag. It does not explicitly name a sibling to contrast against, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the verb (this is the tool for making a new flag), and it gives one conditional hint ('for other types pass initialVariations'). It offers no explicit when-to-use/when-not guidance versus update_flag or any alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_flagDelete feature flagADestructive
PERMANENTLY delete a flag across all environments. Fails with FLAG_HAS_DEPENDENTS if other flags use it as a prerequisite. Prefer archive_flag unless the flag must be fully removed.
| Name | Required | Description | Default |
|---|---|---|---|
| flag | Yes | Flag key or id | |
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, and the description reinforces this by stating the deletion is permanent and applies across all environments. It adds the concrete error condition FLAG_HAS_DEPENDENTS, which is not present in the annotations. This goes beyond the structured metadata without contradicting it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short, information-dense sentences with no filler. The most critical fact (permanent deletion) is front-loaded, followed by the error condition and the preferred alternative. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive two-parameter operation, the description conveys the critical behavioral context: permanence, cross-environment scope, the dependency error, and the recommended alternative. The main gap is the undocumented 'project' parameter, but the overall operational context is otherwise sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: 'flag' is documented in the schema but 'project' is not. The description does not compensate by explaining what 'project' refers to or how it interacts with the deletion. The only added meaning is the implicit focus on the 'flag' parameter, leaving the project parameter semantics under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action as 'PERMANENTLY delete a flag across all environments,' which is specific about the verb, resource, and scope. It also differentiates itself from archive_flag by noting it should be preferred unless full removal is required. This unambiguously identifies the tool's purpose relative to its siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Prefer archive_flag unless the flag must be fully removed,' giving direct guidance on when to use this tool versus the archival alternative. It also mentions the failure condition FLAG_HAS_DEPENDENTS, which informs the user about a prerequisite that blocks deletion. This is strong usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deliver_webhookSend a test event or redeliverA
action=test queues one synthetic flag.toggled event to this subscription alone; action=redeliver (deliveryId from list_webhook_deliveries) re-sends a finished delivery, restarting its retry schedule. Delivery is asynchronous: read list_webhook_deliveries for the outcome. Both are refused for a disabled subscription, and redeliver for a delivery that is still Pending. Requires an Admin token. A 404 on every webhook call means webhooks are not enabled for the organization.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Subscription id (from list_webhooks) | |
| action | Yes | ||
| deliveryId | No | Required for redeliver |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (openWorldHint, destructiveHint), the description discloses asynchronous behavior, retry-schedule restart, admin-token requirement, and the 404 meaning when webhooks are not enabled. These are non-obvious behavioral details an agent must know before invoking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence carries a distinct, necessary fact: action semantics, asynchronous behavior, refusal conditions, auth/error signature. The most important decision (action choice) is front-loaded in the first sentence with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still covers invocation, side effects, error interpretation, and how to observe the result. The immediate response shape is not stated, but the explicit async note and pointer to list_webhook_deliveries make that gap acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description decodes the action enum fully and explains the role of deliveryId, including where it comes from and when it is required. It compensates for the schema's missing action description and adds context beyond the existing id and deliveryId schema fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies two concrete actions ('test' queues a synthetic flag.toggled event; 'redeliver' re-sends a finished delivery) and names the subscription scope. This is enough to distinguish it from read-only siblings like list_webhook_deliveries and list_webhooks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit selection criteria per action, tells the agent to take deliveryId from list_webhook_deliveries and read the same tool for the async outcome, and lists refusal conditions (disabled subscription, Pending redelivery). Clear when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_stale_flagsFind stale flagsARead-only
Find flags that look ready for code cleanup: not updated in N days AND either enabled in every environment (verify rollout is complete before removing — per-rule percentage ramps are not inspected) or disabled in every environment (dead — remove flag and code path). A flag whose expiry date (set_flag_expiry) has passed is always a candidate, however recently it was edited; if it is on in some environments and off in others its reason is past-expiry, meaning the owner has to decide which way to fold it. Every result carries expiresAtUtc, expired and owner (null when unowned). Pass owner to see one person's stale flags ("me" for your own) or "none" for the unowned ones. Every result also carries blockedBy: the live flags that still list it as a prerequisite, which have to be removed and archived before it ([] when none). blockedBy is null when the server did not classify the flag as a removal candidate, which means unknown, not unblocked. A flag that other live flags list as a prerequisite has to be removed LAST, dependents first: archiving it while a dependent still names it makes that dependent fail its prerequisite check and serve its off variation, so archive_flag refuses it with FLAG_HAS_DEPENDENTS. Checks at most 50 candidates per call, expired flags first.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Minimum age in days since last update | |
| owner | No | Only flags with this owner: an email, "me" for the token's own user, or "none" for unowned flags | |
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint already covering safety, the description adds substantial behavior: blockedBy null means unknown not unblocked, the 50-candidate cap with expired-first ordering, per-rule percentage ramps not being inspected, and the mandatory dependents-first removal order that triggers FLAG_HAS_DEPENDENTS. This is rich context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is long and dense, but front-loaded with the core definition and the sentences carry distinct information (candidate rules, return fields, dependency ordering, cap). A few clauses are run-on, but nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must explain returns, which it does thoroughly (expiresAtUtc, expired, owner with null-when-unowned, blockedBy including the null semantics). Combined with the cap, ordering, and dependency warnings, an agent has everything needed to call and act on results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so the schema carries much of the load, but the description adds meaning: owner accepts an email, 'me', or 'none' (schema already lists these but the description gives usage intent), and days is qualified by the rule that a past expiry date always makes a flag a candidate regardless of edit recency. Only project is left unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names a specific verb+resource (find stale flags ready for code cleanup) and precisely defines the scope: not updated in N days AND enabled everywhere or disabled everywhere. The preconditions distinguish it from a generic flag listing and it routes cleanup work implicitly away from archive_flag.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly explains the conditions that make a flag a candidate, the expiry-date override, and how to scope by owner ('me'/'none'). It also warns to verify rollout completeness before removing and explains the FLAG_HAS_DEPENDENTS refusal, but it does not explicitly name an alternative tool or a when-not-to-use case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flag_statusFlag status across environmentsBRead-only
Compact cross-environment view of one flag: enabled state, strategy, default variation, and prerequisites per environment.
| Name | Required | Description | Default |
|---|---|---|---|
| flag | Yes | ||
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes the safety profile, and the description adds that this is a compact view spanning environments with specific fields. This is useful but not a deep disclosure of behavior such as aggregation or omission of full targeting details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that names the scope first and then enumerates the returned fields. Every word earns its place with no repetition or padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read-only tool, the description covers the output fields and cross-environment behavior. It does not describe return format or errors, but the low complexity and readOnlyHint annotation make those omissions minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to compensate, but it only mentions 'one flag' and 'environments' indirectly. The 'project' parameter is never referenced, and no format or relationship details are given beyond the schema's bare names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific resource ('one flag') and a specific scope ('cross-environment'), and lists the fields shown. It is clear enough to distinguish from list_flags and get_targeting, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this tool over get_flag, list_flags, or update_flag_environment_config. The phrase 'compact cross-environment view' implies an overview use case, but no exclusions or alternative routing are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_flagGet feature flagARead-only
Get one feature flag with its variations and metadata, including its owner (null when unowned). Address by flag key or id.
| Name | Required | Description | Default |
|---|---|---|---|
| flag | Yes | Flag key or id | |
| project | Yes | Project key |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds real value beyond that by disclosing what is returned (variations, metadata) and an edge case (owner is null when unowned), which is meaningful given there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action front-loaded and the addressing rule second. The second sentence slightly overlaps the schema's flag description, which is the only minor waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return-value burden, and it does so partially by listing variations, metadata, and owner. A richer account of the response shape would be needed for a 5, but it is adequate for this simple read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented. "Address by flag key or id" restates the schema's own description of the flag parameter and adds no new syntax or format detail, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Get one feature flag") and scopes it to a single entity, which implicitly separates it from list_flags. It does not name a sibling alternative explicitly, so a 5 is not warranted.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The singular "one feature flag" implies the retrieval use case and the description notes the flag/pair addressing, but it never says when to prefer this over list_flags, get_targeting, or update_flag. Usage is only inferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_segmentGet user segmentARead-only
Get one user segment (its conditions, id, and metadata) by key or id.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes | Project key | |
| segment | Yes | Segment key or id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds the response contents (conditions, id, metadata), which is behavior beyond the schema and annotations, providing useful context for what the agent will receive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler. The main action and resource appear first, followed by the return fields and lookup method. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple getter with two well-described parameters, read-only annotation, and no output schema, the description is sufficient. It names the response fields, so an agent knows what to expect without needing additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both 'project' and 'segment' already described in the schema. The description's 'by key or id' merely restates the segment param description, adding no new parameter semantics. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get'), a specific resource ('one user segment'), and the return contents ('conditions, id, and metadata'). It clearly distinguishes from list_segments (plural vs singular) and get_flag (different resource).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when a single segment is needed by key or id, and the sibling list makes the alternative (list_segments) obvious. However, it does not explicitly state when not to use it or name the alternative, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_targetingGet targeting rulesARead-only
Get a flag's current targeting configuration (enabled flag + ordered rules) in one environment.
| Name | Required | Description | Default |
|---|---|---|---|
| flag | Yes | ||
| project | Yes | ||
| environment | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes that this is a safe read operation, lowering the bar. The description adds that the result is current and includes the enabled flag plus ordered rules, but does not disclose error behavior, permissions, or how order is determined. That is adequate but leaves some behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence conveys the object, scope, and return contents without filler. Every phrase adds value, and the parenthetical clarifies what 'targeting configuration' means.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only getter with three self-named parameters and no output schema, the description is mostly complete: it states what the tool returns and the environment scope. The main shortfall is not explaining project/flag identifier semantics, but sibling list tools and property names make this recoverable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the three string parameters. It only contextualizes 'environment' via 'in one environment'; 'project' and 'flag' receive no additional meaning beyond their property names, leaving the agent to infer identifiers and required formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Get'), a specific resource ('a flag's current targeting configuration'), and the result content ('enabled flag + ordered rules') in one environment. This clearly differentiates it from siblings like get_flag and update_targeting.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: use this when you need the current targeting configuration for a specific flag in a specific environment. It does not explicitly name alternatives or exclusions, but the purpose is narrow and unambiguous enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_environmentsList environmentsARead-only
List the environments of a project (e.g. development, staging, production).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| project | Yes | Project key |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include readOnlyHint: true, so the agent already knows this is a read-only operation; the description adds project scoping and example environment names. It does not disclose pagination behavior or the return structure, but for a simple read-only list this is a moderate gap rather than a serious one.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no filler, front-loading the action and object while using the parenthetical example to clarify the concept. Every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with one required parameter, the core purpose and scope are clearly stated and the annotation covers safety. However, the absence of an output schema and any mention of pagination or returned fields leaves the agent uncertain about the result shape, so the description is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, covering only the 'project' parameter, while 'limit' and 'cursor' are undocumented. The description provides no additional meaning for these parameters and only loosely maps 'of a project' to the required project key, leaving the agent to guess at pagination semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('List'), a clear resource ('environments'), and a scoping qualifier ('of a project'), with concrete examples (development, staging, production). It is immediately distinguishable from sibling tools, which cover projects, flags, and segments, so an agent can identify this tool without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is intended for retrieving environments tied to a project, which gives a reasonable usage context. However, it does not explicitly state when to prefer this over alternatives or mention any exclusions or prerequisites other than the project scope.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_flagsList feature flagsARead-only
List feature flags in a project. Filter with search (key/name substring), tag, type, archived, owner. Each flag carries its owner ({id, email, name}, or null when unowned). Paginated via cursor.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | ||
| type | No | ||
| limit | No | ||
| owner | No | Only flags with this owner: an email, "me" for the token's own user, or "none" for unowned flags | |
| cursor | No | ||
| search | No | ||
| project | Yes | Project key | |
| archived | No | true = only archived, false = only active |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already covering the safety profile, the description adds real context: the owner object shape ({id, email, name} or null when unowned) and cursor-based pagination. It still omits the response envelope and default page size/limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with purpose and then filters, with no wasted wording. Slightly dense enumeration but efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description usefully carries the owner-null shape and pagination contract, and the readOnly annotation covers safety. Completeness holds, though filtering defaults and result ordering are not addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 38%, and the schema already documents owner and archived. The description adds one genuinely new detail — that search matches key/name substrings — and confirms cursor pagination, but leaves limit and other param semantics undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List feature flags in a project') with clear scope. However, it does not differentiate itself from siblings like get_flag (single) or find_stale_flags, so an agent gets no explicit routing signal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It enumerates the available filters, which implies when filtering is useful, but gives no explicit when-to-use versus get_flag, find_stale_flags, or other listing tools. Usage is left for the agent to infer rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_projectsList projectsARead-only
List projects in the Featureflip organization. Paginated: pass cursor from next_cursor to continue.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Page size (default 20) | |
| cursor | No | next_cursor from a previous response |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=true already establishes that this is a safe, non-mutating operation. The description adds the pagination behavior (cursor-based continuation), which is useful context beyond the annotation. However, it does not mention any rate limits, authorization requirements, or response format details, which are not required for a simple read-only listing but would add richness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences deliver the essential purpose and the key pagination behavior with zero filler. The primary action is front-loaded, and the pagination note is immediately actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, read-only, paginated list tool with two fully documented optional parameters and no output schema, the description covers everything an agent needs to call it correctly. The annotation covers the safety profile, and the schema covers the parameter details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: limit is described as 'Page size (default 20)' and cursor as 'next_cursor from a previous response' in the schema itself. The description adds little beyond what the schema already provides, though it does reinforce the pagination relationship between the cursor and a prior response. This meets the baseline for full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List projects in the Featureflip organization') and is clearly distinct from all siblings, which target flags, environments, and segments rather than projects. An agent can immediately understand what this tool does and how it differs from other list_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit pagination usage: 'pass cursor from next_cursor to continue.' This tells the agent how to fetch additional pages. It does not state when to use this over alternatives, but none of the siblings list projects, so there is no real alternative to exclude.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_removal_candidatesList removal candidatesARead-only
List the flags the server is confident can be removed from code, the same list the Featureflip flag cleanup GitHub Action works from. Each item has key, reason, status (Dead or Stale), treatment (true = keep the on-branch, false = keep the off-branch) and blockedBy: the live flags that still list it as a prerequisite. Only remove a flag whose blockedBy is empty. A flag that other live flags list as a prerequisite has to be removed LAST, dependents first: archiving it while a dependent still names it makes that dependent fail its prerequisite check and serve its off variation, so archive_flag refuses it with FLAG_HAS_DEPENDENTS. staleness "dead" (the default) returns only dead flags, "stale" adds stale ones. Paginated via cursor.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| project | Yes | Project key | |
| staleness | No | Minimum staleness tier: dead (default) or stale |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations give only readOnlyHint, and the description goes well beyond it: it enumerates the returned fields (key, reason, status, treatment, blockedBy), explains the dead-vs-stale filter semantics and default, notes cursor pagination, and discloses the cross-tool failure mode (FLAG_HAS_DEPENDENTS) that constrains deletion order.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then the item shape, then the ordering constraint. Dense and largely waste-free, though the multi-clause explanation of the dependents rule is slightly long for a list endpoint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description's item-field enumeration is essential, and it is provided. Combined with the staleness/pagination notes and the deletion-ordering prerequisite, an agent has everything needed to call it and act on results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 50% schema description coverage, the description usefully clarifies staleness semantics (dead default, stale adds stale flags) and pagination via cursor. The limit bound and project parameter are not elaborated, but project is self-evident and limit is bounded by the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (list flags the server is confident can be removed) and identifies the exact artifact it mirrors (the Featureflip flag cleanup GitHub Action list). It is clearly distinguishable from list_flags and archive_flag without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong operational guidance: only remove flags whose blockedBy is empty, and remove dependents first with an explicit consequence for violating that order (archive_flag refuses with FLAG_HAS_DEPENDENTS). It does not, however, explain when to prefer this over the sibling find_stale_flags, so the routing between the two candidate-listing tools is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_segmentsList user segmentsARead-only
List reusable user segments in a project. Segments are referenced from targeting rules via userSegmentId.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| project | Yes | Project key |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds useful context about segments being reusable and referenced via userSegmentId, but it does not disclose pagination behavior, result shape, or ordering. This is acceptable given the read-only annotation but not especially rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The core action is front-loaded, and the second sentence earns its place by explaining how segments are referenced, which is useful context for an agent deciding whether to resolve segment IDs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple read-only list operation and identifies the required project context. However, with no output schema and only 33% parameter coverage, it does not explain what fields each segment contains or how pagination via limit/cursor behaves, leaving moderate gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, with project described but limit and cursor left undocumented. The description reinforces the project scope but does not explain how limit and cursor control pagination or output size, so it fails to compensate for the schema gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List'), a specific resource ('reusable user segments'), and scope ('in a project'), and adds the useful detail that segments are referenced from targeting rules via userSegmentId. This clearly differentiates it from siblings like get_segment (single segment) and list_flags (flags, not segments).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys that this is the tool for listing reusable segments within a project, which gives implied context for when to use it. However, it does not explicitly contrast it with get_segment or mention any exclusions, so the guidance is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_webhook_deliveriesList webhook deliveriesARead-only
List one subscription's delivery attempts, most recent first. status is Pending (awaiting an attempt, backing off, or in flight), Succeeded or DeadLettered; lastResponseStatusCode and lastError explain a failure. Paginated via cursor. Requires an Admin token. A 404 on every webhook call means webhooks are not enabled for the organization.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Subscription id (from list_webhooks) | |
| limit | No | Page size (default 20) | |
| cursor | No | next_cursor from a previous response |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true, so the description adds substantial behavioral value: delivery ordering, status enum semantics, failure-field explanations, pagination, admin auth requirement, and the organization-level 404 signal. This goes well beyond what annotations already convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences cover purpose, ordering, statuses, failure explanation, pagination, auth, and a common error signal without redundancy. Every sentence earns its place and the core purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description successfully explains key response semantics: status values, lastResponseStatusCode, lastError, and pagination. It does not describe the exact response envelope, but the agent has enough behavioral and error context to invoke the tool and interpret results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents id, limit, and cursor. The description does not add new parameter-level meaning beyond the schema; it only reinforces pagination context already captured in the cursor and limit descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and identifies a precise resource ('one subscription's delivery attempts'), and 'most recent first' adds scope. This clearly distinguishes it from sibling tools like list_webhooks (subscriptions) and deliver_webhook (manual dispatch).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this targets one subscription, is paginated, requires an Admin token, and has a specific 404 failure meaning. It does not explicitly name alternatives or state when-not-to-use, so it falls short of a 5, but the constraints are clear enough for correct selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_webhooksList webhook subscriptionsARead-only
List the organization's outbound webhook subscriptions (url, event/project/environment filters, enabled state, failure count, signing secret ids — never secret values), plus availableEventTypes: the event types a subscription can filter on. Empty filter lists mean "all". Paginated via cursor. Requires an Admin token. A 404 on every webhook call means webhooks are not enabled for the organization.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Page size (default 20) | |
| cursor | No | next_cursor from a previous response |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavior beyond the readOnlyHint annotation: it notes that secret values are never returned, empty filter lists mean 'all', pagination uses a cursor, an Admin token is required, and a 404 indicates webhooks are not enabled. This is rich, non-obvious behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense with no filler: each sentence adds a distinct fact (contents, filter semantics, pagination, auth, error meaning). The main action and scope are front-loaded, and the supporting details are compact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining what the response contains and how the tool behaves. It covers the listed fields, filter semantics, pagination, required permissions, and a special error condition, which is complete for a list operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters, so the schema already documents limit and cursor. The description adds general pagination context but no parameter-specific meaning beyond what the schema provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses a specific verb ('List') and a clear resource ('the organization's outbound webhook subscriptions'), and enumerates the included fields. It is easily distinguishable from siblings like list_webhook_deliveries because it focuses on subscriptions rather than delivery attempts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: it lists subscriptions, is paginated via cursor, and requires an Admin token. It does not explicitly name an alternative tool or state when not to use it, but the context and sibling names make the intended use clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_variationAdd, update, or remove a variationADestructive
Manage a flag's variations. action=add requires key + value; action=update/remove require variationId (get ids via get_flag). Removing a variation fails with VARIATION_HAS_DEPENDENTS if other flags depend on it.
| Name | Required | Description | Default |
|---|---|---|---|
| key | No | Required for add; immutable afterwards | |
| flag | Yes | ||
| name | No | ||
| value | No | Variation value serialized as a string; required for add | |
| action | Yes | ||
| project | Yes | ||
| description | No | ||
| variationId | No | Required for update/remove |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With destructiveHint=true already signaling mutation risk, the description adds the VARIATION_HAS_DEPENDENTS failure mode and the dependency condition. It also discloses the need to obtain variationId via get_flag, which goes slightly beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
At three sentences, the description is compact and front-loads the core action mapping. The opening sentence is somewhat redundant with the title, but the actionable conditional and failure-mode sentences earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the action-dependent parameter requirements and a key failure mode, which is the main complexity of the tool. It does not describe the response shape or the semantics of optional fields like name/description, but the critical invocation knowledge is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 38%, but the description mostly re-encodes what the schema already says for key, value, and variationId. It adds the useful 'get ids via get_flag' pointer but leaves name and description parameters unexplained, so it only partially compensates for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool manages a flag's variations and enumerates the three concrete actions (add/update/remove). It also differentiates from sibling tools by specifying action-specific requirements, making it unambiguous versus get_flag or update_flag.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear context: use this tool to mutate variations, with action=add/update/remove. It points to get_flag as the source for variationId, which is an explicit prerequisite/alternative, though it does not state formal when-not-to-use exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_webhookCreate, update, or delete a webhook subscription, or rotate its secretADestructive
action=create requires name + url and returns the signing secret. action=update/delete/rotate_secret/retire_secret require id. update changes only the fields you pass and keeps the rest; pass an empty list to widen that filter back to "all". projects take project keys or ids; environments take "/" keys (e.g. "web/production") or ids. eventTypes come from list_webhooks → availableEventTypes. create and rotate_secret return a secret that is NEVER shown again: relay it to the user verbatim so they can store it. rotate_secret adds a new active secret while the old one keeps signing; retire_secret (secretId from list_webhooks) removes one and is refused for the last active secret. A project-restricted service token must list at least one project, all in its scope. Requires an Admin token. A 404 on every webhook call means webhooks are not enabled for the organization.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Subscription id; required for everything except create | |
| url | No | Receiver URL; required for create | |
| name | No | Required for create | |
| action | Yes | ||
| enabled | No | update only; subscriptions are created enabled | |
| projects | No | Project keys or ids to cover; empty or omitted on create = all | |
| secretId | No | Required for retire_secret | |
| eventTypes | No | Event types to send; empty or omitted = all | |
| environments | No | "<project>/<environment>" keys or environment ids; empty or omitted on create = all | |
| idempotency_key | No | Idempotency-Key header for create / rotate_secret |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the destructiveHint annotation by explaining partial-update semantics, the one-time-only display of signing secrets, rotation/retirement behavior, and the 404-equals-disabled-webhooks signal. This gives an agent a reliable model of side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and free of fluff, front-loading the create requirements before the other actions. It is one long paragraph, so a bulleted layout would improve scannability, but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a multi-action tool with 10 params and no output schema, the description covers required params, one-time secret return, auth, token scoping, sibling data sources, and a failure mode. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even with 90% schema coverage, the description adds action-dependent parameter meaning: which params are required per action, how to format projects and environments, that empty lists widen filters, and where to source eventTypes/secretId. These semantics are not inferable from the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The title and description clearly identify the resource (webhook subscriptions) and the specific actions (create, update, delete, rotate_secret, retire_secret). The description's action-level requirements disambiguate it from read-only siblings like list_webhooks and webhook delivery tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear preconditions: create requires name+url, other actions require id, Admin token is required, and project-restricted tokens must list projects in scope. It points to list_webhooks for eventTypes and secretId, but it does not explicitly state when not to use this tool or name a sibling as the alternative for read-only webhook inspection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
restore_flagRestore archived flagB
Restore a previously archived flag.
| Name | Required | Description | Default |
|---|---|---|---|
| flag | Yes | ||
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description communicates that the flag transitions from an archived state back to a restored one, which is useful behavioral context. The destructiveHint=false annotation already covers the safety profile, though side effects and idempotency are not discussed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one short sentence with no filler, but it mostly restates the title and adds little new information. It is concise rather than genuinely informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool, the description is too minimal: it omits parameter semantics, fails to address how project/flag identifiers should be supplied, and gives no guidance on expected behavior or alternatives. An agent can make a reasonable guess but would benefit from richer context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not clarify what 'project' and 'flag' represent or what formats/identifiers are expected. 'flag' is mentioned generically, but 'project' is entirely unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action ('Restore') and a clear object ('previously archived flag'), so an agent can tell this is an un-archive operation. It does not explicitly contrast with archive_flag, but the inverse relationship is strongly implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'previously archived' implies the intended use case and distinguishes restoration from archive/delete operations. However, it does not explicitly name alternatives, state prerequisites, or explain when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_flag_expirySet or clear flag expiryAIdempotent
Set the date a flag is expected to be removed by, or pass expiresAtUtc: null to clear it. Expiry is advisory: evaluation never changes. Once the date passes, find_stale_flags reports the flag as expired. expiresAtUtc takes a date (2026-12-31), which means the end of that day in UTC, the same as picking it in the dashboard, or a full ISO-8601 timestamp, which is stored exactly. Refused with EXPIRY_IN_PAST for a time that has already passed (today's date is allowed), and with FEATURE_NOT_ENABLED while flag expiration is not enabled for the organization. Clearing is always allowed.
| Name | Required | Description | Default |
|---|---|---|---|
| flag | Yes | Flag key or id | |
| project | Yes | ||
| expiresAtUtc | Yes | Date (YYYY-MM-DD, end of that UTC day) or ISO-8601 timestamp, or null to clear the expiry |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only idempotentHint and destructiveHint in the annotations, the description carries the substantive behavioral load and does so well: expiry is advisory only, passing dates are refused with EXPIRY_IN_PAST, disabled orgs get FEATURE_NOT_ENABLED, today is allowed, and clearing is always permitted. That is exactly the failure-mode detail an agent needs before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The primary action and the null escape hatch are front-loaded in the first sentence, and every following sentence adds new information. It reads as one dense block rather than segmented guidance, but there is essentially no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description covers the input formats, the advisory nature of the stored value, downstream visibility via find_stale_flags, and both rejection conditions. An agent has everything required to call this correctly and interpret refusals.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% and the description more than compensates: expiresAtUtc accepts either a bare date (interpreted as end of that UTC day, matching the dashboard) or a full ISO-8601 timestamp stored exactly, and null clears. This adds semantics the schema's one-line property description does not convey, including the timezone normalization.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: setting the date a flag is expected to be removed, with the null case for clearing. It distinguishes itself from sibling mutators like delete_flag and archive_flag by framing the action as scheduling rather than removing, and it names find_stale_flags as the consumer of this state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when the tool is appropriate (advisory expiry, cleared via null) and what it is not (evaluation never changes). It does not explicitly contrast with delete_flag/archive_flag, which is the most likely confusion given the sibling set, so it falls short of full alternative-routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_flag_ownerSet or clear flag ownerAIdempotent
Name the person responsible for a flag, or pass owner: null to leave it unowned. owner is the email of an active member of the organization, or "me" for the token's own user (needs a personal access token). The owner gets the flag's cleanup notices; evaluation never changes. Refused with OWNER_NOT_MEMBER for an email that is not an active member, and with PLAN_FEATURE_UNAVAILABLE below the Pro plan. Clearing is always allowed.
| Name | Required | Description | Default |
|---|---|---|---|
| flag | Yes | Flag key or id | |
| owner | Yes | Owner email, "me" for the token's own user, or null (or "none") to clear the owner | |
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond idempotentHint/destructiveHint, it discloses concrete failure modes (OWNER_NOT_MEMBER, PLAN_FEATURE_UNAVAILABLE below Pro), the side effect (owner receives cleanup notices), that evaluation is unaffected, and that clearing is always permitted even when setting is not.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action and the clearing case, with error behavior appended. Dense but each clause carries distinct information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param write tool with no output schema, the description covers semantics, side effects, and error outcomes well. It stops short of describing the success response, but that is a minor gap given the operation is a simple setter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%; the description compensates by defining what 'owner' accepts (active member email, 'me' with a PAT requirement, null to clear) and the membership constraint that the schema does not state. Only the project/flag parameters are left to the schema, which is adequate since their schema descriptions are minimal.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Name the person responsible for a flag') and explicitly covers both the set and clear cases. It is trivially distinguishable from siblings like update_flag or get_flag.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear conditions for use: pass owner email for a member, 'me' for the token user (with the PAT caveat), or null to clear. It does not name an alternative tool, but the operation is narrow enough that no routing ambiguity exists.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
toggle_flagToggle flag in environmentADestructive
Enable or disable a flag in ONE environment. Affects live evaluation immediately — double-check the environment.
| Name | Required | Description | Default |
|---|---|---|---|
| flag | Yes | Flag key or id | |
| enabled | Yes | ||
| project | Yes | ||
| environment | Yes | Environment key, e.g. production |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The destructiveHint annotation is complemented by the description's explicit statement that the action 'affects live evaluation immediately' and the caution to double-check the environment. This adds valuable behavioral context beyond the annotation, explaining the real-world impact and the need for care.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences that front-load the action and scope, followed by a relevant safety warning. There is no filler, redundancy, or unnecessary detail; every word contributes to the tool's usability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with a destructive annotation and no output schema, the description covers the core action, scope, immediate impact, and provides a caution. It does not discuss return values or prerequisites, but those are less critical here. The description is reasonably complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%, with project and enabled lacking parameter descriptions. The description adds some meaning by implying that 'enabled' is a boolean and that environment selects a single environment, but it does not explain the project parameter or the expected format for flag key/id. It partially compensates for the schema gap but not fully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Enable or disable') on a specific resource ('flag') with a definite scope ('ONE environment'). This distinguishes it from broader sibling tools like update_flag or update_flag_environment_config, and the added impact note about live evaluation further pinpoints its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context through the single-environment scope and the warning to double-check the environment, but it does not explicitly compare against sibling tools like update_flag_environment_config or flag_status, nor does it state when to use this versus an alternative. Guidance is present but only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_flagUpdate feature flag metadataADestructive
Update flag name/description/tags/clientSideVisible. Changes only the fields you pass and keeps the rest; pass description: "" to clear it. Key and type are immutable. Use toggle_flag / update_targeting for behavior changes.
| Name | Required | Description | Default |
|---|---|---|---|
| flag | Yes | Flag key or id | |
| name | No | ||
| tags | No | ||
| project | Yes | ||
| description | No | ||
| clientSideVisible | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only destructiveHint=true declared, the description adds real value: partial-update semantics ('changes only the fields you pass and keeps the rest'), the empty-string convention for clearing description, and the immutability of key and type. It does not address permissions or reversibility, but nothing contradicts the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the mutable-field scope front-loaded, then merge semantics, then sibling exclusions. No filler and every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema mutation tool this covers the important unknowns: which fields change, that others persist, clearing behavior, and immutability. Missing only auth/permission requirements and any confirmation or rate-limit notes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 17%, so the description must compensate and largely does: it enumerates the mutable fields, states that omitted fields are preserved, and specifies the description-clearing convention. It leaves project unspecified and refers to an immutable 'key'/'type' that do not appear as parameters, which is a minor mismatch.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (Update) plus resource (flag) plus the exact field set it can change, which immediately separates it from toggle_flag and update_targeting. An agent can tell what this does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes behavior changes to toggle_flag / update_targeting, which is the key sibling disambiguation for this tool. It stops short of stating when to use this tool versus delete_flag, archive_flag, or update_flag_environment_config, so it is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_flag_environment_configUpdate flag environment configADestructive
Update a flag's per-environment serving config: defaultVariationId (served on fallthrough, required), strategy (SingleVariation | PercentageRollout | TargetedRollout — only SingleVariation is currently supported at the environment level), and prerequisites. Percentage rollouts are configured on targeting rules via update_targeting.
| Name | Required | Description | Default |
|---|---|---|---|
| flag | Yes | ||
| project | Yes | ||
| strategy | No | ||
| environment | Yes | ||
| prerequisites | No | Flags that must evaluate to the expected variation before this flag serves. Omit to leave existing prerequisites untouched; pass a list (possibly empty) to replace them wholesale. | |
| defaultVariationId | Yes | Variation id to serve on fallthrough — required by the API |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint: true, so the safety bar is covered. The description adds valuable behavioral details: defaultVariationId is 'required by the API', only SingleVariation is currently supported at environment level, and prerequisites have an 'omit to leave untouched / pass list to replace wholesale' semantic. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence that front-loads the purpose, then packs the field semantics into a parenthetical list. Nothing is redundant, though the density slightly hurts readability. It earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive operation with six parameters and no output schema, the description covers the key fields but leaves gaps: it doesn't state what happens if strategy is omitted (presumably unchanged, but not stated), nor does it clarify how project/flag/environment should be identified (key vs ID). An agent may need to consult sibling tools or schemas for these details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is low (33%), so the description must compensate. It does explain defaultVariationId (served on fallthrough, required), strategy (enum values and current limitation), and prerequisites (omit vs replace). However, it provides no meaning for project, flag, or environment identifiers, which are equally necessary to invoke the tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource scope: 'Update a flag's per-environment serving config'. It enumerates the exact fields affected (defaultVariationId, strategy, prerequisites) and explicitly routes percentage rollouts to update_targeting, making it easy to distinguish from that sibling and from the broader update_flag tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells the agent when not to use this tool: 'Percentage rollouts are configured on targeting rules via update_targeting.' This is a clear alternative. It doesn't mention when to use this vs update_flag, but the 'per-environment' scope and the listed fields provide enough contextual guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_targetingReplace targeting rulesADestructive
REPLACE all targeting rules of a flag in one environment (full PUT — rules not included are removed). Rules are evaluated in order. Get the current rules first with get_targeting.
| Name | Required | Description | Default |
|---|---|---|---|
| flag | Yes | ||
| rules | Yes | ||
| project | Yes | ||
| environment | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The destructiveHint annotation already flags danger, and the description adds meaningful behavioral detail: it is a full PUT, omitted rules are removed, and rules are evaluated in order. This goes beyond the annotation and sets proper expectations for a write operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences deliver the core action, the destructive consequence, evaluation order, and a safety tip. Every sentence earns its place, and the key 'REPLACE' and 'rules not included are removed' messaging is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive replacement tool, the description covers the essential risks, ordering, and the recommended precondition. It does not detail return values or error cases, but with no output schema and a straightforward rules-replacement model, the description is sufficiently complete; a small gap remains around how to construct valid rule objects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for top-level parameters, yet the description only elaborates on the behavioral meaning of 'rules' rather than defining the format or purpose of project, flag, environment, or rules as parameters. It does not compensate enough for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('REPLACE'), a clear resource ('all targeting rules of a flag in one environment'), and the full-PUT semantics that distinguish it from get_targeting. The phrase 'rules not included are removed' leaves no ambiguity about scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly indicates when to use this tool (full replacement) and explicitly instructs the user to fetch current rules via get_targeting first. It does not explicitly state when not to use it or what alternative to choose for partial updates, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wrap_featureWrap a feature in a new flagA
Create a Boolean feature flag and get back the SDK code snippet to guard the new code path with it. Returns the snippet only — apply the edit yourself. The flag starts DISABLED in every environment; enable it with toggle_flag when ready.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Flag key, e.g. new-checkout | |
| name | No | Display name; defaults to the key | |
| tags | No | ||
| project | Yes | Project key | |
| language | Yes | SDK language of the codebase being edited | |
| description | No | What this flag guards | |
| idempotency_key | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond destructiveHint=false, the description reveals that a flag is actually created, that the tool returns only the snippet and does not modify code, and that the flag starts disabled in all environments. This is valuable behavioral context for an agent. It doesn't cover idempotency or conflict behavior, but there is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences front-load the core behavior, then state the snippet-only constraint and the initial disabled state with the enablement path. No filler or redundant restatement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-param tool with no output schema, the description covers the essential call outcome (create flag, get snippet), the manual edit step, and the disabled initial state with enablement path. It omits snippet format and conflict/idempotency behavior, but the schema plus annotations cover enough for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description doesn't explain any parameters; the schema carries most of the meaning, with descriptions for key, name, project, language, and description, leaving tags and idempotency_key undocumented. At 71% coverage, the description doesn't compensate for the missing two but the schema is largely sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete action ('Create a Boolean feature flag') and a differentiating deliverable ('SDK code snippet'), and clarifies the flag's disabled initial state. This clearly separates it from sibling create_flag even without naming it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives the context for use — wrapping a new code path — and explicitly says the agent must apply the snippet itself, plus points to toggle_flag for later enablement. It doesn't enumerate exclusions versus create_flag, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.10- Added
list_removal_candidates
5 tool updates
v0.1.9- Changed
create_flag2 fields changed- changed
Input schema / properties / expiresAtUtc / descriptionPrevious value: -"Optional ISO-8601 UTC instant the flag is expected to be removed by (see set_flag_expiry)"New value: +"Optional date (YYYY-MM-DD, end of that UTC day) or ISO-8601 timestamp the flag is expected to be removed by (see set_flag_expiry)" - added
Input schema / properties / ownerEmailAdded value: +{ + "description": "Optional owner: the email of an active member (Pro plan and above, see set_flag_owner). Omitted, a personal access token makes its own user the owner and a service token leaves the flag unowned", + "type": "string" +}
- Changed
find_stale_flags1 field changed- added
Input schema / properties / ownerAdded value: +{ + "description": "Only flags with this owner: an email, \"me\" for the token's own user, or \"none\" for unowned flags", + "type": "string" +}
- Changed
list_flags1 field changed- added
Input schema / properties / ownerAdded value: +{ + "description": "Only flags with this owner: an email, \"me\" for the token's own user, or \"none\" for unowned flags", + "type": "string" +}
- Changed
set_flag_expiry1 field changed- changed
Input schema / properties / expiresAtUtc / descriptionPrevious value: -"ISO-8601 UTC instant, or null to clear the expiry"New value: +"Date (YYYY-MM-DD, end of that UTC day) or ISO-8601 timestamp, or null to clear the expiry"
- Added
set_flag_owner
6 tool updates
v0.1.8- Changed
create_flag1 field changed- added
Input schema / properties / expiresAtUtcAdded value: +{ + "description": "Optional ISO-8601 UTC instant the flag is expected to be removed by (see set_flag_expiry)", + "type": "string" +}
- Added
deliver_webhook - Added
list_webhook_deliveries - Added
list_webhooks - Added
manage_webhook - Added
set_flag_expiry
1 tool update
v0.1.7- Changed
archive_flag1 field changed- added
Input schema / properties / forceAdded value: +{ + "description": "Archive even though traffic is still evaluating the flag. Only for clients that can never be updated (old mobile app versions), where traffic will never drain on its own.", + "type": "boolean" +}
19 tool updates
v0.1.6- First observed
archive_flag - First observed
create_flag - First observed
delete_flag - First observed
find_stale_flags - First observed
flag_status - First observed
get_flag - First observed
get_segment - First observed
get_targeting - First observed
list_environments - First observed
list_flags - First observed
list_projects - First observed
list_segments - First observed
manage_variation - First observed
restore_flag - First observed
toggle_flag - First observed
update_flag - First observed
update_flag_environment_config - First observed
update_targeting - First observed
wrap_feature
TDQS
Scored across 26 tools
Most tools target distinct resources or actions, but find_stale_flags and list_removal_candidates both surface flags for cleanup with overlapping descriptions, and create_flag and wrap_feature both create flags. Descriptions help differentiate, yet an agent could still misselect between these pairs.
Tool names are consistently snake_case and nearly all follow a verb_noun pattern (list_flags, get_flag, update_targeting, etc.). The only minor deviation is flag_status, which is noun-first, but it remains readable and consistent in style.
With 26 tools, the server is on the heavy side for its purpose. Several operations could be consolidated (e.g., flag_status into get_flag, find_stale_flags and list_removal_candidates into one parameterized tool), making the surface larger than necessary despite a broad domain.
The set covers core feature flag lifecycle (create, read, update, delete, archive, restore), targeting, variations, cleanup, and webhooks well. Minor gaps exist for segments and environments (only list/get, no create/update/delete), but these are likely managed elsewhere and don't block core workflows.
Maintenance
Related MCP Connectors
Independent directory of agentic AI tools — search, compare & recommend via MCP. Read-only.
The OpenRouter for tools. One MCP connection gives any AI agent 254 hosted tools, pay per call.
An MCP for marketers that gives agents tools for SEO, AEO, socials and ads data. Generous free tier; works in Claude, ChatGPT, Cursor and other MCP clients.
60+ Meta Ads tools for AI agents: audits, campaign management, audiences and CAPI tracking.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables interaction with LaunchDarkly's feature flag platform through AI clients. Supports managing feature flags, AI configs, and their variations with operations like create, update, delete, and targeting configuration.30,145 npm27MIT
- AlicenseAqualityDmaintenanceEnables AI assistants to manage Featureflow feature flags, including creating and updating features, controlling feature states across environments, and managing projects, environments, and targeting rules through natural language.2210 npmMIT

DevCycle MCP Serverofficial
AlicenseNot gradedqualityBmaintenanceEnables AI coding assistants like Cursor and Claude to manage DevCycle feature flags directly from the development environment.11,185 npm20MIT- AlicenseNot gradedqualityCmaintenanceEnables querying and operating on a Segment workspace through the Segment Public API, with ~40 tools for reading and mutating sources, destinations, functions, tracking plans, and more.MIT