Skip to main content
Glama

πŸ›‘οΈ ADO Guard MCP

A safety-first Model Context Protocol server for Azure DevOps. Let Claude, VS Code Copilot or Cursor work with your boards, pull requests and pipelines without handing them the keys.

CI TypeScript MCP License

β–Ά Try the live playground Β· no install, runs the real engine in your browser


Why

Giving an LLM a Personal Access Token is all or nothing: if the agent can comment on a work item, it can also delete one, and anything sensitive in a ticket goes straight into the model's context. ADO Guard puts a policy engine between the model and Azure DevOps:

Guardrail

What it does

πŸ”’ Read-only by default

Zero config = no write tools exist. Disabled tools are never registered, so the model can't even see them.

πŸ‘οΈ Dry run

Write tools return a before β†’ after diff instead of changing anything.

βœ… Human approval

Risky calls return a preview and a one-time token bound to the exact arguments. The agent must show the preview and retry with the token after the user approves. Tokens are single-use and expire after 5 minutes.

🎯 Scoping

Project allowlist, protected fields (e.g. AreaPath) and protected branches (main, release/*) for pipeline runs.

πŸ”‘ Secret redaction

PATs, bearer tokens, cloud keys, private keys and password= style secrets are stripped from every response. Emails can be masked too.

🚦 Rate limiting

Sliding-window cap on writes per minute stops a looping agent.

πŸ“œ Audit log

Every call (allowed, denied, dry-run, pending or failed) is written to a JSONL file and exposed via a tool.

Related MCP server: Azure DevOps MCP Server

How it works

Claude / VS Code / Cursor ──MCP (stdio)──▢ ado-guard-mcp
                                              β”‚
                                   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                   β”‚     GuardEngine      │◀── policy.json + env
                                   β”‚ expose β†’ validate β†’  β”‚
                                   β”‚ scope β†’ cap β†’ preview│──▢ audit.jsonl
                                   β”‚ β†’ approve β†’ throttle β”‚
                                   β”‚ β†’ execute β†’ redact   β”‚
                                   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                              β”‚ PAT (never shown to the model)
                                   Azure DevOps REST API 7.1

The engine (src/core) is plain TypeScript with no Node APIs. The MCP server is a thin adapter around it, and the playground bundles the same code for the browser with a mock organization.

Quickstart

Runs locally over stdio. There's nothing to host.

Try it with demo data (no Azure DevOps account needed):

npx @modelcontextprotocol/inspector -e ADO_GUARD_MOCK=true -e ADO_GUARD_MODE=read-write \
  npx -y github:abdulbasit7010/ado-guard-mcp

Claude Desktop (claude_desktop_config.json):

{
  "mcpServers": {
    "azure-devops": {
      "command": "npx",
      "args": ["-y", "github:abdulbasit7010/ado-guard-mcp"],
      "env": {
        "ADO_ORG_URL": "https://dev.azure.com/your-org",
        "ADO_PAT": "<personal access token>",
        "ADO_GUARD_MODE": "read-write",
        "ADO_GUARD_PROJECTS": "Atlas,Helix",
        "ADO_GUARD_AUDIT_FILE": "/Users/you/.ado-guard/audit.jsonl"
      }
    }
  }
}

VS Code (.vscode/mcp.json):

{
  "inputs": [{ "id": "ado_pat", "type": "promptString", "description": "Azure DevOps PAT", "password": true }],
  "servers": {
    "azure-devops": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "github:abdulbasit7010/ado-guard-mcp"],
      "env": {
        "ADO_ORG_URL": "https://dev.azure.com/your-org",
        "ADO_PAT": "${input:ado_pat}",
        "ADO_GUARD_POLICY": "${workspaceFolder}/ado-guard.policy.json"
      }
    }
  }
}

Configuration

Variable

Default

Description

ADO_ORG_URL

(required)

e.g. https://dev.azure.com/your-org

ADO_PAT

(required)

Personal Access Token. Give it the narrowest scopes you need.

ADO_GUARD_MOCK

false

Use the built-in fictional org instead of a real one

ADO_GUARD_POLICY

Path to a JSON policy file (example)

ADO_GUARD_MODE

read-only

read-only or read-write

ADO_GUARD_DRY_RUN

false

Preview writes without executing

ADO_GUARD_ALLOW_DESTRUCTIVE

false

Enable delete_work_item and run_pipeline

ADO_GUARD_PROJECTS

all

Comma-separated project allowlist

ADO_GUARD_AUDIT_FILE

in-memory

Append audit entries as JSON lines

Environment variables override the policy file. Everything else (requireApproval, protectedFields, protectedBranches, maxResults, rateLimit, redact, tools.allow/deny) is set in the policy file, and every field has a safe default. See src/core/policy.ts.

Tools

Tool

Risk

Notes

list_projects

read

Filtered by the allowlist

search_work_items

read

Type / state / assignee / title filters, built into escaped WIQL

get_work_item

read

Full detail, HTML description stripped

list_pull_requests / get_pull_request

read

Reviewers + votes, linked work items, changed files

list_pipelines / list_pipeline_runs

read

create_work_item

write

update_work_item

write

Protected fields refused. Dry run shows a field diff.

add_work_item_comment / add_pull_request_comment

write

run_pipeline

destructive

Protected branches refused before approval is even requested

delete_work_item

destructive

Moves to the recycle bin

guard_get_policy / guard_get_audit_log

read

Lets the agent explain what it may do and what it did

Tools carry MCP annotations (readOnlyHint, destructiveHint) so clients can apply their own confirmations on top.

The approval flow

// 1. agent β†’ delete_work_item { "project": "Atlas", "id": 106 }
{ "outcome": "pending_approval",
  "preview": { "summary": "Delete Task #106 \"Add p95 latency alert for /charges\" (moves to recycle bin)" },
  "approvalToken": "approve-k3x9q2ma",
  "message": "Approval required. Show this preview to the user..." }

// 2. user says yes β†’ agent retries with the token
// agent β†’ delete_work_item { "project": "Atlas", "id": 106, "confirm": "approve-k3x9q2ma" }
{ "id": 106, "deleted": true }

// Same token, different id β†’ denied. Same token again β†’ denied. After 5 minutes β†’ denied.

Development

npm install
npm test              # 29 tests: engine, REST client request shapes, end-to-end over stdio
npm run dev           # start the server on mock data
npm run site:dev      # playground at http://localhost:5173/ado-guard-mcp/
src/
  core/          browser-safe: engine, policy, tools, redaction, limits, audit, mock org
  node/          REST client (Azure DevOps 7.1), config loader, file audit sink
  server.ts      MCP adapter (registerTool + annotations)
  index.ts       stdio entry point
site/            Vite + React playground deployed to GitHub Pages
test/            Vitest suites, including a real MCP client ↔ server round trip

Security notes

  • The guard is a second line of defence. Scope your PAT (e.g. Work Items: Read & Write only) first.

  • The PAT is only sent to Azure DevOps. It's never included in tool output or audit entries.

  • Approval tokens only help if a human actually reviews the preview. MCP clients that show tool results to the user (Claude Desktop, VS Code) make this natural.

Roadmap

  • MCP elicitation for in-client approve/reject buttons where supported

  • Wiki and repo file tools behind the same policy

  • Per-tool rate limits and a time-window policy (e.g. no pipeline runs on Fridays)

  • Publish to npm


Built by Malik Abdul Basit. I build enterprise platforms and AI developer tooling.

Available Tools

9 tools
get_pull_requestGet pull requestB
Read-only

Get one pull request with reviewers, linked work items and changed files.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
projectYesAzure DevOps project name
repositoryYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is covered structurally. The description adds value by disclosing that the result bundles reviewers, linked work items and changed files, but says nothing about permissions, visibility restrictions, or not-found behavior for an open-world read.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the action and the payload with zero filler. Every clause earns its place by naming the associated data the caller receives.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description usefully sketches the return payload (reviewers, work items, changed files), but with three required parameters and two of them undocumented it stops short of being sufficient. It is serviceable for a simple read tool but leaves gaps around identifier semantics and lookup failure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% – 'repository' and 'id' have no descriptions. The description adds no meaning about what any parameter expects (e.g. numeric PR id scoping), so the coverage gap is not compensated for. Only 'project' is documented, and that is in the schema, not the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Get') plus resource ('pull request') with scope narrowed to 'one', which implicitly distinguishes it from list_pull_requests. It also enumerates the associated data returned (reviewers, linked work items, changed files), so an agent knows this is a single-item detail fetch rather than a list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: fetching a single PR by project/repository/id. There is no explicit statement of when to prefer this over list_pull_requests, nor any prerequisite or failure-mode guidance. Adequate but leaves the agent to infer the selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_work_itemGet work itemB
Read-only

Get full details of one work item, including description and tags.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
projectYesAzure DevOps project name

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=true, so the safety and side-effect profile is fully covered without the description. The description adds only modest value by hinting at returned content ('description and tags'), but says nothing about error behavior for missing IDs or pagination, which is acceptable at this bar.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence that front-loads the action and resource, with the scope qualifier and returned fields trailing efficiently. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter read tool whose annotations already cover the safety profile, the description is minimally adequate, but with no output schema and an undocumented 'id' parameter it should do more to clarify the identifier semantics and confirm what a 'full details' payload contains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%: 'project' is documented as an Azure DevOps project name, but 'id' has no description anywhere. The description says nothing to clarify that 'id' is a work item identifier or its expected form, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Get full details of one work item') and scopes it to a single item, which implicitly contrasts with the list-oriented sibling search_work_items. However, it never names or explicitly contrasts that alternative, so the differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The singular 'one work item' implies this is the single-item retrieval path as opposed to search_work_items or list_projects, but there is no explicit when-to-use statement, no prerequisite note (e.g., needing a known ID), and no exclusion guidance. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

guard_get_audit_logShow audit logB
Read-only

Show the most recent tool calls and whether they were allowed, denied or awaiting approval.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered by structured data. The description usefully adds that entries carry an outcome classification (allowed/denied/awaiting approval) rather than raw events, but says nothing about ordering, retention window, or whether the log is scoped to the calling agent or the whole workspace. With annotations carrying the main burden, a 3 is fair.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the operative noun and the returned content both land immediately. It is efficient, though the brevity leaves the gaps noted in other dimensions unfilled.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple optional-parameter read tool with a safe annotation profile and no output schema, the description covers what the tool is and roughly what comes back. It stops short of the details an agent would want before invoking: result ordering, the scope of the log, and the effect of the limit parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter at 0% schema description coverage: the schema gives bounds and a default (1-100, default 20) but no narrative meaning. The description only obliquely implies recency ordering via 'most recent' and never explains what limit controls or what happens when it is omitted. It does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb plus resource ('Show' + 'audit log'), and it goes beyond the title by naming the payload contents: recent tool calls with their allow/deny/pending-approval disposition. It does not distinguish itself from the nearest sibling, guard_get_policy, so an agent must infer that one covers policy and the other covers call outcomes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to reach for this tool versus guard_get_policy, who can see the log, or whether it is a privileged inspection surface. Usage is only implied by the phrase 'most recent tool calls'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

guard_get_policyShow guardrail policyA
Read-only

Show the active guardrail policy so you can explain to the user what you're allowed to do.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds that the policy is 'active' and that its purpose is to explain permissions, but says nothing about the return format or how policy content is structured.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with no filler; the action and its rationale are both stated without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations about return shape, the description never indicates what the policy content looks like (rules, limits, formatting), which is the main thing an agent needs to relay it to a user. Adequate but leaves a clear gap for a zero-param read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies; there is no parameter behavior the description needs to compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Show) and resource (the active guardrail policy), and no sibling tool overlaps with policy inspection, so an agent can tell it apart immediately. The scope word 'active' further narrows what is returned.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear use context β€” call it to explain your permitted actions to the user β€” which tells the agent when this tool is relevant. It does not name exclusions or alternatives, but the sibling set contains nothing policy-related to disambiguate against.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_pipeline_runsList pipeline runsB
Read-only

List recent runs of a pipeline with their state and result.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNoMax results (capped by policy.maxResults)
projectYesAzure DevOps project name
pipelineIdYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds that the listing is limited to 'recent' runs and includes state/result, but says nothing about how far back 'recent' reaches, ordering, or the interaction with the schema's policy.maxResults cap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the resource and scope come first and nothing is repeated from the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should carry more of the return contract, but it only gestures at 'state and result'. Given a simple three-parameter read tool with annotations covering safety, this is the minimum viable level rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%: 'project' and 'top' are documented in the schema, but 'pipelineId' has no description anywhere. The description mentions no parameters at all, so it contributes nothing toward closing that gap or clarifying the default/cap behavior of 'top'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List ... runs of a pipeline') plus the payload shape ('state and result'), so the agent knows exactly what comes back. It does not distinguish itself from the nearby 'list_pipelines' sibling, which lists pipelines rather than runs, leaving that inference to the agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisite (e.g. that the pipeline must exist or which project scope is required), and no mention of alternatives such as list_pipelines. The agent must infer the use case from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_pipelinesList pipelinesC
Read-only

List pipelines defined in a project.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesAzure DevOps project name

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds nothing further β€” no note on result volume, pagination, or permissions for a tool flagged as open-world, where large result sets are plausible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler, which is appropriate for a one-parameter read tool. It is efficient, though arguably under-specified rather than optimally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with a fully documented single parameter and annotations covering safety, the definition is minimally adequate. It omits any indication of result size or pagination behavior, which matters given openWorldHint=true.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'project' parameter is already documented as 'Azure DevOps project name'. The phrase 'in a project' merely restates that parameter, adding no syntax, format, or default information beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description pairs the verb 'List' with the resource 'pipelines' and scopes it to a project, so the operation is unambiguous. It does not, however, distinguish itself from the sibling 'list_pipeline_runs', which an agent could easily confuse with it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as list_pipeline_runs or guard_get_policy. The agent must infer usage purely from the name and the required 'project' parameter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_projectsList projectsB
Read-only

List Azure DevOps projects the server is allowed to access.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety and openness profile is covered. The description adds the useful constraint that only projects the server is allowed to access are returned, which is behavioral context beyond annotations. However, it doesn't cover pagination, result ordering, or what 'allowed to access' entails.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the verb and resource and adds the essential scoping constraint. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool with annotations covering safety and openness and no output schema, the description is adequate but minimal. It doesn't explain return structure or how the 'allowed to access' filter is determined, which could matter to an agent deciding whether to rely on this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, so the baseline is 4. The description appropriately does not discuss parameters and the schema is empty as expected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb (List) and resource (Azure DevOps projects) and adds scope ('the server is allowed to access'). It doesn't explicitly distinguish itself from siblings like list_pull_requests or list_pipelines, but the resource is distinct enough that confusion is unlikely.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when to use this tool versus alternatives, no prerequisites, and no mention of what an agent might do with the results. The description gives no guidance on context of use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_pull_requestsList pull requestsC
Read-only

List pull requests in a project, optionally for one repository.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNoMax results (capped by policy.maxResults)
statusNoactive
projectYesAzure DevOps project name
repositoryNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, destructiveHint=false, and openWorldHint, so the safety profile is covered. The description adds nothing behavioral beyond scope: it omits that results default to status=active, that results are capped/paginated, and what the return shape is. With annotations carrying the safety burden, this still leaves notable behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words, correctly leading with the primary action and scope. It is arguably under-specified rather than verbose, so conciseness itself is fine.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema, four parameters, and only one required field, the description is thin. It omits the default active-status filtering and result-cap behavior that an agent should know before calling, though the schema partly compensates.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: 'top' and 'project' are documented in the schema, while 'status' (with its active default and enum) and 'repository' have no schema descriptions. The description only clarifies that repository is optional, adding marginal value over the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List pull requests') plus scope ('in a project, optionally for one repository'). It does not explicitly distinguish itself from the sibling get_pull_request, but the list-vs-get distinction is inferable from names.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'optionally for one repository' hints at the narrow-vs-broad query choice, but there is no explicit when-to-use guidance, no mention of when to prefer get_pull_request, and no statement of prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_work_itemsSearch work itemsA
Read-only

Find work items in a project by type, state, assignee or text in the title. Returns compact summaries; use get_work_item for full detail.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNoMax results (capped by policy.maxResults)
textNoText to match in the title
typeNoe.g. "Bug", "User Story", "Task"
stateNoe.g. "Active", "New", "Resolved"
projectYesAzure DevOps project name
assignedToNoDisplay name or email, or "@me"

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful behavioral context not in the annotations: the response is compact summaries rather than full records. It still omits pagination/truncation behavior for a capped result set, but the return-shape disclosure is a real addition.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste, with the core capability front-loaded and the routing hint second. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only search tool with no output schema, the description covers what it searches, the shape of the result, and where to go for detail. It does not mention result caps or how to page through matches, which is a minor gap given the top parameter and policy cap already exist in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (top, text, type, state, project, assignedTo) is already documented in the schema, including defaults and caps. The description only restates the filterable fields without adding format, syntax, or matching semantics, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (find) and resource (work items) plus the filterable dimensions (type, state, assignee, title text). It explicitly differentiates from the sibling get_work_item by noting that this returns compact summaries while that tool returns full detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative (get_work_item) and the condition that selects it (need for full detail), which is clear routing guidance. It does not state exclusions or prerequisites (e.g., whether a project must exist, pagination behavior), so it stops short of explicit when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv0.1.0
    • First observedget_pull_request
    • First observedget_work_item
    • First observedguard_get_audit_log
    • First observedguard_get_policy
    • First observedlist_pipeline_runs
    • First observedlist_pipelines
    • First observedlist_projects
    • First observedlist_pull_requests
    • First observedsearch_work_items

TDQS

A3.6/5.0

Scored across 9 tools

Disambiguation5/5

Each tool targets a distinct resource and action: list vs. get pairs (work items, PRs) are clearly differentiated, and search_work_items vs. get_work_item are explicitly distinguished in the descriptions. The two guard_ tools (audit log, policy) are also unambiguous.

Naming Consistency5/5

All tools follow a consistent snake_case verb_noun pattern (list_projects, get_work_item, search_work_items, list_pipeline_runs). The guard_ prefix on two tools acts as a clear namespace rather than an inconsistency.

Tool Count5/5

Nine tools is well-scoped for a guarded read surface over Azure DevOps. Each tool earns its place with no redundant or filler operations.

Completeness3/5

The read surface covers projects, work items, PRs and pipelines reasonably, but there are notable gaps: no repository listing, no single-pipeline or single-run detail, and no write operations (create/update/comment) anywhere. Some of this may be intentional for a guardrail server, but agents will hit dead ends for common detail and mutation workflows.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers