hitl-gate-mcp
This server acts as a Human-in-the-Loop (HITL) approval gate for AI agents, intercepting dangerous operations and requiring explicit human approval before they proceed. It integrates directly into the Cursor IDE via MCP elicitation.
assess_and_gate: Evaluate whether a proposed action is dangerous. If risky, automatically creates an approval ticket and triggers an in-IDE approval form for the user to approve or reject. The agent must not proceed unless the ticket isapproved.list_dangerous_ops: Browse the built-in catalog of operations considered dangerous (e.g., delete files,git push --force,DROP/TRUNCATE,curl|bash,chmod 777, production deploys, bulk payments, etc.).request_approval: Manually create an approval ticket for an action (specifying risk level:low,medium,high,critical) without going through automatic risk assessment.get_approval_status: Check the current status of a specific approval ticket (approved,rejected,pending,expired,cancelled).list_pending: List all approval tickets currently awaiting a human decision.list_approval_history: View an audit log of past approval decisions, filterable by status and count, showing time, action type, requester, approver, status, and reason.
Approval records are persisted to a configurable data directory (HITL_DATA_DIR), and the server optionally supports a standalone web panel in addition to the Cursor IDE form.
Provides Human-in-the-Loop approval gates for dangerous operations within VS Code (GitHub Copilot) IDE environment, allowing AI agents to request human approval before executing risky actions.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@hitl-gate-mcpassess dropping the users table"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
hitl-gate-mcp
给 AI Agent 用的 Human-in-the-Loop 审批门 MCP:危险操作先评估、再在 IDE 里点选批准/拒绝,并可在对话中查看审批记录。
运行时:Node.js ≥ 18
传输:stdio MCP
默认审批:IDE 内建表单(MCP elicitation);不可用时走聊天文案 / 可选网页面板
客户端:Cursor 与 VS Code(GitHub Copilot) 同一套接入步骤
推荐接入(两步,双端相同)
① 配置 MCP(让工具连上)
两边最终都跑:npx -y hitl-gate-mcp。差异只在配置文件路径。
Cursor(macOS 推荐,复制即用) — .cursor/mcp.json:
{
"mcpServers": {
"hitl_mcp": {
"command": "/bin/zsh",
"args": ["-lic", "npx -y hitl-gate-mcp"],
"env": {
"HITL_ELICIT": "1",
"HITL_ENABLE_PANEL": "0"
}
}
}
}用登录 shell 启动,自动带上 nvm / Homebrew 的 Node,不必写绝对路径。
若写裸"command": "npx",Cursor 从图标启动时经常报spawn npx ENOENT。
VS Code — .vscode/mcp.json:
{
"servers": {
"hitl_mcp": {
"type": "stdio",
"command": "npx",
"args": ["-y", "hitl-gate-mcp"],
"env": {
"HITL_ELICIT": "1",
"HITL_ENABLE_PANEL": "0"
}
}
}
}VS Code 若同样 ENOENT,把 command/args 改成与上面 Cursor 相同的 /bin/zsh + -lic 写法即可。
init 也会生成示例:.cursor/mcp.json.example、.vscode/mcp.json.example。
保存后在 MCP 面板确认 hitl_mcp 已连接。
② 手动 init(装上硬约束,提高遵从率)
在目标项目根目录:
hitl-gate-mcp init
# 或
npx hitl-gate-mcp init默认 --client all,一次写入 Cursor + VS Code 钩子:
客户端 | 写入内容 |
Cursor |
|
VS Code / Copilot |
|
只要其中一个客户端:
hitl-gate-mcp init --client cursor
hitl-gate-mcp init --client vscode已存在则跳过;覆盖加 --force。预览 --dry-run。
npm install -g hitl-gate-mcp # 一次
hitl-gate-mcp init # 任意项目诚实边界:只配 MCP、不跑
init→ 依赖模型遵从instructions,不保证每次都调闸门。
MCP 不能拦截 IDE 原生 Delete/Terminal。
Related MCP server: @looppause/mcp
审批弹窗不可用时(VS Code 更常见)
Agent 应把返回的
user_prompt_zh贴到对话里,停止执行可选:MCP env 设
HITL_ENABLE_PANEL=1,浏览器打开http://127.0.0.1:8787批/拒再调
get_approval_status,仅approved后继续
方式二:源码本地运行
git clone <本仓库地址>
cd <克隆下来的目录名>
npm install
npm run build把 MCP 的 command/args 改成:
node /绝对路径/到/本项目/dist/index.jsHITL_DATA_DIR 建议用绝对路径。示例见 mcp.local.example.json。
怎么用
危险操作审批
Agent 在副作用前调用
assess_and_gate(带code_context)五档风险;仅 高/致命 开审批
表单批准 / 拒绝;聊天「可以」不算批准
高/致命附带 爆炸半径;执行后
submit_execution_report
查看审批记录
说:「看审批记录」 → list_approval_history → 展示 summary_zh 表格。
可用工具
工具 | 作用 |
| 五档风险 + 爆炸半径;仅高/致命开单 |
| 批后计划 vs 实际对照 |
| 内置危险操作表 |
| 手动开单 |
| 查工单 |
| 待审列表 |
| 审批历史 |
环境变量
变量 | 默认 | 说明 |
|
| 是否使用 IDE 内审批表单 |
|
| 本机网页备用面板 |
|
| 审批数据目录(建议绝对路径) |
|
| 面板端口 |
|
| 面板是否自动开浏览器 |
本地自检
npm run smokeLicense
MIT
Available Tools
6 toolsassess_and_gateA
BUILTIN POLICY: Assess risk; if dangerous, create a ticket and request approval via Cursor in-IDE form (MCP elicitation). Returns approved/rejected when user decides in Cursor. Fallback: pending + chat/panel instructions. Call BEFORE side effects. Do NOT execute while pending/rejected.
| Name | Required | Description | Default |
|---|---|---|---|
| action | No | Optional explicit action id, e.g. delete_files | |
| intent | Yes | User request or proposed action in natural language | |
| params | No | Optional params to lock into the approval ticket | |
| requester | No | ||
| auto_create | No | Auto-create ticket when risky (default true) | |
| ttl_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses the workflow (assess, then if dangerous create ticket and await approval), and warns not to execute while pending/rejected. However, it does not describe the assessment criteria or the exact output format beyond 'approved/rejected'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense paragraph that conveys all essential information with no wasted words. It front-loads the core action: assess and gate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a policy gate tool with 6 parameters and no output schema, the description covers the essential workflow and constraints. It lacks details on return values and synchronization behavior, but is sufficient for guiding agent decisions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%, so baseline is 3. The description adds context for 'intent' (the user request) but does not explain 'action', 'params', 'requester', 'auto_create', or 'ttl_seconds' beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool assesses risk and gates dangerous actions by creating a ticket and requesting approval. It uses specific verbs like 'assess risk' and 'create a ticket', and the distinction from siblings (which are query tools) is implicit but clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly advises to 'Call BEFORE side effects' and 'Do NOT execute while pending/rejected', providing clear when-to-use and when-not-to-use guidance. It also mentions a fallback, but does not explicitly list alternatives to this tool vs siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_approval_statusA
Get approval ticket status. Only proceed when status is approved. Stop when rejected or expired.
| Name | Required | Description | Default |
|---|---|---|---|
| ticket_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of explaining behavioral traits. It reveals that the tool returns a status string and that the agent must conditionally act on it. This is adequate for a simple read operation, though it could mention error handling for invalid ticket_ids.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only two sentences, front-loaded with key information, and contains no unnecessary words. Every part serves a purpose: identifying the tool and providing action criteria.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there is no output schema, the description should explicitly state the possible return values (e.g., 'approved', 'rejected', 'expired') and the output format. It implies these values but does not specify them, leaving some ambiguity about the exact response structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage and the tool description does not add any meaning to the single parameter 'ticket_id'. It does not explain what the parameter represents or provide any formatting guidance, leaving the agent to rely solely on the parameter name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('get') and resource ('approval ticket status'), and provides conditional guidance on how to use the result. It distinguishes from sibling tools like list_pending and request_approval, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells the agent when to proceed (if status is approved) and when to stop (if rejected or expired), giving clear context for decision-making. However, it does not mention when to use this tool over alternatives, such as assess_and_gate or list_pending.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_approval_historyA
List approval audit history for the user. Call when the user asks to 看审批记录 / view approval history / audit log. Present summary_zh in chat. Optional status filter and limit.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max records to return (default 20, max 100) | |
| status | No | Filter by status; default all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description should carry full behavioral disclosure. It does not explicitly state that the tool is read-only or safe, nor does it mention side effects, permissions, or rate limits. The instruction to 'Present summary_zh in chat' implies output handling but not behavioral constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences: the first defines the tool's core action, the second adds usage guidance and a post-call instruction. Every phrase serves a purpose, and the most critical information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (two optional parameters, no output schema), the description covers the purpose, trigger phrases, and a key response field (summary_zh). However, it does not describe the full return format or pagination behavior, which would be helpful but not essential.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description merely restates the optional status filter and limit without adding new semantics or examples beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies the action (list) and the resource (approval audit history for the user). It explicitly mentions trigger phrases (看审批记录, view approval history, audit log) and instructs to present summary_zh, leaving no ambiguity about the tool's function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to call the tool ('when the user asks to...'). However, it does not explicitly exclude cases where alternative tools (e.g., get_approval_status, list_pending) would be more appropriate, nor does it compare to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_dangerous_opsB
List the built-in dangerous operation catalog used by assess_and_gate (MCP-embedded policy).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, and the description does not disclose behavioral traits such as whether the operation is read-only, requires authentication, or has rate limits. Short description lacks necessary context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence with no extraneous information. Efficiently conveys the essential purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no parameters and no output schema, the description is minimally complete. It explains what is listed and its relation to 'assess_and_gate', but lacks details on read-only nature or catalog dynamics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist, so baseline score of 4 applies. The description does not add any parameter-level detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists a 'built-in dangerous operation catalog' and ties it to 'assess_and_gate', which distinguishes it from sibling tools focused on approval workflows. However, it does not describe the catalog's format or contents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. The description only notes that the catalog is 'used by assess_and_gate', implying reference use, but no explicit when-to-use or when-not-to-use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_pendingB
List all pending approval tickets waiting for human decision.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description should explicitly state behavioral traits. It does not mention whether the operation is read-only, any authentication needs, pagination, or rate limits. The description is too brief to inform safe usage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that is efficient and front-loaded with the key action and resource. No extraneous words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lacks crucial context: it does not specify the output format (e.g., list of IDs, full objects), scope (user-specific or global), or any prerequisites. Given no output schema and no annotations, the description should compensate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so baseline 4 applies. The description does not need to add parameter information, and the schema coverage is 100% trivially.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List' and the resource 'pending approval tickets waiting for human decision,' which is specific and distinct from sibling tools like list_approval_history or get_approval_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. For example, it does not mention that this shows only pending items while list_approval_history shows all statuses, or when to use request_approval instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_approvalA
Manually request human approval (skip assessment). Uses Cursor elicitation when available. Prefer assess_and_gate for normal flow.
| Name | Required | Description | Default |
|---|---|---|---|
| risk | No | ||
| action | Yes | Action name, e.g. delete_files | |
| params | No | ||
| summary | Yes | One-line human-readable summary | |
| requester | No | ||
| ttl_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description fails to disclose behavioral traits like destruction potential, permissions, response format, or process details for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, clear and front-loaded, but slightly verbose with internal implementation detail ('Cursor elicitation').
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complex tool (6 params, nested objects, no output schema) but description omits output, error handling, and behavior after request.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 33%; description adds no meaning to any of the 6 parameters, does not compensate for undocumented params.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Describes requesting human approval and skipping assessment, clearly distinguishes from sibling 'assess_and_gate'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use (skip assessment) and when to prefer alternative (assess_and_gate for normal flow), plus mentions Cursor elicitation condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool has a distinct purpose: assess_and_gate is the main gating action, get_approval_status checks a specific ticket, list_approval_history provides audit history, list_dangerous_ops catalogs risky operations, list_pending shows pending tickets, and request_approval allows manual bypass. No two tools have ambiguous overlap.
All tool names follow a consistent verb_noun pattern (e.g., list_pending, request_approval). The slight deviation in 'assess_and_gate' is minor and still clear.
With 6 tools, the server covers the essential operations for human-in-the-loop gating without being excessive or sparse. Each tool handles a specific, necessary function.
The tool set covers the full lifecycle: assessment (assess_and_gate), manual request (request_approval), status checking (get_approval_status), pending list (list_pending), history (list_approval_history), and catalog (list_dangerous_ops). No obvious gaps for the stated purpose.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Preventive human-approval write-gate for AI agents: writes commit only after a human approves.
Human-in-the-loop review and approval for AI agents. Audit trail, approval policies, native MCP.
Human-in-the-loop for AI coding agents — ask questions, get approvals via Slack.
Authorize consequential AI agent actions before execution
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceHuman-in-the-Loop authorization gateway for AI Agents. Securely pause MCP workflows and route high-risk actions to human approvers via Slack or Email.1151MIT
- AlicenseNot gradedqualityCmaintenancePauses AI agent execution and routes approval requests to humans via Slack or email, with cryptographically signed proof of the human's decision.52MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to request human approvals with customizable forms, webhooks, and team features.69MIT
- FlicenseNot gradedqualityCmaintenanceHuman-in-the-loop approval gate for MCP agents, classifying tool calls by risk and requiring Slack-based human approval for medium/high risk actions.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/JerryFish123/Jerry_hitl_mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server