stata-mcp
Stata MCP 서버
Windows Stata Automation COM을 통해 Claude Code가 Stata GUI를 제어할 수 있게 해주는 MCP 서버입니다.
A Claude Code MCP server that controls the Stata GUI through Windows Stata Automation COM.
中文
기능
Stata Automation COM을 통해 실제 Stata GUI 창에서 명령을 실행합니다.
do-file 절대 경로별로 Stata 세션을 유지합니다: 동일한 do-file은 동일한 Stata 창을 재사용합니다.
서로 다른 do-file은 서로 다른 Stata 창을 열어 여러 분석 작업을 동시에 쉽게 탐색할 수 있습니다.
경로가 없는 명령은 가장 최근에 실행된 do-file 세션으로 전송됩니다.
do-file 쓰기, 읽기, 추가 및 실행을 지원합니다.
Claude가 출력 결과를 분석할 수 있도록 Stata 텍스트 로그 읽기를 지원합니다.
Claude가 do-file을 수정하는 데 도움이 되도록 현재 데이터 구조, 결측치 상황 및 샘플 미리보기를 읽는 기능을 지원합니다.
데모 효과

디렉토리 구조
D:/Stata18/mcp/
├── README.md
├── pyproject.toml
├── .gitignore
├── stata_mcp.py # 兼容启动器
├── src/
│ └── stata_mcp/
│ ├── __init__.py
│ └── server.py # MCP 服务器主文件
├── runtime/
│ ├── dofiles/ # Claude/MCP 默认生成 do 文件
│ └── logs/ # Claude/MCP 默认读取或生成 log
└── examples/ # 示例 do 文件환경 요구 사항
Windows
Stata 18 MP, Automation COM 등록 완료
Python 3.10+
Python 패키지:
mcp,pywin32Claude Code
권장 설치 위치
본 프로젝트를 Stata 설치 디렉토리 아래의 mcp 폴더에 두는 것을 권장합니다. 예:
D:/Stata18/mcpStata가 다른 위치에 설치된 경우에도 해당 디렉토리 아래에 두는 것이 좋습니다. 예:
C:/Program Files/Stata18/mcp고급 사용자는 임의의 안정적인 디렉토리에 배치할 수도 있습니다. 어디에 배치하든 Claude Code MCP 구성의 스크립트 경로는 실제 stata_mcp.py를 가리켜야 합니다.
설치 단계
본 저장소를 다운로드하거나 Stata 설치 디렉토리 아래의
mcp폴더로 복제(clone)합니다.Python 의존성 설치:
pip install mcp pywin32가상 환경에 설치할 수도 있습니다:
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install mcp pywin32관리자 권한으로 PowerShell을 열고 Stata Automation을 등록합니다. 본인의 Stata 설치 경로에 맞춰 명령어를 수정하세요:
Start-Process -FilePath "D:\Stata18\StataMP-64.exe" -ArgumentList "/Register" -WaitCOM 사용 가능 여부 확인:
python -c "import win32com.client; s=win32com.client.Dispatch('stata.StataOLEApp'); s.DoCommand('display 12345')"예상 결과: Stata GUI가 열리거나 연결되며 12345가 표시됩니다.
Claude Code에 추가 (권장)
claude mcp add 사용을 권장합니다:
claude mcp add stata -- python D:\Stata18\mcp\stata_mcp.py다른 경로에 설치한 경우 실제 경로로 대체하세요. 예:
claude mcp add stata -- python "C:\Program Files\Stata18\mcp\stata_mcp.py"추가 후 Claude Code를 재시작하여 MCP 도구 목록에 stata 도구가 나타나는지 확인합니다.
수동으로 Claude Code MCP 구성
claude mcp add를 사용하지 않는 경우, MCP JSON을 수동으로 구성할 수도 있습니다:
"stata": {
"command": "python",
"args": ["D:\\Stata18\\mcp\\stata_mcp.py"]
}메인 파일을 직접 가리킬 수도 있습니다:
"stata": {
"command": "python",
"args": ["D:\\Stata18\\mcp\\src\\stata_mcp\\server.py"]
}일반적으로 호환성 런처인 stata_mcp.py를 사용하는 것이 좋습니다. 이렇게 하면 향후 프로젝트 내부 구조가 변경되어도 구성이 더 안정적입니다.
구성 가능한 환경 변수
일반 사용자는 설정할 필요가 없습니다. 설치 경로 또는 COM 이름이 특수한 경우 설정할 수 있습니다:
변수 | 기본값 | 역할 |
| 프로젝트 루트 자동 인식 | MCP 프로젝트 디렉토리 |
| MCP 상위 디렉토리 | Stata 설치 디렉토리 |
|
| README 등록 안내용 |
|
| Stata Automation ProgID |
도구
도구 | 기능 |
| 최근 do-file 세션에서 명령 실행 |
| do-file 절대 경로로 Stata 세션 실행/재사용 |
| do-file 쓰기; 일반 파일명은 |
| do-file 내용 읽기 |
| do-file 내용 추가 |
| Stata 텍스트 로그 읽기 |
| Stata GUI에서 외부 패키지 설치 |
| Stata GUI에서 |
| Stata GUI에서 |
| 현재 데이터 구조, 결측치 요약 및 샘플 미리보기를 읽어 Claude에게 반환 |
| MCP 및 Stata 세션 상태 확인 |
do-file 및 세션 규칙
stata_run_dofile(path)는 do-file 절대 경로를 세션 식별자로 사용합니다.동일한 do-file 경로는 동일한 Stata 창을 재사용합니다.
서로 다른 do-file 경로는 서로 다른 Stata 창을 생성합니다.
stata_run,stata_get_results,stata_get_data_info,stata_get_data_schema는 기본적으로 가장 최근에 실행된 do-file 세션으로 전송됩니다.아직 do-file을 실행한 적이 없는 경우, 경로가 없는 명령은 먼저 do-file을 실행하라는 메시지를 표시합니다.
현재 데이터 구조 및 샘플 읽기
stata_get_data_schema는 최근 do-file 세션에서 텍스트 로그 스냅샷을 생성하고 그 내용을 Claude에게 반환합니다. 기본적으로 다음을 포함합니다:
describecodebook, compactmisstable summarizelist in 1/20, abbreviate(20)noteslabel dir
선택적 매개변수:
매개변수 | 기본값 | 설명 |
|
| 샘플 미리보기 행 수, 최대 1000 |
|
| compact codebook 포함 여부 |
|
| 샘플 미리보기 포함 여부 |
|
| 결측치 요약 포함 여부 |
이를 통해 Claude는 변수명, 유형, 라벨뿐만 아니라 실제 데이터의 일부를 볼 수 있어 do-file 수정을 더 정확하게 지원할 수 있습니다.
로그 전략
stata_get_data_schema는 현재 데이터 구조를 나타내므로 자체 스냅샷 로그를 덮어씁니다.다른 분석 결과는 Claude가 읽어야 할 경우 Stata가 텍스트 로그를 쓰게 한 다음
stata_read_log를 사용하여 읽어야 합니다.권장 규칙: 동일한 Claude/MCP 세션 내에서는 추가(append)하고, 다음 새 세션에서 동일한 do-file을 처음 실행할 때는 덮어써서 로그가 무한히 커지는 것을 방지합니다.
문제 해결
pywin32를 사용할 수 없음
실행:
pip install pywin32COM이 Stata 인스턴스를 생성할 수 없음
Stata Automation을 다시 등록하세요:
Start-Process -FilePath "D:\Stata18\StataMP-64.exe" -ArgumentList "/Register" -WaitGit Bash에서 /Register가 포함된 명령을 직접 실행하지 마세요. Git Bash가 /Register를 경로로 재작성할 수 있기 때문입니다.
stata_run이 do-file 세션이 없다고 알림
먼저 do-file을 실행하세요:
stata_run_dofile(path="D:/Stata18/mcp/examples/browse_test.do")그 후에는 경로가 없는 명령이 이 최근 세션으로 전송됩니다.
Claude가 Stata GUI 결과를 직접 볼 수 없음
Claude는 화면 내용을 읽을 수 없습니다. Stata가 텍스트 로그를 쓰게 한 후 stata_read_log를 통해 읽거나, stata_get_data_schema를 사용하여 현재 데이터 구조 스냅샷을 읽어야 합니다.
Related MCP server: Stata MCP Server
English
Features
Runs Stata commands in the real Stata GUI through Stata Automation COM.
Keeps one Stata session per absolute do-file path: the same do file reuses the same Stata window.
Different do files open different Stata windows, so multiple analysis tasks can stay visible.
Pathless commands are sent to the most recently used do-file session.
Supports writing, reading, appending, and running do files.
Supports reading Stata text logs so Claude can analyze command output.
Supports reading the current data structure, missing-value summary, and sample rows to help Claude edit do files.
Project layout
D:/Stata18/mcp/
├── README.md
├── pyproject.toml
├── .gitignore
├── stata_mcp.py # compatibility launcher
├── src/
│ └── stata_mcp/
│ ├── __init__.py
│ └── server.py # main MCP server
├── runtime/
│ ├── dofiles/ # default do files generated by Claude/MCP
│ └── logs/ # default logs generated or read by Claude/MCP
└── examples/ # example do filesRequirements
Windows
Stata 18 MP with Automation COM registered
Python 3.10+
Python packages:
mcp,pywin32Claude Code
Recommended install location
It is recommended to place this project inside your Stata installation directory, for example:
D:/Stata18/mcpIf Stata is installed somewhere else, place the project under that Stata directory, for example:
C:/Program Files/Stata18/mcpAdvanced users may place it in any stable directory. Wherever you install it, the Claude Code MCP configuration must point to the actual stata_mcp.py path.
Installation
Download or clone this repository into the
mcpfolder under your Stata installation directory.Install Python dependencies:
pip install mcp pywin32You may also use a virtual environment:
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install mcp pywin32Open PowerShell as Administrator, then register Stata Automation. Adjust the path to your Stata installation:
Start-Process -FilePath "D:\Stata18\StataMP-64.exe" -ArgumentList "/Register" -WaitVerify COM:
python -c "import win32com.client; s=win32com.client.Dispatch('stata.StataOLEApp'); s.DoCommand('display 12345')"Expected result: the Stata GUI opens or is connected, and displays 12345.
Add to Claude Code (recommended)
Use claude mcp add:
claude mcp add stata -- python D:\Stata18\mcp\stata_mcp.pyIf you installed the project somewhere else, replace the path with your actual path. Example:
claude mcp add stata -- python "C:\Program Files\Stata18\mcp\stata_mcp.py"Restart Claude Code after adding the MCP server, then confirm that the stata tools are available.
Manual Claude Code MCP configuration
If you do not use claude mcp add, you can configure MCP manually with JSON:
"stata": {
"command": "python",
"args": ["D:\\Stata18\\mcp\\stata_mcp.py"]
}You can also point directly to the main server file:
"stata": {
"command": "python",
"args": ["D:\\Stata18\\mcp\\src\\stata_mcp\\server.py"]
}The compatibility launcher stata_mcp.py is usually recommended because it keeps your Claude Code configuration stable if the internal project layout changes later.
Environment variables
Most users do not need these. If your installation path or COM ProgID is unusual, you can configure:
Variable | Default | Purpose |
| auto-detected project root | MCP project directory |
| parent of MCP directory | Stata installation directory |
|
| used in README registration examples |
|
| Stata Automation ProgID |
Tools
Tool | Purpose |
| Run commands in the most recent do-file session |
| Run/reuse a Stata session by absolute do-file path |
| Write a do file; simple names go to |
| Read a do file |
| Append content to a do file |
| Read a Stata text log |
| Install external packages in the Stata GUI |
| Display |
| Display |
| Return current data structure, missing summary, and sample rows to Claude |
| Show MCP and Stata session status |
Do-file and session rules
stata_run_dofile(path)uses the absolute do-file path as the session key.The same do-file path reuses the same Stata window.
Different do-file paths create different Stata windows.
stata_run,stata_get_results,stata_get_data_info, andstata_get_data_schemaare sent to the most recent do-file session by default.If no do file has been run yet, pathless commands will ask you to run a do file first.
Reading current data structure and sample rows
stata_get_data_schema creates a text log snapshot in the most recent do-file session and returns it to Claude. By default, it includes:
describecodebook, compactmisstable summarizelist in 1/20, abbreviate(20)noteslabel dir
Options:
Option | Default | Description |
|
| Number of sample rows, max 1000 |
|
| Include compact codebook |
|
| Include sample preview |
|
| Include missing-value summary |
This lets Claude see not only variable names, types, and labels, but also a small real-data sample, which helps it write better Stata code.
Log strategy
stata_get_data_schemaoverwrites its own schema snapshot log because it represents the current data state.For other analysis output, ask Stata to write a text log, then use
stata_read_logto read it.Recommended rule: append within the same Claude/MCP session; overwrite the first time the same do file is run in a new session, so logs do not grow forever.
Troubleshooting
pywin32 is unavailable
Run:
pip install pywin32COM cannot create a Stata instance
Register Stata Automation again:
Start-Process -FilePath "D:\Stata18\StataMP-64.exe" -ArgumentList "/Register" -WaitDo not run the /Register command directly in Git Bash, because Git Bash may rewrite /Register as a path.
stata_run says there is no do-file session
Run a do file first:
stata_run_dofile(path="D:/Stata18/mcp/examples/browse_test.do")After that, pathless commands are sent to the most recent session.
Claude cannot directly see Stata GUI output
Claude cannot read screen contents directly. To let Claude analyze Stata output, write a Stata text log and read it with stata_read_log; or use stata_get_data_schema for a current data snapshot.
Available Tools
12 toolsstata_append_dofileA
向已有 do 文件追加内容
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | 已有 do 文件的完整路径 | |
| content | Yes | 要追加的 Stata 代码内容 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden of disclosing behavior. It does convey that the operation is non-destructive (append) and targets an existing file, but it does not state what happens if the file is missing, how line separators are handled, or what the result of appending looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly worded sentence that front-loads the verb and target with no filler. It could have used the space to add usage or behavior notes, but as written it is appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity append operation with fully documented parameters and no meaningful return value, the description is largely sufficient. The only gaps are edge-case behaviors such as missing files, which are secondary to selecting and invoking the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both path and content already have clear descriptions. The tool description adds no additional parameter-level meaning, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb, 'append' (追加), and a specific resource, an existing Stata do-file (已有 do 文件). This clearly differentiates it from sibling tools like stata_write_dofile (create/overwrite) and stata_read_dofile (read) without needing to open the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase '已有' implies the file must already exist, and 'append' implies adding to rather than replacing content, so usage is inferable. However, there is no explicit when-to-use/when-not-to-use guidance or mention of alternatives such as stata_write_dofile for creating/overwriting.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stata_get_data_infoB
运行 describe,并通过 session 最新运行日志将当前数据集基本信息直接返回给调用端
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does reveal the mechanism: run `describe`, read the session's latest run log, and return basic info to the caller. However, it does not mention prerequisites (e.g., a dataset must be loaded), potential side effects on the session, or what happens on failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the core action (`运行 describe`) and then explains the mechanism and result. There is no redundant information or filler; every phrase contributes meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a zero-parameter tool, but it leaves gaps: the exact meaning of '基本信息' is vague, there is no mention of required session/dataset state, and it does not clarify how this differs from `stata_get_data_schema`. Given there is no output schema or annotations, some additional context would help an agent invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. There is no parameter information missing, and the description does not need to compensate for undocumented inputs. The empty schema is consistent with the description's focus on an action without arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the action: it runs Stata's `describe` command and returns basic information about the current dataset to the caller. It conveys the resource and outcome specifically, though it does not explicitly distinguish itself from the closely related sibling `stata_get_data_schema`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives like `stata_get_data_schema` or `stata_get_results`. The intended context is implied by the tool name and description, but there are no explicit usage conditions or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stata_get_data_schemaB
使用独立且每次覆盖的 schema text log,读取当前 Stata 数据集结构、缺失摘要和样本预览并返回给调用端
| Name | Required | Description | Default |
|---|---|---|---|
| sample_rows | No | 样本预览行数,默认 20 | |
| include_sample | No | 是否包含样本预览,默认 true | |
| include_missing | No | 是否包含缺失值摘要,默认 true | |
| include_codebook | No | 是否包含 codebook, compact,默认 true |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the behavioral burden. It does disclose a meaningful implementation detail — the schema text log is independent and overwritten on every call. It does not mention potential side effects, required session state, or what happens if no dataset is loaded, but it provides more than a bare statement of purpose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The entire description is one sentence with no filler or repetition. It is compact, though the opening clause about the schema text log delays the main action slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description gives only a broad sense of return content. It covers the data structure, missing summary, and sample preview, but omits usage boundaries, return format, or error conditions. It is adequate for a moderately simple read tool, but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so parameters are already documented. The description loosely aligns with the parameters by mentioning sample preview and missing summary, but it adds no detail about defaults, combinations, or behavior beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (read) and resource (current Stata dataset structure, missingness summary, and sample preview), and adds a unique mechanism (independent schema text log). It is clear, though it does not explicitly contrast with the similarly named sibling stata_get_data_info.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: call this when you need the current dataset's schema, missing summary, or sample preview. However, it gives no explicit guidance about when to choose this over stata_get_data_info or stata_get_results, so alternatives are not addressed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stata_get_resultsA
运行 return list 或 ereturn list,并通过 session 最新运行日志将 r() 或 e() 存储结果直接返回给调用端
| Name | Required | Description | Default |
|---|---|---|---|
| type | Yes | 结果类型:r 或 e |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It discloses that the tool executes commands and relies on the session's latest run log to extract results, which is useful context. However, it does not mention prerequisites (e.g., an active session or prior command producing r()/e()) or potential side effects of running commands, so it is only partially transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the action and then explains the delivery mechanism. There is no redundant phrasing; all included details (command execution, session log, result type) earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a simple input schema and no output schema, so the description needs to explain return values and prerequisites. It describes what is returned (r()/e() stored results) but does not specify the output format or note that a Stata session with recent results must exist, leaving modest gaps for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents the `type` parameter (enum r/e) with 100% coverage. The description adds meaning by associating the type with the specific commands (`return list` vs `ereturn list`) and clarifying that returned values correspond to r() or e() stored results, going beyond the bare enum documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: it runs `return list` or `ereturn list` and returns the stored r()/e() results to the caller. It names the exact Stata commands and the result type, which clearly distinguishes it from siblings like `stata_run` and `stata_read_log`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for retrieving Stata stored results after commands, but it does not explicitly state when to use it over alternatives such as `stata_read_log` or provide any exclusions. The use case is inferable but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stata_install_packageA
在已打开的 Stata GUI 中安装外部包,如 estout、ivreg2、rdrobust 等
| Name | Required | Description | Default |
|---|---|---|---|
| source | No | 安装来源:ssc(默认)或 net | |
| package | Yes | 包名称,例如 estout、ivreg2、rdrobust |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the prerequisite that a Stata GUI must already be open, but it does not mention side effects of installation, network requirements, overwriting behavior, or failure conditions. For a mutating operation, this is a significant transparency gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that conveys the action, target, environment, and examples with no filler. Every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter installer, the description plus fully documented schema provide enough information to invoke the tool correctly. The main gaps are behavioral side effects and expected output, but these are not blocking for a basic install operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents both parameters: package names and source with an enum and default. The description adds example packages but no additional meaning beyond what the schema provides, matching the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('install'), a clear resource ('external packages'), and the environment ('in the already open Stata GUI'), with concrete examples like estout, ivreg2, and rdrobust. This clearly distinguishes the tool from sibling operations such as running commands, writing do-files, or reading results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when an external Stata package needs to be installed in an already-open GUI session. However, it does not explicitly compare it to alternatives such as stata_run, which could also execute an install command, nor does it state when not to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stata_read_dofileA
读取 do 文件内容,返回给 Claude 检查、解释或继续修改
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | do 文件的完整路径 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It does say the tool reads and returns content, which implies a non-mutating operation, but it does not explicitly state that it does not execute the do-file or trigger Stata, a relevant distinction from sibling tools like stata_run_dofile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no redundant or filler content. It states the action, the resource, and the intended purpose efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool, the description provides the essential context: what it reads and why. It could be slightly more complete by explicitly stating that it does not execute the file, but the use case is still clear enough for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents the only parameter, 'path', with the description 'do 文件的完整路径', giving 100% coverage. The tool description adds no parameter-specific meaning beyond the schema, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: '读取 do 文件内容' (read do-file content). This clearly distinguishes it from sibling tools like stata_write_dofile and stata_append_dofile, and the purpose '检查、解释或继续修改' further clarifies its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended usage is implied by the purpose phrase '检查、解释或继续修改' — use this tool when Claude needs to inspect or modify a do-file. However, it does not explicitly name alternatives or state when not to use it, leaving usage guidance implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stata_read_logA
读取 Stata text log。path 留空时读取最近 session 的 last_log_path;推荐 output_format='dict',它会把日志解析成命令-结果对,便于 AI 判断报错位置。
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | log 文件的完整路径;留空则读取最近 session 的 last_log_path | |
| tail_lines | No | 只读取最后 N 行,留空则读取全部 | |
| output_format | No | 输出格式:full=完整文本,core=去除日志框架行,dict=JSON 命令-结果对 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It discloses the default path fallback and the parsing behavior of output_format='dict' into command-result pairs, which adds real value beyond the schema. It does not mention error conditions such as what happens when no prior session exists, but the disclosed behaviors are meaningful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with front-loaded purpose and actionable defaults/recommendations. No filler, no repetition of schema content, and every clause adds useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-parameter read tool with full schema coverage, the description covers purpose, default behavior, and recommended output format. It lacks an output schema and annotation safety profile, but the core invocation decision is adequately specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All three parameters are fully described in the schema, so the baseline is 3. The description adds important semantics for path (empty means recent session's last_log_path) and output_format (dict is recommended and parses logs into command-result pairs), which is value beyond the schema. tail_lines remains schema-only but already has a clear one-line description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific verb and resource: reading a Stata text log, and the log-vs-dofile distinction in sibling names makes the target reasonably clear. It does not explicitly contrast with stata_get_results or stata_read_dofile, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides concrete usage context: leave path empty to target the most recent session's last_log_path, and prefer output_format='dict' when the AI needs to locate errors. It does not explicitly say when to use alternative tools, but the practical guidance is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stata_runA
在最近的 Stata MCP session 对应 GUI 中将 commands 作为一个完整代码块执行;每次覆盖该 session 的最新运行 text log,并自动返回实际输出。可先用 stata_session(action='set_recent') 切换目标 session。
| Name | Required | Description | Default |
|---|---|---|---|
| commands | Yes | 要执行的 Stata 命令,多行用换行符分隔 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It explicitly mentions that the tool overwrites the session's latest run text log and returns actual output, which are important side effects. It does not mention error handling or permission requirements, but the disclosed behavior is clear and adds value beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, with the primary action front-loaded and side effects and session-switching guidance following. It is two sentences with no fluff, and each clause carries information. It is slightly verbose due to phrasing, but overall efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, no output schema, no annotations), the description covers the essential context: the session to use, the side effect (log overwrite), the return of output, and how to switch sessions. It does not cover error handling, but that is not expected for a simple command executor. It is complete enough for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds a nuance by describing 'commands' as a 'complete code block', but does not provide additional syntax or format details beyond what the schema already states. It adds minimal semantic value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the primary action: executing 'commands' as a complete code block in the most recent Stata session. It also mentions the side effect of overwriting the session's log and returning output, which adds specificity. However, it does not explicitly differentiate from sibling tools like 'stata_run_dofile', so it lacks explicit sibling distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a usage hint by mentioning that one can switch the target session using 'stata_session(action='set_recent')', which is helpful context. But it does not explicitly state when to use this tool versus alternatives like 'stata_run_dofile' or when not to use it. The guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stata_run_dofileA
在 Stata GUI 中运行一个 do 文件。该工具以 session_id 表示一个复现任务/同一个 Stata GUI;同一任务的原始 do、检查 do、续跑 do 应使用同一个 session_id。每次调用都会覆盖该 session 在项目 .stata-mcp/cache/ 下的最新运行 text log,并自动返回日志内容;不会修改用户原始 do 文件。
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | do 文件的完整路径,例如 D:/project/analysis.do | |
| role | No | 该 do 在 session 中的角色:entry/source/auxiliary,默认 entry | |
| log_mode | No | 兼容参数;1.1 版始终使用 replace 覆盖 session 最新运行日志 | |
| session_id | No | 可选;同一复现任务固定使用同一个 session_id。不传时兼容旧行为:用 do 文件路径作为 session key |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It transparently discloses that each call overwrites the session's latest running log in .stata-mcp/cache/, automatically returns the log content, and does not modify the user's original do file. These are important side effects and constraints. It does not mention potential blocking behavior or error handling, but the disclosed information is substantial and accurate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is composed of a few sentences, with the primary purpose stated first. It includes essential behavioral details (log overwriting, return value, non-modification) without unnecessary verbosity. The structure is clear and efficient, though it could benefit from a slight reordering to front-load the most critical behavioral note about log overwriting.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there is no output schema, the description explains what the tool returns (log content). It covers the key aspects an agent needs to know: setting up session_id for related tasks, the behavior of the log, and the safe handling of the original file. It does not mention return format or error handling in detail, but for a run tool with this parameter set, it is fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides 100% coverage with descriptions for all parameters. The description adds value beyond the schema by explaining the session_id semantics (same session for related tasks) and clarifying that log_mode is a compatibility parameter that always uses 'replace' in version 1.1. This enhances the agent's understanding of how to use the parameters correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: running a do file in Stata GUI. It specifies the resource (do file) and the environment (Stata GUI). It is distinct enough, though it does not explicitly differentiate from the sibling 'stata_run' tool, which could be a potential confusion point. Overall the purpose is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides useful usage context: it explains the session_id concept (same session for original/check/continuation do files) and notes the behavior of overwriting the latest log. However, it does not give explicit guidance on when to use this tool versus alternatives like stata_run or stata_append_dofile. The context is helpful but not a full usage guideline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stata_sessionB
管理 Stata MCP session:列出、查询、销毁、切换最近 session。session_id 代表同一复现任务/同一个 Stata GUI,并关联 entry/source/current do、每个 do 的 log_paths 和 last_log_path。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | 操作:list/get/destroy/set_recent | |
| session_id | No | get/destroy/set_recent 需要指定的 session_id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are absent, so the description must carry the full behavioral burden. It does explain what session_id represents and what it is associated with, but it does not disclose the side effects of destroy, what get returns, how set_recent changes behavior, or whether these operations are safe or destructive. A session-management tool with a destroy action needs clearer behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence that front-loads the tool's purpose and action list, then adds essential session-semantics context. There is no fluff or repetition of the schema. It could be slightly better structured by separating the action enumeration from the session_id semantics, but it remains appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description should clarify return values and behavioral outcomes for each action. It adequately covers what session_id means and which actions need it, but it does not explain what list/get return, what destroy removes, or what set_recent actually switches. The tool is simple enough that this is a moderate gap, not a fatal one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by explaining that session_id identifies the same reproduction task/GUI and is linked to entry/source/current do and log paths. It also clarifies that get/destroy/set_recent require session_id, which is not reflected as a conditional requirement in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource (Stata MCP session) and enumerates the four supported operations: list, get, destroy, and set_recent. This maps directly to the action enum and distinguishes the tool from siblings focused on running or editing do-files. '管理' alone would be vague, but the explicit action list makes the purpose concrete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit guidance on when to use this tool versus alternatives, and it does not explain when one action should be preferred over another. The action enum and the session_id note imply some usage contexts, but the description never states conditions like 'use list to discover active sessions' or 'destroy removes the session and its associated state.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stata_statusA
检查 Stata MCP 服务器状态、内存 session 和项目局部 .stata-mcp/cache/task_registry.json 中记录的 session/do/log 关系
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden. '检查' strongly implies a read-only inspection, and naming the project-local cache file adds useful concrete context about the data source. However, it does not explicitly state side-effect-freeness, permission requirements, or behavior when the registry file is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence that front-loads the action and lists its targets without filler. The inclusion of the specific cache file path is dense but relevant, and no sentence or phrase is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the zero-parameter schema and the straightforward status-checking nature, the description adequately covers what the tool does and what it inspects. It does not spell out the exact response format, but for a status/inspection tool that information is largely inferable from the listed targets.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
This tool has zero parameters and an empty input schema, so there is no parameter documentation burden. The description still adds semantic value by explaining what the status check covers, which is sufficient for an agent to invoke the tool correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the concrete verb '检查' (check/inspect) and identifies three specific objects: Stata MCP server status, in-memory sessions, and session/do/log relationships recorded in .stata-mcp/cache/task_registry.json. This resource-level specificity clearly distinguishes it from siblings like stata_run or stata_session, even without naming them. It is not a tautology and tells an agent exactly what the tool inspects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance is provided, and no sibling alternatives are named. The intended usage is implied by the content: call this tool to check server/session/cache status. However, it lacks explicit routing or exclusion conditions that would make this dimension stronger.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stata_write_dofileA
将 Stata 代码写入 do 文件并保存到磁盘,返回文件路径;相对文件名会写入最近 session 项目的 .stata-mcp/dofiles/,不会使用 MCP runtime。
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | do 文件的完整内容 | |
| filename | No | 文件名(不含扩展名)或绝对路径;留空则自动生成时间戳文件名 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description carries the disclosure burden and does add meaningful behavior: it writes to disk, returns the file path, routes relative filenames to .stata-mcp/dofiles/, and notes that MCP runtime is not used. It does not disclose overwrite behavior or session prerequisites, but the disclosed traits are substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense, front-loaded senrence that includes action, return value, path semantics, and a runtime caveat. Every clause adds information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with full schema coverage, this is nearly complete: it specifies the return value, the default location behavior, and the runtime exception. It could be more complete about overwrite behavior or what happens if no session project exists, but those are edge cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers 100% of parameters, so the baseline is 3. The description adds no new parameter-level meaning beyond what the schema already states about filename and auto-generation; it mostly restates the path behavior already present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('write'), a resource ('do file'), and the outcome ('saved to disk, returns file path'). The description clearly distinguishes this from sibling tools like stata_run_dofile or stata_append_dofile by focusing on writing and saving.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for creating or overwriting do-files and clarifies path resolution for relative filenames, but it does not explicitly say when to prefer this tool over alternatives such as stata_append_dofile or stata_run_dofile. No when-not-to-use guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
v1.1.0- First observed
stata_append_dofile - First observed
stata_get_data_info - First observed
stata_get_data_schema - First observed
stata_get_results - First observed
stata_install_package - First observed
stata_read_dofile - First observed
stata_read_log - First observed
stata_run - First observed
stata_run_dofile - First observed
stata_session - First observed
stata_status - First observed
stata_write_dofile
TDQS
Scored across 12 tools
Each tool targets a distinct artifact (do file, log, session, package, data metadata), and the descriptions clarify the boundaries well. A few pairs like stata_run/stata_run_dofile and stata_get_data_info/stata_get_data_schema are close enough that an agent could initially pick the wrong one.
Almost all tools follow a consistent stata_verb_noun pattern in snake_case. Minor deviations are stata_session and stata_status, which are noun-only names rather than verb_object, and stata_run which lacks an explicit object.
12 tools is well within the ideal scope for a domain-specific server. Each tool covers a meaningful part of the do-file, session, log, and data-inspection workflow without unnecessary bloat.
The tool surface covers the full write-do-file, append, read, run, read-log, install-package, and retrieve-results workflow coherently. Minor gaps exist around explicit session creation and do-file deletion, but arbitrary command execution via stata_run mitigates most dead ends.
Maintenance
Related MCP Connectors
Use your Mac, Windows or Linux computer from ChatGPT, Claude or Codex: files, commands, documents.
Control Unreal Engine to browse assets, import content, and manage levels and sequences. Automate…
Streamline your Attio workflows using natural language to search, create, update, and organize com…
Automate eSignature workflows and signing tasks via natural language commands.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceProvides a bridge between Stata statistical software and code editors like VS Code and Cursor, enabling users to run Stata commands directly from the editor, view output in real-time, and get AI-powered assistance with Stata coding.506MIT
- AlicenseNot gradedqualityDmaintenanceConnects AI agents to a local Stata installation, enabling execution of Stata code, data inspection, graph generation, and result verification through natural language interactions.731 PyPI86AGPL 3.0
- FlicenseNot gradedqualityDmaintenanceEnables an agent to execute Stata do scripts or inline commands locally and retrieve structured results including status, error diagnostics, and output text for consumption by the model.1-
- AlicenseAqualityBmaintenanceEnables LLM/agent applications to execute Stata code, load and inspect data, obtain structured regression results and graphs, and manage background tasks through the Model Context Protocol.10MIT