spss-studio-mcp
This server lets AI agents use IBM SPSS Statistics on macOS as a statistics engine and chart factory, with safe execution and structured results.
Run arbitrary SPSS syntax and return Markdown/JSON structured results with statistical summaries
Perform common statistical analyses: t-tests, ANOVA, regression, logistic/ordinal regression, GLM, mixed models, correlations, factor analysis, reliability, nonparametric tests, MANOVA, survival analysis (Cox, Kaplan-Meier), discriminant analysis, clustering
Run mediation and moderation analyses (Baron & Kenny + Sobel test; mean-centred interaction regression)
Export publication-ready charts as 300-dpi PNG/TIFF (histograms, scatter, bar, line, boxplot, error bar, Q-Q plot, KM curve, area, density, bar with error)
Read, inspect, summarize, and convert SPSS .sav files without needing SPSS installed (list files/variables, metadata, data preview, file summary, CSV import)
Validate syntax safely with dry-run mode, dangerous-command blocking, directories allowlist, audit logging, and status checks
Discover registry-backed methods and inspect their schemas/support metadata
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@spss-studio-mcpRun an independent-samples t-test on experiment_study.sav (grouping: group, dependent: posttest)"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
SPSS Studio MCP for Mac
Make SPSS an Agent's "statistics engine + chart factory": paper-ready charts, deep result parsing, verified methods, and safe execution.
This MCP is Mac Only, for Windows versions, see flupke91/spss-studio-mcp.
English | 繁體中文(香港)
spss-studio-mcp is an MCP (Model Context Protocol) server for IBM SPSS
Statistics. It gives Agent clients such as Codex / Claude Code / Cursor:
Paper-ready charts: 11
spss_chart_*tools export PNG / TIFF (1950×1500 @300 dpi) in one call and return the file path for direct submission to journals;Deep result parsing: OMS text → Markdown tables + structured JSON + statistical summaries for 16 analysis families (t / F / r / B / Wald / α / χ² / p / effect sizes);
Verified methods: all 37 analysis tools plus 11 supporting tools are verified on real SPSS Statistics 32 for Mac;
Mediation / moderation:
spss_mediation(Baron & Kenny three-step regression + Sobel test) andspss_moderation(mean-centred interaction regression);Safety layer: dangerous-command blocking, data path allowlist,
dry_runpreflight, and JSONL audit logging.
Everything is verified on IBM SPSS Statistics 32.0.0.0 for macOS — 105 unit tests pass, including a self-contained real-machine reproduction manifest. Per-case results: docs/macos_verification.md.
Installation
Requirements: macOS · Python 3.10+ · IBM SPSS Statistics for Mac.
SPSS 32.0.0.0 is verified and auto-detected under /Applications; without it
only the file-reading tools work.
git clone https://github.com/CongJyu/spss-studio-mcp.git
cd spss-studio-mcp
bash scripts/install_macos.sh # .venv + deps + status + Claude Code config
bash scripts/install_macos.sh --codex # same, but writes the Codex config
bash scripts/install_macos.sh --local # Claude Code → ~/.claude/settings.local.jsonThe installer creates .venv, installs -e ".[dev]", runs a status check, and
writes the MCP client config.
1. Verify
.venv/bin/spss-studio-mcp status
# pyreadstat : OK v1.3.6
# pandas : OK v3.0.5
# SPSS batch : OK - /Applications/IBM SPSS Statistics/IBM SPSS Statistics.app/Contents/bin/spssengineSPSS batch : NOT FOUND means auto-detection failed — set SPSS_INSTALL_PATH
(see below and docs/macos_verification.md).
2. Configure your client
Client | Command | Written to |
Claude Code |
|
|
Codex |
|
|
Other MCP hosts |
| prints a JSON snippet to paste manually |
Both commands merge into the existing file and leave a timestamped backup
(*.backup.YYYYMMDD_HHMMSS). configure-claude --local targets
~/.claude/settings.local.json instead. Restart the client, then ask it to
run spss_check_status to confirm the server is connected.
Manual MCP configuration
The installer symlinks spss-studio-mcp into ~/.local/bin, so the plain
command name resolves in any client. Copy the matching snippet below — no path
editing needed.
If you passed
--no-link, or~/.local/binis not on yourPATH, substitute the absolute path to your clone (/absolute/path/to/spss-studio-mcp/.venv/bin/spss-studio-mcp).
Claude Code — ~/.claude.json, or .mcp.json at a project root:
{
"mcpServers": {
"spss": {
"type": "stdio",
"command": "spss-studio-mcp",
"args": ["serve", "--transport", "stdio"]
}
}
}Codex — ~/.codex/config.toml. This file is TOML, not JSON:
[mcp_servers.spss]
command = "spss-studio-mcp"
args = ["serve", "--transport", "stdio"]Codex aborts a tool call after 60 s by default. Add tool_timeout_sec = 600
inside the table if an analysis needs longer.
Gemini CLI — ~/.gemini/settings.json, or .gemini/settings.json at a
project root:
{
"mcpServers": {
"spss": {
"command": "spss-studio-mcp",
"args": ["serve", "--transport", "stdio"]
}
}
}Gemini CLI's default request timeout is 10 minutes, which already covers a full
analysis run, so no timeout key is needed.
If status reports SPSS batch : NOT FOUND, add the detection path to the
entry's env:
"env": {
"SPSS_INSTALL_PATH": "/Applications/IBM SPSS Statistics/IBM SPSS Statistics.app/Contents/bin"
}Restart the client after editing — MCP configuration is read at startup only.
Manual install
python3 -m venv .venv
.venv/bin/python -m pip install -e ".[dev]" # -e . for runtime only
.venv/bin/spss-studio-mcp configure-claude # or configure-codex / setup-infoEnvironment variables
Variable | Default | Purpose |
| auto-detected | SPSS |
|
| Per-analysis timeout, seconds |
|
| Engine startup timeout; licensing and Python init can be slow |
|
| Set to |
| — |
|
|
| Audit log path |
Full walkthrough: docs/tutorial.md.
Related MCP server: mcp-macos-cua
Paper-Ready Charts (Core Feature)
Tool | Purpose |
| Histogram / histogram with normal density |
| Scatter plot of two variables |
| Bar chart of category means / with 95% CI error bars |
| Time-series line / area charts |
| Grouped box-and-whisker plot |
| Mean ± CI error bar chart |
| Normal Q-Q plot |
| Kaplan-Meier survival curve |
spss_chart_histogram_density(
variable="engagement_total",
title="Engagement total distribution (with normal density)",
image_format="PNG", # PNG / TIFF
width_px=1950, height_px=1500, dpi=300,
data_file="examples/data/survey_study.sav",
)
# → returns the image file path, ready for submissionFormat note: chart output is PNG and TIFF only. Windows vector EMF export was removed because SPSS for macOS cannot produce EMF metafiles (its
OMS FORMAT=DOCarchive contains only raster PNG wrapped in.eps). Useimage_format="PNG"or"TIFF".
Structured Results & Statistical Summaries
spss_structured_result(
syntax="T-TEST GROUPS=group(1 2) /VARIABLES=posttest.",
data_file="examples/data/experiment_study.sav",
)
# → {markdown, json: {tables, summary}, files, warnings}The Markdown returned by run_syntax ends with a ### Statistical Summary
block containing a plain-language conclusion plus key statistics.
Extraction details for the 16 analysis families: docs/result_parsing.md.
Sample Datasets
examples/data/ ships five research-style datasets (fixed random seeds,
fully reproducible):
File | Scenario | Key variables |
| Survey: 200 students' learning engagement |
|
| Experiment: 120 participants, pre/post memory training |
|
| Survival: 150 follow-up cases |
|
| Mediation: 300 employees |
|
| Longitudinal: 60 people, 3 waves |
|
Safe Execution
Dangerous commands (
HOST/ERASE/DELETE FILE, ...) are blocked at the start of a line;Data files must live under
examples/, the system temp dir, or a directory declared inSPSS_ALLOWED_DIRS;spss_run_syntax(..., dry_run=True)validates without executing;Audit logs are written to
logs/audit.jsonlby default (override withSPSS_AUDIT_LOG).
See docs/security.md.
Documentation
License
MIT (upstream: flupke91/spss-studio-mcp, MIT).
MIT (upstream: Exekiel179/SPSS-MCP, MIT).
Available Tools
51 toolsspss_anovaSpss AnovaB
Run SPSS one-way ANOVA (ONEWAY). Optionally includes post-hoc tests (e.g., TUKEY, BONFERRONI, LSD). Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| factor | Yes | ||
| post_hoc | No | ||
| dependent | Yes | ||
| file_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It usefully discloses a hard prerequisite ('Requires IBM SPSS Statistics to be installed') and that post-hoc testing is conditional, but says nothing about error behavior, whether the dependent must be scale, or how the factor is treated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no filler; the core operation leads and the prerequisite is placed last. Efficient, though it spends words on the optional feature while omitting required-parameter context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. The prerequisite is covered, but for a 4-parameter tool with 0% schema description coverage the definition leaves the required inputs and input-data expectations undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description explains only the post_hoc parameter by giving example values (TUKEY, BONFERRONI, LSD), which the schema lacks. The three required parameters (file_path, dependent, factor) receive no explanation in either the schema or the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and statistical resource ('Run SPSS one-way ANOVA (ONEWAY)'), which lets an agent distinguish it from siblings like spss_repeated_measures_anova, spss_manova, and spss_glm_univariate. It does not name those siblings explicitly, so differentiation relies on the agent's knowledge of the procedures.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance or exclusions relative to the many alternative ANOVA/GLM tools in the sibling list. It only notes that post-hoc tests are optional, which is a capability note rather than routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_chart_areaSpss Chart AreaA
Export a publication-ready area chart (PNG/TIFF) of y against a time/ordinal x variable, via GGRAPH + OMS.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| dpi | No | ||
| title | No | ||
| x_label | No | ||
| y_label | No | ||
| width_px | No | ||
| data_file | No | ||
| height_px | No | ||
| image_format | No | PNG |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses that the operation exports a file and the underlying mechanism (GGRAPH + OMS), but does not state where the file is written, whether it requires an open/loaded dataset, whether it mutates state, or required permissions. The PNG/TIFF disclosure is already redundant with the image_format enum in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb and resource first and zero wasted words. Tightly sized given the absence of annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers chart type, axis semantics, and output format. However, for a 10-parameter chart-generation tool it omits the file destination, dataset prerequisites, and any parameter detail, leaving it only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 10 parameters, so the description must compensate. It adds genuine meaning for the two required params by specifying x as time/ordinal and y as the plotted value, but leaves the other eight (dpi, title, x_label, y_label, width_px, height_px, data_file, image_format) completely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Export) and resource (area chart) and pins the semantics of the two axes as 'y against a time/ordinal x variable'. This cleanly distinguishes it from the many sibling chart tools (line, bar, scatter, boxplot) by chart type without needing to open a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the chart type: an agent can infer 'use this when you want an area chart'. There is no explicit when-to-use, no prerequisites (e.g., active dataset required), and no routing versus similar siblings like spss_chart_line or spss_chart_bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_chart_barSpss Chart BarC
Export a publication-ready bar chart (PNG/TIFF) of a categorical variable against the mean (or sum) of a continuous variable, via GGRAPH + OMS IMAGE.
| Name | Required | Description | Default |
|---|---|---|---|
| dpi | No | ||
| stat | No | mean | |
| title | No | ||
| value | Yes | ||
| x_label | No | ||
| y_label | No | ||
| category | Yes | ||
| width_px | No | ||
| data_file | No | ||
| height_px | No | ||
| image_format | No | PNG |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the output medium (PNG/TIFF image file) and the engine (GGRAPH + OMS IMAGE), but says nothing about where the output is written, whether data must be pre-loaded, permissions, or overwrite behavior, and none of the 11 sizing/formatting parameters are characterized.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence with the outcome front-loaded ('Export a publication-ready bar chart'), followed by the variable mapping and output formats. Nothing is padded, though the implementation aside ('GGRAPH + OMS IMAGE') is of marginal value to an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter tool with zero schema documentation and no annotations, the description is thin. The output schema covers return values, but nothing tells the agent which dataset is charted (data_file semantics), how sizing works, or why one would pick this chart over its siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 11 parameters, so the description must compensate and largely does not. It clarifies the 'stat' semantics (mean or sum) and the image formats (PNG/TIFF) plus the category/value roles, but leaves dpi, width_px, height_px, title, x_label, y_label, and data_file entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Export') and resource ('bar chart') with a clear analytical scope: a categorical variable against the mean or sum of a continuous variable. It implicitly distinguishes itself from other chart siblings (histogram, scatter, bar_error) by naming the categorical-vs-continuous mapping, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this tool over spss_chart_bar_error, spss_chart_histogram, or spss_chart_area, and no prerequisites (e.g. data must already be loaded). The agent must infer selection from the chart type alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_chart_bar_errorSpss Chart Bar ErrorC
Export a publication-ready bar chart of category means with confidence-interval error bars (PNG/TIFF) via GGRAPH + OMS.
| Name | Required | Description | Default |
|---|---|---|---|
| ci | No | ||
| dpi | No | ||
| title | No | ||
| value | Yes | ||
| x_label | No | ||
| y_label | No | ||
| category | Yes | ||
| width_px | No | ||
| data_file | No | ||
| height_px | No | ||
| image_format | No | PNG |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It says 'Export' and '(PNG/TIFF)' but never states where the file is written, whether an existing file is overwritten, whether a data file must already be loaded, or what permissions/session state are needed. 'GGRAPH + OMS' is an implementation detail, not agent-useful behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the key noun phrase front-loaded. No filler, though at only one sentence for an 11-parameter tool it is under-informative rather than over-long.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not required, but with 11 parameters, 0% schema coverage, no annotations, and multiple look-alike sibling chart tools, a single sentence leaves the agent without enough to call it correctly or safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 11 parameters, so the description must compensate. It only implicitly touches category/value (means), ci (error bars), and image_format (PNG/TIFF); the other seven parameters (dpi, title, labels, width_px, height_px, data_file) are undocumented anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (Export) and a precisely qualified resource: a bar chart of category means with confidence-interval error bars, in PNG/TIFF. This meaningfully distinguishes it from plain spss_chart_bar and spss_chart_errorbar without naming them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to choose this tool over the sibling charts (spss_chart_bar, spss_chart_errorbar, spss_chart_boxplot). No prerequisites, no mention of data_file requirements, and no indication of what kind of input data is expected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_chart_boxplotSpss Chart BoxplotB
Export a publication-ready box-and-whisker plot (PNG/TIFF) via GGRAPH + OMS. Provide a continuous variable and optionally a categorical grouping variable.
| Name | Required | Description | Default |
|---|---|---|---|
| dpi | No | ||
| title | No | ||
| x_label | No | ||
| y_label | No | ||
| category | No | ||
| variable | Yes | ||
| width_px | No | ||
| data_file | No | ||
| height_px | No | ||
| image_format | No | PNG |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses that output is a file in PNG/TIFF via GGRAPH + OMS, but not where the file is written, whether an active dataset/file must already be loaded, what happens if data_file is omitted, or what the tool returns. For a 10-parameter export tool with no annotation coverage this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and output, with zero filler. Nothing redundant and nothing padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained, but with 0% schema coverage across 10 parameters and no annotations, an agent still lacks guidance on most inputs (resolution, dimensions, format, source data file, labels). The description is materially incomplete for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are 10 parameters, so the description must compensate. It only characterizes 2 of them (variable as continuous, category as optional grouping); dpi, title, x_label, y_label, width_px, height_px, image_format, and data_file are left entirely undocumented. The two it does cover add real meaning (continuous vs categorical role), but coverage is far too thin.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Export), a specific resource (box-and-whisker plot), and the output formats (PNG/TIFF), which cleanly distinguishes it from siblings like spss_chart_histogram or spss_chart_scatter. It doesn't explicitly name those siblings, but the chart type alone is enough to differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives input preconditions (a continuous variable, optionally a categorical grouping variable), which is useful usage context, but it never says when to choose a boxplot over alternatives or what data requirements/prerequisites exist. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_chart_errorbarSpss Chart ErrorbarC
Export a publication-ready error bar chart (PNG/TIFF): mean with confidence-interval whiskers per category, via GGRAPH + OMS.
| Name | Required | Description | Default |
|---|---|---|---|
| ci | No | ||
| dpi | No | ||
| title | No | ||
| value | Yes | ||
| x_label | No | ||
| y_label | No | ||
| category | Yes | ||
| width_px | No | ||
| data_file | No | ||
| height_px | No | ||
| image_format | No | PNG |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the output type (PNG/TIFF file), the statistical construction (mean with CI whiskers), and the generation path (GGRAPH + OMS), but says nothing about where the file is written, whether an active dataset or data_file is required, or how missing data is handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the outcome ('publication-ready error bar chart') with no filler. The trailing 'via GGRAPH + OMS' is implementation detail of marginal value to an agent choosing a tool, but it is not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for an 11-parameter chart tool with zero annotation coverage and zero schema descriptions, the single sentence is too thin to be complete. The relationship to spss_chart_bar_error and the file-output behavior are the missing pieces.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 11 parameters, so the description must compensate and largely does not. It only implicitly touches 'ci' (confidence interval) and the image_format enum (PNG/TIFF) via 'PNG/TIFF'; dpi, width_px, height_px, title, x_label, y_label, and data_file are left entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('Export a publication-ready error bar chart') and specifies the statistical content (mean with confidence-interval whiskers per category). However, it never distinguishes itself from the near-identical sibling spss_chart_bar_error, leaving the agent to guess which of the two to invoke.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives, no prerequisites, and no exclusions. The existence of a confusingly similar sibling (spss_chart_bar_error) makes the absence of routing guidance a real gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_chart_histogramSpss Chart HistogramC
Export a publication-ready histogram (PNG/TIFF) via GGRAPH + OMS IMAGE. Provide the continuous variable; the chart file path is returned for direct submission.
| Name | Required | Description | Default |
|---|---|---|---|
| dpi | No | ||
| title | No | ||
| x_label | No | ||
| y_label | No | ||
| variable | Yes | ||
| width_px | No | ||
| data_file | No | ||
| height_px | No | ||
| image_format | No | PNG |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose useful behavior beyond the schema: the underlying GGRAPH + OMS IMAGE mechanism, the supported output formats, and that a chart file path is returned for submission. It omits other behavioral facts, though (whether files are written to disk, where, any session/permission requirements).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with the action and output front-loaded and no filler. It is efficient, though the brevity trades away needed detail rather than being optimally balanced.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with zero schema coverage and no annotations, this description is too thin: it covers only the required variable and the return value (which the output schema already handles). An agent lacks enough context to set dimensions, format, or supply a data_file.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the description must compensate and explains only the required 'variable' (continuous). It says nothing about dpi, width_px/height_px, title, x_label/y_label, data_file, or image_format, leaving most parameters undocumented in prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Export a publication-ready histogram (PNG/TIFF)'. This is clearly distinguishable from scatter/bar/line siblings. However, it fails to differentiate from the nearest sibling spss_chart_histogram_density, which an agent must choose between; that omission holds it back from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Only the terse directive 'Provide the continuous variable' hints at usage; there is no statement of when to use this versus spss_chart_histogram_density or other chart tools, no prerequisites (e.g., a loaded dataset/session), and no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_chart_histogram_densitySpss Chart Histogram DensityA
Export a publication-ready histogram overlaid with a normal density curve (PNG/TIFF) via GGRAPH + OMS. Provide a continuous variable.
| Name | Required | Description | Default |
|---|---|---|---|
| dpi | No | ||
| title | No | ||
| x_label | No | ||
| y_label | No | ||
| variable | Yes | ||
| width_px | No | ||
| data_file | No | ||
| height_px | No | ||
| image_format | No | PNG |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It discloses the output artifact type and formats and the underlying GGRAPH + OMS mechanism, which is useful. However it says nothing about authentication, file destination, overwrite behavior, or what happens if the variable is non-continuous.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core purpose and output format front-loaded, followed by the key input constraint. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. But for a 9-parameter chart-export tool with 0% schema coverage and no annotations, the definition leaves the majority of tunable behavior undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the description must compensate, but it only identifies the required 'variable' parameter. Nine configurable options (dpi, width_px, height_px, image_format, title, x_label, y_label, data_file) are left entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Export), the resource (histogram overlaid with a normal density curve), the output formats (PNG/TIFF), and the implementation backend (GGRAPH + OMS). The 'density curve' overlay clearly distinguishes it from the sibling spss_chart_histogram.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Provide a continuous variable' gives one usage constraint, implying this is for continuous data. But there is no guidance on when to choose this over spss_chart_histogram or the other chart siblings, and no when-not-to-use conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_chart_km_curveSpss Chart Km CurveC
Export a publication-ready Kaplan-Meier survival curve (PNG/TIFF) via the KM procedure. Provide time and status variables; group is optional.
| Name | Required | Description | Default |
|---|---|---|---|
| dpi | No | ||
| time | Yes | ||
| event | No | ||
| group | No | ||
| title | No | ||
| status | Yes | ||
| width_px | No | ||
| data_file | No | ||
| height_px | No | ||
| image_format | No | PNG |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses that the output is a publication-ready file in PNG/TIFF, but says nothing about where the file is written, whether it overwrites, what permissions or data setup are required, or how failures surface. The bulk of behavioral context is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the artifact and the required inputs front-loaded and no filler. Brevity is reasonable, though given ten undocumented parameters the terseness edges toward under-specification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but with zero annotation coverage and 0% schema description coverage on a 10-parameter file-producing tool, the description leaves too much unstated (data_file, resolution/size controls, event coding) for an agent to invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 10 parameters, so the description must compensate, yet it only touches time, status, group, and the PNG/TIFF format. It leaves dpi, event (and its event-code value), title, width_px, height_px, and especially data_file entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('Export a publication-ready Kaplan-Meier survival curve') and names the underlying KM procedure, so the agent knows exactly what artifact is produced. It does not, however, distinguish itself from the sibling spss_kaplan_meier analysis tool or the other spss_chart_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Provide time and status variables; group is optional' is parameter guidance rather than usage guidance. There is no statement of when to pick this tool over spss_kaplan_meier or another chart sibling, and no prerequisites or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_chart_lineSpss Chart LineC
Export a publication-ready line chart (PNG/TIFF) via GGRAPH + OMS IMAGE. Provide the x (time/ordinal) and y variables; the file path is returned.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| dpi | No | ||
| title | No | ||
| x_label | No | ||
| y_label | No | ||
| width_px | No | ||
| data_file | No | ||
| height_px | No | ||
| image_format | No | PNG |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden; it discloses the underlying mechanism (GGRAPH + OMS IMAGE), the format options, and that a file path is returned. However, it omits permissions, overwrite behavior, output location control, and validation constraints, leaving meaningful behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action and format. Little waste, though the parenthetical format list and GGRAPH/OMS jargon add minor noise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained, but for a 10-parameter chart tool with 0% schema coverage the description is far too thin. It omits sizing, labeling, resolution, and data-source parameters an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 10 parameters, yet the description explains only x and y. Eight parameters (dpi, title, x_label, y_label, width_px, height_px, data_file, image_format) are undocumented in both schema and description, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Export) and resource (line chart), plus the output formats PNG/TIFF, so the agent can distinguish it from spss_chart_scatter or spss_chart_bar by the chart type. It stops short of explicit sibling differentiation, but the chart type is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It says to provide x and y variables but gives no guidance on when to prefer a line chart over the sibling chart tools or what prerequisites are needed. Usage is only implied by the input hint.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_chart_qqplotSpss Chart QqplotB
Export a publication-ready normal Q-Q plot (PNG/TIFF) of a continuous variable via PPLOT, captured through the OMS chart pipeline.
| Name | Required | Description | Default |
|---|---|---|---|
| dpi | No | ||
| title | No | ||
| variable | Yes | ||
| width_px | No | ||
| data_file | No | ||
| height_px | No | ||
| image_format | No | PNG |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the underlying mechanism (PPLOT via the OMS chart pipeline) and that output is publication-ready, but omits where the exported file is written, how it is returned, whether an active dataset or data_file is required, and any permission/latency characteristics. For an export tool this leaves key operational behavior undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, densely informative sentence with the action and artifact front-loaded; no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but for a 7-parameter export tool with no annotations and 0% schema coverage the description is too thin — it omits parameter behaviors, dataset requirements, and where artifacts land. An agent would have to guess how to configure and consume the call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 7 parameters, so the description must compensate. It only hints at 'variable' (must be continuous) and image format (PNG/TIFF); dpi, width_px, height_px, title, and data_file are entirely undocumented anywhere, leaving more than half the parameters without semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Export') and a specific resource ('normal Q-Q plot of a continuous variable'), with output formats (PNG/TIFF) named inline. This is plainly distinguishable from the sibling chart tools (histogram, scatter, boxplot, etc.) without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never says when to prefer this Q-Q plot over alternatives such as spss_normality_outliers, spss_chart_histogram_density, or spss_read_data plus external plotting. No prerequisites, no exclusions, no routing guidance; the usage context is only inferable from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_chart_scatterSpss Chart ScatterB
Export a publication-ready scatter plot (PNG/TIFF) via GGRAPH + OMS IMAGE. Provide x and y variables; the chart file path is returned for direct submission.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| dpi | No | ||
| title | No | ||
| x_label | No | ||
| y_label | No | ||
| width_px | No | ||
| data_file | No | ||
| height_px | No | ||
| image_format | No | PNG |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries more burden. It discloses useful implementation context (GGRAPH + OMS IMAGE) and that a chart file path is returned, but it does not mention permissions, whether existing files are overwritten, or other operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and output, with no wasted words. The essential information is immediately clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 10 parameters and no annotations, the description is too sparse. It covers the basic required inputs and output path, but leaves most parameter meanings and behavioral details undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate but only mentions x and y. It adds no meaning for the other 8 parameters such as dpi, title, labels, dimensions, data_file, or image_format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Export), a clear resource (scatter plot), and the output formats (PNG/TIFF). It distinguishes this tool from sibling chart tools like spss_chart_histogram or spss_chart_bar by naming the chart type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It says to provide x and y variables, but gives no guidance on when to choose a scatter plot over other chart types or any preconditions. Usage is only implied by the tool name and required inputs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_check_statusSpss Check StatusA
Check the SPSS MCP server status: which capabilities are available (SPSS installed vs file-only mode), SPSS path, library versions, and configuration. Call this first to understand what tools are available.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden and does disclose the key behavioral insight: whether SPSS is fully installed or the server is in file-only mode, which fundamentally changes what other tools can do. It does not explicitly state that the call is side-effect-free or local, leaving a small gap, but the returned information is well described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core status-check purpose and then the when-to-call guidance. Every phrase earns its place without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a zero-parameter tool, no annotations, and an output schema that documents return values, the description covers everything needed: what is checked, why it matters (installed vs file-only), and when to call it. Nothing required for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The empty schema means there is no parameter semantics to document, and the description correctly focuses on return content instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Check the SPSS MCP server status') and enumerates exactly what is reported: available capabilities, installed vs file-only mode, SPSS path, library versions, and configuration. This distinguishes it from all sibling tools, which perform analysis rather than environment introspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Call this first to understand what tools are available,' giving a clear invocation context. No alternative sibling competes for this purpose, so no exclusion guidance is needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_cluster_hierarchicalSpss Cluster HierarchicalC
Run hierarchical cluster analysis with dendrogram. Supports multiple linkage methods and distance measures. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | WARD | |
| measure | No | SEUCLID | |
| file_path | Yes | ||
| variables | Yes | ||
| dendrogram | No | ||
| id_variable | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the external dependency ('Requires IBM SPSS Statistics to be installed'), but says nothing about whether this executes syntax that mutates the session, what permissions are needed, or how execution failures surface.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action. Efficient, though the dependency note could be positioned after the capability details for better priority ordering.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not required. For a fairly complex 6-parameter statistical tool with zero schema documentation, however, the description leaves method/measure selection and variable input expectations unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 6 parameters. The description gestures at 'multiple linkage methods and distance measures' (method, measure) and the dendrogram flag, but adds no enum semantics, no explanation of file_path/variables/id_variable, and no indication of defaults or required fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Run hierarchical cluster analysis with dendrogram.' This clearly identifies the statistical procedure and distinguishes it from siblings like spss_twostep_cluster and spss_factor, though it never names the differentiating sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to choose hierarchical clustering over spss_twostep_cluster, nor prerequisites such as variable type requirements or when linkage/distance choices should be varied. The reader must infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_compute_scale_scoreSpss Compute Scale ScoreC
Compute a scale score (SUM or MEAN) from multiple item variables, with optional reverse coding and minimum valid item count. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | ||
| method | No | mean | |
| file_path | Yes | ||
| min_valid | No | ||
| reverse_max | No | ||
| reverse_min | No | ||
| new_variable | Yes | ||
| reverse_items | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions the SPSS installation requirement, but does not say whether the tool modifies the source file, writes a new variable to disk, or how reverse coding and min_valid interact with the computation. Side effects and failure behavior are left unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words, front-loading the core action before the prerequisite. Both sentences earn their place and the text is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters, 0% schema description coverage, no annotations, and a composite-scoring operation that likely mutates data, the description is too thin. Because an output schema exists, return values need not be explained, but critical operational details about parameter meaning and side effects are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only clarifies method (SUM or MEAN), reverse coding, and minimum valid item count. It does not explain file_path, new_variable, items, reverse_items, or the reverse_min/reverse_max pair, leaving most of the 8 parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Compute a scale score (SUM or MEAN) from multiple item variables') and adds scope details like reverse coding and minimum valid item count. It does not explicitly distinguish itself from siblings such as spss_reliability_alpha or spss_run_syntax, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no exclusions, and no named alternatives. The only contextual note is the prerequisite that IBM SPSS Statistics must be installed, which is not usage guidance. An agent must infer when this tool is preferred over other SPSS procedures.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_correlationsSpss CorrelationsC
Run SPSS CORRELATIONS to compute Pearson or Spearman correlation matrix. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | pearson | |
| file_path | Yes | ||
| variables | Yes | ||
| two_tailed | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the external SPSS dependency, which is genuinely useful, but says nothing about runtime behavior, side effects, error handling, or whether the underlying data is modified. For a computation tool with zero annotation coverage this is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and the method choice, with no filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but with no usage guidance and 0% parameter documentation the agent lacks enough to invoke this correctly beyond guessing at file_path/variables semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across four parameters. The description mentions Pearson/Spearman, but that merely restates the enum already present in the schema; file_path, variables, and two_tailed receive no explanation, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: run SPSS CORRELATIONS and compute a Pearson or Spearman correlation matrix. This clearly distinguishes it from regression, crosstabs, and other statistical siblings in the list, though it does not name an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only guidance is the prerequisite that IBM SPSS Statistics must be installed. There is no indication of when to choose correlations over, say, regression or crosstabs, nor any exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_cox_regressionSpss Cox RegressionC
Run Cox proportional hazards regression for survival analysis. Supports time-dependent covariates, stratification, and model diagnostics. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | ENTER | |
| strata | No | ||
| file_path | Yes | ||
| predictors | Yes | ||
| categorical | No | ||
| save_survival | No | ||
| time_variable | Yes | ||
| status_variable | Yes | ||
| status_event_value | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It notes the SPSS dependency and mentions feature support (time-dependent covariates, stratification, diagnostics), but says nothing about whether files are read or written, side effects of save_survival, or runtime/permission constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core purpose, then capabilities, then the environment requirement. No filler, though it is arguably too terse given the parameter surface.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but for a 9-parameter statistical tool with zero annotation coverage and zero schema descriptions, the description is far too thin on parameter meaning, prerequisites, and behavioral expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% with 9 parameters and none documented in the schema. The description only loosely gestures at strata ('stratification') and diagnostics, leaving file_path, time_variable, status_variable, status_event_value, predictors, method, categorical, and save_survival entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run Cox proportional hazards regression for survival analysis'), which clearly separates it from sibling tools like spss_logistic_regression and spss_kaplan_meier. It does not explicitly name an alternative, but the method name is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to choose this over spss_kaplan_meier (non-parametric survival) or spss_logistic_regression, nor any stated prerequisites beyond SPSS being installed. The agent must infer the use case from the method name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_crosstabsSpss CrosstabsA
Run SPSS CROSSTABS to create a contingency table between two categorical variables. Optionally includes chi-square test and row/column percentages. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| row_variable | Yes | ||
| column_variable | Yes | ||
| include_col_pct | No | ||
| include_row_pct | No | ||
| include_chisquare | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the external dependency ('Requires IBM SPSS Statistics to be installed') and the optional outputs (chi-square, row/column percentages), but it does not state whether the input file is modified, whether the operation is read-only, or what error conditions look like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, followed by optional behavior and the prerequisite. Every sentence adds distinct information, and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. The description covers the tool's purpose, optional outputs, and dependency, which is largely sufficient for a moderately complex statistical procedure. The main gap is that it does not describe the expected input file format or whether defaults apply.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains that row_variable and column_variable are categorical variables and maps include_chisquare, include_row_pct, and include_col_pct to chi-square and percentage outputs. It does not clarify file_path format or that the boolean options default to true, but the core parameter meanings are covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Run SPSS CROSSTABS') and resource ('contingency table between two categorical variables'), making the tool's function immediately clear. However, it does not explicitly distinguish itself from statistical siblings such as spss_frequencies or spss_descriptives, which prevents a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the stated scope: use when analyzing two categorical variables. The prerequisite ('Requires IBM SPSS Statistics to be installed') is noted, but there is no explicit guidance on when to choose this over alternatives like spss_frequencies or spss_nonparametric_tests, and no exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_descriptivesSpss DescriptivesB
Run SPSS DESCRIPTIVES for numeric variables. Returns N, mean, std deviation, min, max, and optional statistics. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| variables | Yes | ||
| statistics | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry behavioral disclosure and does add real value: it lists the returned statistics and flags that IBM SPSS Statistics must be installed. It does not cover error behavior, how missing values or non-numeric variables are handled, or whether the call mutates state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the action and scope front-loaded, then output and the environment prerequisite. No filler, though the return-value listing partially duplicates the output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be restated, and the environment prerequisite is a useful addition. However, with three undocumented parameters and no annotations, the definition leaves gaps an agent would need to resolve before calling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters. The description hints at the 'statistics' parameter being optional and mentions numeric variables, but 'file_path' (a required parameter) is never explained and no value/format guidance is given for any parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run SPSS DESCRIPTIVES') scoped to 'numeric variables', which helps an agent distinguish it from frequency/crosstab procedures. It stops short of explicitly naming sibling alternatives like spss_frequencies or spss_file_summary, so differentiation relies on inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for numeric variables' implies when the tool applies, but there is no explicit when-to-use versus alternatives such as spss_frequencies or spss_file_summary, and no prerequisites about data state. An agent must infer the routing decision.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_discriminantSpss DiscriminantC
Run discriminant analysis to classify cases into groups. Supports stepwise selection and cross-validation. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| groups | Yes | ||
| method | No | DIRECT | |
| priors | No | EQUAL | |
| file_path | Yes | ||
| predictors | Yes | ||
| save_class | No | ||
| save_scores | No | ||
| group_values | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the SPSS dependency and mentions stepwise selection and cross-validation, but says nothing about whether save_class/save_scores write files, permission needs, or side effects of a statistical write operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no waste. The requirement is stated early and capabilities follow, though there is room for a brief field-level clarification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter statistical tool with zero schema description coverage and no annotations, the description is far too thin. The output schema covers return values, but parameter meaning and mutation side effects are left undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 8 parameters, so the description must compensate and does not. It gestures at method (stepwise) and implies groups/predictors, but leaves method enums, priors, group_values, and the save_* flags entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run discriminant analysis') plus the goal ('classify cases into groups'). Discriminant analysis is a distinct method among the statistical siblings, though the description never explicitly distinguishes it from logistic_regression or manova.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, prerequisites for the data, or alternatives among siblings. It only states the environment prerequisite (SPSS installed), not when this method is preferred over others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_factorSpss FactorC
Run SPSS FACTOR analysis (principal components or principal axis factoring). Includes eigenvalues, variance explained, and rotated factor matrix. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | PC | |
| rotation | No | VARIMAX | |
| file_path | Yes | ||
| n_factors | No | ||
| variables | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It usefully states the external dependency (IBM SPSS Statistics required) and lists the result contents (eigenvalues, variance explained, rotated factor matrix), but it omits whether the operation is read-only, any file/output side effects, or performance considerations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action and result, with the installation prerequisite stated last. No filler, though the return-value sentence is somewhat redundant given the output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter statistical procedure with no annotations and no schema parameter descriptions, the definition is incomplete. It leaves required parameters (file_path, variables) and n_factors unexplained, so an agent cannot confidently invoke it beyond the method choice.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It hints at the method parameter by naming principal components and principal axis factoring, but it does not explain file_path, variables, n_factors, or the rotation parameter's options and effects.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Run SPSS FACTOR analysis') and names the two factor extraction methods. It is clear what the tool does, though it does not explicitly distinguish itself from siblings such as spss_cluster_hierarchical or spss_discriminant.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites beyond the SPSS installation, and no indication of when to choose this over other dimension-reduction or multivariate siblings. Usage is only implied by the tool name and the mention of factor analysis.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_file_summarySpss File SummaryA
Get a summary of an SPSS .sav file: case count, variable count, variable list, and basic descriptive statistics computed locally (no SPSS needed). Does not require SPSS to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one meaningful behavioral trait: the summary is computed locally and does not require an SPSS installation. That is genuinely useful for a suite of SPSS tools that presumably need a backend. It says nothing about error behavior for a missing or malformed file, performance on large .sav files, or any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the operation and its outputs in one dense sentence. The trailing sentence restates 'no SPSS needed' from the parenthetical, which is mild redundancy, but overall it is short and earns most of its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so enumerating return contents goes beyond necessity, and nothing critical is missing for a single-parameter read tool. Gaps like behavior on unreadable files or large-file performance are minor for this complexity level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single file_path parameter, so the description must compensate. It partially does by implying the argument is a path to an SPSS .sav file, which constrains the expected format, but it adds no guidance on absolute vs relative paths or how the file is located.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get a summary of an SPSS .sav file') and enumerates exactly what is returned: case count, variable count, variable list, and basic descriptive statistics. It is clear enough to distinguish from deep-analysis siblings, though it does not explicitly contrast with overlapping tools like spss_read_metadata or spss_descriptives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the content list (a fast file-level overview usable before deeper analysis), and the 'computed locally' note hints at when it is appropriate. However, no alternative sibling is named and no condition for choosing this over spss_read_metadata or spss_descriptives is stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_frequenciesSpss FrequenciesB
Run SPSS FREQUENCIES on one or more variables. Returns frequency tables with counts, percentages, and optional statistics. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| variables | Yes | ||
| statistics | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses an environment prerequisite ('Requires IBM SPSS Statistics to be installed') and the return shape, but says nothing about error behavior, whether file_path is read-only, or how optional statistics affects execution.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler, and the core action is front-loaded before the return description and prerequisite. Nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained in depth, and the prerequisite is a useful addition. But with 0% schema coverage and no annotations, the missing explanation of file_path and the absence of any alternative-tool guidance leave notable gaps for a 3-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It partially does: 'one or more variables' clarifies the variables param accepts multiple, and 'optional statistics' flags the third param as non-required. However, file_path is never explained, leaving the most important parameter undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run SPSS FREQUENCIES on one or more variables') and describes the output ('frequency tables with counts, percentages, and optional statistics'). It does not distinguish itself from close siblings like spss_descriptives or spss_crosstabs, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The implied use case is clear (getting frequency tables for one or more variables), but there is no explicit when-to-use guidance, no exclusions, and no routing toward the very similar spss_descriptives or spss_crosstabs siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_genlinSpss GenlinC
Run generalized linear model (GENLIN) with flexible distribution and link functions. Supports Poisson, binomial, gamma, negative binomial, and other distributions. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| link | No | ||
| scale | No | ||
| dependent | Yes | ||
| file_path | Yes | ||
| predictors | Yes | ||
| categorical | No | ||
| distribution | No | NORMAL | |
| save_predicted | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the external dependency ('Requires IBM SPSS Statistics to be installed'), but says nothing about whether files are modified, how save_predicted behaves, or side effects. Partial behavioral context only.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action. No wasted phrasing, though the final sentence is a practical constraint rather than descriptive filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 8 parameters, zero schema descriptions, and no annotations, the description is too thin — the agent gets no help mapping its arguments. The output schema covers return values, but the input side is essentially undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% — no parameter has any description. The description names distributions (which merely restates the existing enum) and alludes to link functions, but leaves file_path, dependent, predictors, categorical, scale, and save_predicted entirely unexplained. It does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb and resource ('Run generalized linear model (GENLIN)') and characterizes the scope via flexible distribution/link functions. It does not explicitly differentiate from close siblings like spss_glm_univariate, spss_logistic_regression, or spss_genlinmixed, so the agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance or alternatives are given. In a family crowded with regression models (logistic, ordinal, glm_univariate, genlinmixed), the absence of routing hints is a real gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_genlinmixedSpss GenlinmixedC
Run generalized linear mixed model combining GLM with random effects. Supports non-normal outcomes with hierarchical structure. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| link | No | ||
| subject | No | ||
| dependent | Yes | ||
| file_path | Yes | ||
| distribution | No | NORMAL | |
| fixed_effects | Yes | ||
| random_effects | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It contributes one operational fact, the SPSS installation requirement, but says nothing about side effects, permissions, convergence behavior, or whether data is modified. With no annotations and a complex statistical operation, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three concise sentences with the core purpose front-loaded. The capability note and installation prerequisite each add distinct value. It is efficient, though it could be slightly more informative without becoming bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 7-parameter statistical tool with no schema descriptions and no annotations, the description is incomplete. It gives purpose and an installation prerequisite, but omits parameter semantics and usage guidance versus related tools like spss_genlin and spss_mixed. The output schema covers return values, but invocation details remain underspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no information about any of the 7 parameters. Terms like file_path, dependent, fixed_effects, random_effects, distribution, link, and subject are not clarified or mapped to their meaning. The description does not compensate for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Run generalized linear mixed model combining GLM with random effects.' It adds distinguishing scope with 'non-normal outcomes' and 'hierarchical structure,' which helps separate it from spss_genlin and spss_mixed. It does not explicitly name those siblings, so it falls short of full 5-level differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: for hierarchical data with non-normal outcomes. It also states the prerequisite that IBM SPSS Statistics must be installed. However, it gives no explicit when-not guidance or alternatives such as spss_genlin or spss_mixed, leaving selection logic largely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_get_method_schemaSpss Get Method SchemaA
Get the JSON schema for a registry-backed SPSS method. Useful for structured orchestration and parameter inspection before execution.
| Name | Required | Description | Default |
|---|---|---|---|
| tool_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden, and it does convey that the method must be registry-backed and that this is a read-style inspection. However it omits error behavior for unknown method names, and says nothing about the safety/permission profile beyond the implied read-only nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and followed by the use case. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and this is a simple one-parameter read tool. Still, the definition leaves the agent without guidance on valid method identifiers or the relationship to sibling lookup tools, which is the main gap given no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single parameter (tool_name) has no description in the schema. The description hints the argument is a registered method name, but it never states the expected format or how to discover valid values (e.g., via spss_list_supported_methods), so it only partially compensates for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Get the JSON schema for a registry-backed SPSS method") that an agent can distinguish from siblings like spss_list_supported_methods and spss_get_method_support. It is clear what it returns, though it never explicitly contrasts itself with those adjacent tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use it ("before execution", "parameter inspection", "structured orchestration"), so the agent knows this is a pre-flight inspection step. It does not name alternatives or state when not to use it, keeping it just below a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_get_method_supportSpss Get Method SupportB
Get support metadata for a registry-backed SPSS method, including command family, support tier, coverage assertions, and documentation tags.
| Name | Required | Description | Default |
|---|---|---|---|
| tool_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the return content categories and 'Get' implies a read-only, side-effect-free lookup, which is useful. It does not state any auth requirements, error behavior on unknown method names, or whether an unknown tool_name returns empty vs. an error.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the verb and scope front-loaded and no filler. Slightly dense in its list of returned categories, but every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. For a simple single-param lookup this is close to adequate, but the lone parameter is undocumented in both schema and description, and there is no sibling routing guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One required parameter (tool_name) with 0% schema description coverage, and the description never explains it. The phrase 'registry-backed SPSS method' implies tool_name is a registered method identifier, but format, allowed values, and how to discover valid names are left undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get support metadata for a registry-backed SPSS method') and enumerates the metadata categories returned (command family, support tier, coverage assertions, documentation tags). It is distinguishable from siblings like spss_get_method_schema and spss_list_supported_methods, though the boundary is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus spss_list_supported_methods, spss_get_method_schema, or spss_read_metadata. There are no prerequisites or exclusion conditions, leaving the agent to guess from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_glm_univariateSpss Glm UnivariateB
Run univariate general linear model (GLM) with factorial designs. Supports estimated marginal means, contrasts, and post-hoc tests. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| emmeans | No | ||
| factors | Yes | ||
| posthoc | No | ||
| dependent | Yes | ||
| file_path | Yes | ||
| covariates | No | ||
| posthoc_method | No | ||
| save_predicted | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose one useful environmental prerequisite (SPSS installation). It says nothing about side effects such as file mutation or the save_predicted parameter writing new columns, nor about output format or failure modes, leaving meaningful behavioral gaps for a statistical procedure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no filler; the core capability leads and the prerequisite trails. Efficiently sized for the information offered, though some brevity comes at the cost of missing parameter detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. But with zero annotation coverage, zero schema descriptions, and 8 parameters, the description is not complete enough for reliable invocation – it omits argument semantics and any sibling differentiation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 8 parameters, so the description must compensate. It loosely maps to emmeans, posthoc/posthoc_method, and factors ('factorial designs'), but leaves file_path, dependent, covariates, and save_predicted completely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run univariate general linear model (GLM) with factorial designs') and names the capabilities it covers (EMMs, contrasts, post-hoc). It is distinguishable from multivariate siblings like spss_manova, though it never explicitly contrasts itself with spss_anova or spss_genlin, which an agent could easily confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implied usage is present – factorial designs plus EMMs/contrasts/post-hoc signal when the tool is appropriate, and it notes the prerequisite that IBM SPSS Statistics must be installed. However, it names no alternative sibling (spss_anova, spss_genlin, spss_mixed) or the conditions that would route an agent elsewhere.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_import_csvSpss Import CsvA
Convert a CSV file to SPSS .sav format directly using pandas + pyreadstat — no IBM SPSS Statistics installation required. Much faster than going through SPSS syntax because it bypasses the SPSS engine entirely. Saves the .sav file next to the CSV by default, or to a custom output_path.
| Name | Required | Description | Default |
|---|---|---|---|
| csv_path | Yes | ||
| encoding | No | utf-8 | |
| delimiter | No | , | |
| output_path | No | ||
| column_labels | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the mechanism (pandas + pyreadstat) and the default output location ('saves the .sav file next to the CSV'), but it omits key side-effect details such as overwrite behavior, error handling, and any permissions or validation requirements for a file-writing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and uses three efficient sentences. The second sentence is a useful usage comparison rather than fluff, though it has a slight marketing tone; overall it is well-structured and not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given five parameters, no annotations, and 0% schema description coverage, the description is incomplete. It covers the conversion purpose and output location, but fails to explain key parameters like encoding, delimiter, and column_labels, and does not address overwrite or error behavior, so an agent must guess about important call details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description only partially compensates. It explains the default behavior of output_path ('saves the .sav file next to the CSV by default, or to a custom output_path') but says nothing about encoding, delimiter, column_labels, or even the required csv_path, leaving four of five parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Convert a CSV file to SPSS .sav format directly using pandas + pyreadstat.' It also distinguishes the tool from sibling operations by noting it bypasses the SPSS engine and requires no IBM SPSS installation, so an agent can tell it apart from spss_run_syntax or analysis tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear alternative and context: 'Much faster than going through SPSS syntax because it bypasses the SPSS engine entirely.' This implicitly tells the agent to use this tool for CSV-to-SAV conversion instead of relying on SPSS syntax, but it does not explicitly state when not to use the tool or other alternatives among the siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_kaplan_meierSpss Kaplan MeierC
Run Kaplan-Meier survival analysis with log-rank test. Produces survival curves and compares groups. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| strata | No | ||
| file_path | Yes | ||
| percentiles | No | ||
| time_variable | Yes | ||
| compare_method | No | LOGRANK | |
| status_variable | Yes | ||
| status_event_value | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully states a dependency ('Requires IBM SPSS Statistics to be installed') and that survival curves are produced, but says nothing about permissions, computational cost, failure modes for invalid data, or how mutation-like behavior is handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core action before the dependency note. No filler, though the brevity here reflects under-specification rather than disciplined economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. However, for a 7-parameter analytical tool with 0% schema coverage and zero annotations, the description is materially incomplete: an agent cannot tell what value status_event_value expects or what strata/percentiles do.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 7 parameters, so the description must compensate yet it adds almost nothing. Only 'log-rank test' loosely maps to compare_method and 'compares groups' loosely maps to strata; the required time_variable, status_variable, and status_event_value, plus percentiles, are entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Run Kaplan-Meier survival analysis with log-rank test,' which is a distinct statistical method. It clearly distinguishes itself by naming the method, though it does not explicitly contrast with the nearby sibling spss_cox_regression or spss_chart_km_curve.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says it 'compares groups' but gives no when-to-use guidance, no prerequisites beyond the SPSS install, and no comparison to comparable siblings like Cox regression or the KM chart tool. A user would have to infer that this performs the analysis whereas spss_chart_km_curve only plots.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_list_filesSpss List FilesB
List SPSS .sav files in a directory. Useful for discovering available datasets when the user hasn't specified a file path.
| Name | Required | Description | Default |
|---|---|---|---|
| directory | Yes | ||
| recursive | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden but only implies this is a safe read/discovery operation. It says nothing about recursion behavior, path resolution, permission requirements, or what happens with an empty/invalid directory. The safety profile is inferable but not disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, zero waste, with the core action front-loaded and the usage hint following. Nothing extraneous.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained, and the purpose is adequately conveyed for a simple listing tool. But both parameters are undocumented and there are no annotations, leaving meaningful gaps around the 'recursive' behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and largely fails. It implies the directory is a filesystem location but adds no format details (absolute vs relative path), and the 'recursive' parameter is entirely unaddressed in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (.sav files in a directory), which is clearly distinct from sibling tools like spss_list_variables or spss_list_supported_methods. It reads well without opening the schema, though it doesn't explicitly name a sibling to differentiate against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a concrete when-to-use condition: 'discovering available datasets when the user hasn't specified a file path.' That is real routing context. However, it names no alternatives and gives no when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_list_supported_methodsSpss List Supported MethodsA
List registry-backed SPSS methods available for structured execution. Use this to discover cold methods that have schemas, templates, and coverage assertions.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a safe read-only listing and discloses that returned methods come with schemas, templates, and coverage assertions, but it never states that it is non-mutating, costs nothing, or what the response shape contains—minor gaps for a discovery tool whose output schema covers the return format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with the core purpose and the intended use front-loaded, and no filler. The undefined jargon 'cold methods' costs a little clarity but the structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument listing tool with an output schema, the description needn't explain return values, and it correctly signals the discovery role. The only residual gap is the unexplained 'cold methods' term and the absence of sibling routing, which is minor given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies. No misleading parameter guidance is present.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List registry-backed SPSS methods') and hints at the differentiating scope ('available for structured execution'), which separates it from the analysis siblings. However, it doesn't explicitly name or contrast with the closest siblings spss_get_method_schema and spss_get_method_support, so an agent must infer the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives implied usage ('Use this to discover cold methods...'), which suggests a discovery step before execution. But it offers no explicit when-not guidance and never names alternatives like spss_get_method_schema, so routing between the three method-related tools is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_list_variablesSpss List VariablesB
List all variable names and their labels from an SPSS .sav file. Optionally filter by a search term. Does not require SPSS to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| search | No | ||
| file_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses that SPSS does not need to be installed, and 'List' implies a read-only, side-effect-free operation, but it says nothing about file access errors, encoding, or performance on large .sav files.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, with no filler or redundancy. Every clause contributes information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is a simple read operation with an output schema, so return values need no explanation. The description is nearly sufficient for correct invocation; only the absence of sibling differentiation and error behavior keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the schema supplies no parameter meaning on its own. The description compensates for the 'search' parameter by explaining it acts as an optional filter, and 'from an SPSS .sav file' implies the file_path argument, but neither parameter's format or expected values are specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'List all variable names and their labels from an SPSS .sav file.' It is clear what the tool returns, but it does not differentiate itself from close siblings like spss_read_metadata or spss_file_summary, which plausibly also surface variable metadata.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never says when to choose this over spss_read_metadata, spss_read_data, or spss_file_summary. The only conditional content ('Optionally filter by a search term') describes a parameter, not a usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_logistic_regressionSpss Logistic RegressionB
Run binary or multinomial logistic regression. Supports stepwise selection, categorical predictors, and model diagnostics. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | ENTER | |
| contrast | No | ||
| dependent | Yes | ||
| file_path | Yes | ||
| predictors | Yes | ||
| categorical | No | ||
| print_options | No | ||
| save_predicted | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the supported modeling features and the environment requirement, which is real value, but says nothing about side effects (whether save_predicted writes new variables into the dataset, whether output files are created), permissions, or failure modes for stepwise selection.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no filler; the capability statement leads and the environment prerequisite trails. It is efficient, though it stops short of using the remaining space for the parameter guidance this tool actually lacks.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. But for an 8-parameter statistical procedure with zero schema descriptions, the description leaves the agent without enough information to call it correctly (e.g., accepted values for contrast or print_options, meaning of file_path, behavior of save_predicted).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 8 parameters, so the schema documents nothing. The description only loosely gestures at three of them (stepwise -> method, categorical -> categorical, diagnostics -> print_options); contrast, file_path, dependent, predictors, and save_predicted get no meaning, format, or accepted values, so the agent cannot reliably populate them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('Run binary or multinomial logistic regression'), and the 'binary or multinomial' qualifier distinguishes it from the sibling spss_ordinal_regression and from the generic spss_regression. An agent can route to this tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names supported features ('stepwise selection, categorical predictors, and model diagnostics') and states a hard prerequisite ('Requires IBM SPSS Statistics to be installed'), which implies when the tool is usable. However, it never says when to prefer this over spss_regression or spss_ordinal_regression, nor any exclusions or data requirements.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_manovaSpss ManovaC
Run multivariate analysis of variance (MANOVA) for multiple dependent variables. Tests multivariate effects and provides univariate follow-ups. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | SSTYPE3 | |
| factors | Yes | ||
| file_path | Yes | ||
| covariates | No | ||
| dependents | Yes | ||
| factor_ranges | No | ||
| print_univariate | No | ||
| print_multivariate | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses only the SPSS installation requirement and that both multivariate effects and univariate follow-ups are produced; it omits read-only vs. mutating nature, any permission or resource constraints, and runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the core purpose front-loaded and no filler. It is efficient, though slightly under-specified rather than truly concise in the informative sense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for an 8-parameter statistical tool with zero schema description coverage and no annotations the description is far too thin. Critical inputs like method, factors, covariates and range specifications are left entirely undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 8 parameters, and the description explains none of them. Nothing is said about method (SSTYPE1-4), factors, covariates, factor_ranges, or the print_univariate/print_multivariate toggles, so the agent must guess from parameter names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Run) and resource (multivariate analysis of variance / MANOVA) plus the key scope qualifier: multiple dependent variables. This implicitly separates it from spss_anova (single DV), but it never names that sibling or the several other ANOVA-family tools, so the distinction must be inferred.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a prerequisite (SPSS must be installed) but no when-to-use criteria and no alternatives. With spss_anova, spss_glm_univariate and spss_repeated_measures_anova as siblings, the agent gets no guidance on which to pick for a given data shape.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_mediationSpss MediationA
Run a Baron & Kenny three-step mediation analysis (X -> M -> Y) with regression, reporting paths a/b/c/c', the indirect effect a*b, and a Sobel test. Does not bundle the PROCESS macro.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| mediator | Yes | ||
| file_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the analytical method and its components (regression, indirect effect, Sobel test) and a scope limitation (no PROCESS macro), which is meaningful behavioral context. It omits preconditions such as variable type requirements, how missing data is handled, and whether the file must already be loaded, leaving gaps for a mutation-free but assumption-heavy analysis tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core method and outputs, followed by the scope caveat. No filler, no repetition of the tool name or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. But with zero annotations and zero schema coverage, the description should do more to cover invocation preconditions (data loading state, variable measurement level, assumption checks). It is adequate for identifying the method but thin on operational prerequisites.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all four required parameters, so the description must compensate. The 'X -> M -> Y' notation usefully maps x=independent, mediator=mediator, y=outcome, covering three of four params. file_path is left unexplained, though its intent is inferable from the spss_* family, so compensation is partial rather than complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific statistical procedure (Baron & Kenny three-step mediation, X -> M -> Y), enumerates the exact output paths (a/b/c/c', indirect effect a*b, Sobel test), and implicitly delineates itself from the sibling spss_moderation by virtue of being mediation rather than moderation. An agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The scope exclusion 'Does not bundle the PROCESS macro' is a useful boundary note, effectively routing users who want PROCESS bootstrapping elsewhere. However, it gives no positive when-to-use guidance relative to spss_regression or spss_moderation, and no prerequisites. Usage is only implied through the tool name and method labels.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_mixedSpss MixedC
Run linear mixed-effects model (multilevel model) with random effects. Supports nested and crossed random effects, repeated measures structures. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | REML | |
| subject | No | ||
| repeated | No | ||
| dependent | Yes | ||
| file_path | Yes | ||
| fixed_effects | Yes | ||
| repeated_type | No | ||
| covtype_random | No | ||
| random_effects | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it only discloses one operational prerequisite ('Requires IBM SPSS Statistics to be installed'). It says nothing about whether the tool writes files, how convergence/failure is handled, what preprocessing the data needs, or roughly what it returns — significant gaps for a 9-parameter modeling tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the model type, followed by supported structures and the runtime dependency. No filler, though the capability sentence and the installation sentence are the only substantive content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. But for a 9-parameter statistical tool with zero annotation coverage and zero schema descriptions, the definition leaves too much undefined — no parameter guidance, no usage context, and no behavioral detail beyond the SPSS dependency.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the schema documents only names, types, and one enum, leaving the description to compensate — and it largely does not. Terms like 'random effects' and 'repeated measures structures' loosely gesture at random_effects/repeated, but there is nothing on file_path, dependent, fixed_effects, subject, method (REML vs ML), repeated_type, or covtype_random.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run linear mixed-effects model (multilevel model) with random effects') and names supported structures (nested/crossed, repeated measures). An agent can identify it as the mixed-modeling tool. It does not, however, explicitly route against close siblings like spss_genlinmixed, spss_repeated_measures_anova, or spss_glm_univariate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, and no alternative is named. The capability phrases ('supports nested and crossed random effects') hint at applicable data shapes but never tell the agent when this tool is the right choice over the repeated-measures ANOVA or GENLINMIXED siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_moderationSpss ModerationA
Run a mean-centred moderation regression (Y ~ X + W + X*W) and report the interaction term that tests the moderating effect of W.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| file_path | Yes | ||
| moderator | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose one meaningful trait beyond the schema: that predictors are mean-centred (which differs from a plain regression) and which interaction term it reports. However, it says nothing about data prerequisites (e.g., an SPSS .sav file, variable scale requirements) or failure behavior, leaving notable gaps for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that specifies the method, the model equation, and the reported statistic with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the method is concisely specified. But with zero annotations and 0% parameter description coverage, the definition omits prerequisites such as the expected file type/path semantics and any data requirements, leaving it only minimally complete for a four-required-parameter analysis tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must define all four parameters. The formula maps X, W, and Y to x, moderator, and y well enough to infer roles, but file_path is never addressed and the string/variable-name nature of the inputs is left implicit. Partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (run), a specific model (mean-centred moderation regression Y ~ X + W + X*W), and the exact output of interest (the interaction term testing W's moderating effect). This clearly distinguishes it from siblings like spss_regression and spss_mediation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The statistical formulation implicitly signals when to use it (testing whether W moderates X→Y), but there is no explicit when-to-use vs when-not guidance and no named alternative such as spss_mediation for indirect-effect questions. Usage is inferable only from the method name and formula.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_nonparametric_testsSpss Nonparametric TestsC
Run common nonparametric tests in SPSS: Mann-Whitney U, Wilcoxon signed-rank, or Kruskal-Wallis. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| test_type | Yes | ||
| variables | Yes | ||
| group_values | No | ||
| grouping_variable | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses only the installation requirement; it says nothing about whether the source file is modified, what happens on invalid variable/group combinations, or any execution constraints. The presence of an output schema covers return values, but mutation and side-effect behavior remain opaque.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler, and the core capability is front-loaded before the prerequisite. Nothing is wasted, though the brevity is achieved partly by omitting information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter statistical tool with no annotations, zero schema coverage, and test-specific parameter requirements, the description is too thin. The output schema excuses it from explaining results, but it should at minimum explain the file_path/variables/grouping requirements that differ across the three tests.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and none of the five parameters are explained in the description. It never clarifies that file_path is an .sav path, what 'variables' should contain, or that grouping_variable/group_values are required for Mann-Whitney and Kruskal-Wallis but not Wilcoxon. Only the test_type enum is reflected, indirectly, by the listed test names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('run') and resource (nonparametric tests) and enumerates the three supported tests, matching the test_type enum. It implicitly separates itself from siblings like spss_t_test and spss_anova by labeling the family as nonparametric, though it never names those alternatives explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the environment prerequisite (SPSS installed) but gives no when-to-use guidance, no indication of when a nonparametric test is preferable to the parametric siblings (spss_t_test, spss_anova), and no conditions for choosing among the three offered tests. An agent must infer all routing decisions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_normality_outliersSpss Normality OutliersA
Run SPSS EXAMINE to check normality and outliers for numeric variables, with optional diagnostic plots. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| plots | No | ||
| file_path | Yes | ||
| variables | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses an environmental prerequisite ('Requires IBM SPSS Statistics to be installed') and that plots are optional, but it does not state whether the procedure mutates the file, what happens on missing values, or any performance or permissions characteristics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and target, with the prerequisite placed last. No wasted words, though it is slightly sparse given the tool's analytical complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers purpose, the plots option, and the environment requirement. It is nearly complete for an analysis tool, with the main gap being undocumented file_path semantics and no output-shape hint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It partially does: 'numeric variables' maps to the variables array and 'optional diagnostic plots' maps to the plots flag (consistent with its default of true), but file_path is never mentioned and no format or element-type guidance is given for variables.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Run), procedure (SPSS EXAMINE), and resource (normality and outliers for numeric variables), which is clear and actionable. However, it does not differentiate itself from closely related siblings like spss_descriptives or spss_chart_qqplot, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by naming the analytical goal (checking normality and outliers) and noting diagnostic plots are optional, but it gives no explicit when-to-use context, no exclusions, and no guidance on choosing it over spss_descriptives or the chart tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_ordinal_regressionSpss Ordinal RegressionB
Run ordinal regression (PLUM) for ordered categorical outcomes. Supports multiple link functions and parallel lines test. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| link | No | LOGIT | |
| dependent | Yes | ||
| file_path | Yes | ||
| predictors | Yes | ||
| categorical | No | ||
| test_parallel | No | ||
| save_predicted | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully adds the environmental prerequisite ('Requires IBM SPSS Statistics to be installed') and names supported features (multiple link functions, parallel lines test), but it omits key behavioral traits such as whether the tool modifies data, whether it saves predicted values when save_predicted=true, and how results are produced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The purpose is front-loaded, followed by capability and prerequisite details in a logical order.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 7 parameters, no schema descriptions, and no annotations, so the description must do more. It covers purpose and the SPSS installation requirement, and an output schema exists to explain return values, but it leaves parameter semantics, usage conditions, and side effects largely undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not compensate. It mentions 'multiple link functions' and 'parallel lines test' generically, but this adds no meaning beyond the enum and boolean already present in the schema. Required parameters (file_path, dependent, predictors) and optional parameters (categorical, save_predicted) receive no semantic explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Run') and resource ('ordinal regression (PLUM)') and scopes it to 'ordered categorical outcomes,' which distinguishes it from linear regression. However, it does not explicitly differentiate itself from close siblings like spss_logistic_regression or spss_regression, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for ordered categorical outcomes' implies when the tool is appropriate, but there is no explicit when-to-use guidance, no when-not-to-use exclusions, and no named alternatives such as spss_logistic_regression. Usage is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_read_dataSpss Read DataB
Read rows of data from an SPSS .sav file as a Markdown table. Optionally filter to specific variables and limit row count. Does not require SPSS to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| max_rows | No | ||
| file_path | Yes | ||
| variables | No | ||
| apply_value_labels | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the return format (Markdown table) and the notable constraint that SPSS need not be installed, which is real behavioral context. It does not say what happens on missing files, whether output is truncated when max_rows is hit, or any permission/environment caveats.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and output format, then options, then the environment caveat. Nothing is padded, though the middle sentence is somewhat thin relative to the gaps left elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is not required. However, for a 4-parameter tool with 0% schema coverage, the description leaves key behaviors (truncation signaling, value-label semantics) undocumented, so it is only minimally sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, so the description must compensate and largely does not. It explains the intent behind 'variables' and row limiting, but says nothing about max_rows' default of 50, its truncation semantics, or the meaning and effect of apply_value_labels (raw codes vs. labeled values).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read rows of data from an SPSS .sav file') plus the output format ('as a Markdown table'), which is more informative than the title restatement. It implicitly separates itself from metadata/summary siblings, but never names them, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The mention of optional variable filtering and row limiting implies when the tool is useful for targeted inspection, but there is no explicit when-to-use / when-not-to-use guidance and no reference to alternatives like spss_read_metadata or spss_file_summary that also read from a .sav file.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_read_metadataSpss Read MetadataB
Read variable names, types, labels, and value labels from an SPSS .sav file. Returns a detailed Markdown report of the file's structure. Does not require SPSS to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses that it is a read operation and that SPSS does not need to be installed, which is real behavioral context, but it omits failure modes (e.g., missing/invalid .sav file) and any permission or size considerations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and free of redundancy. Each sentence contributes (what it reads, what it returns, and the environment requirement).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-parameter read tool with an output schema that documents return values, the description covers the essentials: what is extracted and the no-SPSS-install requirement. Only the lack of sibling disambiguation keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description adds no meaning to file_path beyond its obvious role as the path to the .sav file. With a single self-evident required parameter, the gap is minor, so this lands at mid-range rather than low.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (variable names, types, labels, value labels from an SPSS .sav file) and specifies the return format (Markdown report of file structure). It is clear what the tool does, though it does not explicitly distinguish itself from close siblings like spss_list_variables or spss_file_summary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no exclusions, and no reference to neighboring tools such as spss_list_variables or spss_file_summary that could be confused with inspecting a file's structure. Usage is only implied by the tool name and purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_regressionSpss RegressionB
Run SPSS linear regression. Specify a dependent variable and one or more predictors. Returns coefficients, R-squared, ANOVA table, and significance tests. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | ENTER | |
| dependent | Yes | ||
| file_path | Yes | ||
| predictors | Yes | ||
| include_diagnostics | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the IBM SPSS Statistics dependency and the statistical outputs (coefficients, R-squared, ANOVA, significance tests), but says nothing about whether files are modified, permissions, pagination, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences with the action front-loaded and no filler. Efficient, though the return-value sentence is somewhat redundant given an output schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The existence of an output schema means the return-value details are redundant rather than necessary, and the missing coverage of 'method' and 'include_diagnostics' is a real gap for a 5-parameter statistical tool. Adequate but incomplete for its complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 parameters, so the description must compensate. It only clarifies 'dependent' and 'predictors' and never mentions 'method' (ENTER default), 'include_diagnostics', or 'file_path' semantics, leaving two parameters entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Run) and resource (SPSS linear regression) and identifies the required inputs (dependent variable, predictors). The word 'linear' implicitly distinguishes it from sibling spss_logistic_regression and spss_ordinal_regression, though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to pick this over siblings like spss_logistic_regression, spss_anova, or spss_glm_univariate, nor on method selection or file prerequisites. The only usage note is the environment requirement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_reliability_alphaSpss Reliability AlphaC
Run SPSS RELIABILITY analysis (Cronbach's alpha). Returns scale reliability and item statistics for psychometric workflows. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ALPHA | |
| file_path | Yes | ||
| variables | Yes | ||
| scale_name | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions the prerequisite 'Requires IBM SPSS Statistics to be installed' but omits whether the operation is read-only, what side effects it has, or how output is handled, leaving significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each front-loaded and purposeful: purpose, return content, and prerequisite. There is no wasted language or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be detailed, but the description is still incomplete for a tool with four undocumented parameters and no usage guidance. The prerequisite is helpful but insufficient to cover the missing parameter semantics and behavioral context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 4 parameters and 0% schema description coverage, the description fails to mention any of them. It does not explain file_path, variables, model, or scale_name, leaving the agent to infer their meaning solely from the schema structure.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Run) and resource (SPSS RELIABILITY analysis), clarifying it computes Cronbach's alpha. It does not explicitly distinguish itself from siblings like spss_factor or spss_compute_scale_score, so a 4 is appropriate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for psychometric workflows' hints at context but does not specify when to use this tool versus alternatives such as spss_factor or spss_compute_scale_score. No exclusions or prerequisites beyond the SPSS installation are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_repeated_measures_anovaSpss Repeated Measures AnovaB
Run SPSS repeated-measures ANOVA (within-subject GLM). Provide within-factor name, number of levels, and one variable per level. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| levels | Yes | ||
| file_path | Yes | ||
| variables | Yes | ||
| include_pairwise | No | ||
| within_factor_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It usefully discloses the hard requirement that IBM SPSS Statistics be installed, plus the shape of the analysis (within-subject). It does not say whether the tool writes output files, what permissions are needed, or how it behaves on unbalanced designs, and with no annotations that gap is notable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the operation and followed by the required inputs and prerequisite. No wasted words, though the parameter sentence is a bare list rather than explanatory.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the SPSS-installed prerequisite is covered. Still missing are guidance on the undocumented file_path/include_pairwise parameters and any routing signal against the many sibling statistical methods, leaving the definition only adequate for a complex analysis tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It maps three of five parameters (within-factor name, levels, variables) but says nothing about file_path or include_pairwise, leaving those semantics entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Run) and resource (SPSS repeated-measures ANOVA / within-subject GLM), which is enough to separate it from the between-subjects spss_anova and from spss_manova or spss_glm_univariate. It never names a sibling explicitly, so it falls just short of the top tier.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical 'within-subject GLM' implies the design this tool is for, which is implicit when-to-use guidance. It gives no explicit condition for choosing it over spss_anova, spss_mixed, or spss_manova, and no prerequisites beyond the software install.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_run_syntaxSpss Run SyntaxB
Execute arbitrary SPSS syntax commands and return the output as Markdown. Optionally specify a data_file to automatically prepend GET FILE. By default, this also persists .spv (SPSS viewer) and .sps (executed syntax) files. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| syntax | Yes | ||
| dry_run | No | ||
| data_file | No | ||
| select_if | No | ||
| filter_variable | No | ||
| save_syntax_file | No | ||
| save_viewer_output | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose meaningful traits: default persistence of .spv and .sps files (a side effect), the GET FILE prepend behavior, and the hard requirement that IBM SPSS Statistics be installed. It omits what dry_run actually does and how errors/failures are surfaced, which matters for an arbitrary-code execution tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and scope, then optional behavior, then prerequisites. No filler. Slightly more could be trimmed only by cutting the useful GET FILE detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. However, with 7 parameters at 0% schema coverage and a powerful code-execution tool, the unexplained dry_run and filtering parameters leave the definition short of what an agent needs to invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% across 7 parameters, so the description must compensate. It explains syntax, data_file (prepends GET FILE) and the two save_* defaults, covering roughly half the surface, but leaves dry_run, select_if, and filter_variable completely undefined in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Execute) and resource (arbitrary SPSS syntax) and clarifies the return format (Markdown). The word 'arbitrary' implicitly distinguishes it from the many structured sibling tools like spss_frequencies or spss_regression, but the description never explicitly says 'use this for raw syntax instead of the convenience wrappers.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance and no comparison against alternatives. Given 38 siblings including spss_validate_syntax and dozens of higher-level method wrappers, the description should tell the agent when raw execution is the right choice, but it does not.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_structured_resultSpss Structured ResultC
Run SPSS syntax and return a unified structured result as JSON: {markdown, json: {tables, summary}, files, warnings}. The summary extracts key statistics (t/F/r/B/p) for t-test, ANOVA, correlation, regression, descriptives and frequencies.
| Name | Required | Description | Default |
|---|---|---|---|
| syntax | Yes | ||
| data_file | No | ||
| select_if | No | ||
| filter_variable | No | ||
| save_syntax_file | No | ||
| save_viewer_output | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the return structure and summary behavior, which is useful, but does not state execution side effects, permission requirements, data-modification risk, or what the save flags do.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both load-bearing, with the core action and output shape front-loaded. The second sentence efficiently details the summary extraction without filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be fully explained, and the description does cover the return envelope. But with six parameters, no annotations, no parameter descriptions, and no usage guidance relative to many siblings, the description is materially incomplete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Six parameters exist and schema description coverage is 0%, but the description explains none of them. It says 'Run SPSS syntax' which only restates the required 'syntax' parameter name, and adds no meaning for data_file, select_if, filter_variable, save_syntax_file, or save_viewer_output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: run SPSS syntax and return a unified structured result. The output shape and summary-statistic extraction are made explicit, which helps identify it as a structured-output variant. However, it does not explicitly distinguish itself from the sibling spss_run_syntax tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no when-to-use guidance, no exclusions, and no named alternatives. An agent cannot tell from the description when to choose this tool over spss_run_syntax or the many analysis-specific siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_t_testSpss T TestC
Run SPSS t-test. Supports one_sample, independent, and paired test types. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| test_type | Yes | ||
| variables | Yes | ||
| test_value | No | ||
| grouping_variable | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It does disclose a real behavioral constraint — the external IBM SPSS Statistics dependency — which is valuable context. But it says nothing about permissions, side effects, or execution characteristics beyond that, and an output schema exists so return values needn't be covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no filler; the core action leads. Slightly under-specified rather than overly concise, but structurally sound.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter statistical tool with no annotations and 0% schema description coverage, the description is too thin. It never explains what the parameters mean or how test_value/grouping_variable relate to test_type, leaving the agent to guess.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 parameters. The description names the enum values for test_type, but that merely repeats the schema, and it is silent on file_path, variables, test_value, and grouping_variable. With zero schema coverage, the description needed to compensate and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb+resource ('Run SPSS t-test') and it names the three supported test variants, which is concrete. However, it offers no differentiation from siblings like spss_anova or spss_nonparametric_tests, so an agent can't tell when a t-test is the right choice versus those.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description lists the test types but gives no guidance on when to choose one_sample vs independent vs paired, nor any prerequisites beyond SPSS being installed. No mention of alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_twostep_clusterSpss Twostep ClusterB
Run two-step cluster analysis with automatic cluster number determination. Handles large datasets and mixed variable types. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| distance | No | LIKELIHOOD | |
| file_path | Yes | ||
| continuous | No | ||
| categorical | No | ||
| max_clusters | No | ||
| num_clusters | No | ||
| outlier_handling | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose one important operational requirement ('Requires IBM SPSS Statistics to be installed'), but is silent on runtime cost, determinism, how outlier_handling or num_clusters vs max_clusters interact, and what errors to expect. Useful prerequisite context, but significant behavioral gaps remain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with no filler; the core purpose and its differentiators are front-loaded, with the environment prerequisite placed last. Efficient, though the terseness leaves room that the parameter gaps should have used.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but for a 7-parameter analysis tool with zero annotation and zero schema-description coverage the entry is too thin. Key calling decisions (distance metric meaning, num_clusters vs max_clusters precedence, variable-list construction) are left entirely undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 7 parameters, yet the description mentions none of them. It does not explain distance choices, what the continuous/categorical lists expect, or the relationship between num_clusters (fixed k) and max_clusters (automatic search ceiling) — precisely the ambiguity an agent would need resolved.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run two-step cluster analysis') plus the distinctive capability ('automatic cluster number determination. Handles large datasets and mixed variable types'), which implicitly separates it from spss_cluster_hierarchical. It never names the sibling explicitly, so an agent must infer the routing from characteristics rather than a direct pointer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'Handles large datasets and mixed variable types' implies the conditions under which this tool is preferred over hierarchical clustering, but there is no explicit when-to-use/when-not guidance and no mention of alternatives by name. Usage is inferable but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spss_validate_syntaxSpss Validate SyntaxA
Validate SPSS syntax without executing it. Checks for basic syntax errors. Requires IBM SPSS Statistics to be installed.
| Name | Required | Description | Default |
|---|---|---|---|
| syntax | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose useful traits: the operation is non-executing (safe side-effect profile), checking is limited to 'basic' syntax errors, and IBM SPSS Statistics must be installed as a prerequisite. It stops short of saying what is returned on failure, though the output schema covers results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and the key differentiator. No filler, though the dependency sentence could be folded in more tightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter validator with an output schema, the description covers the essential context: what it does, its non-executing nature, and the environment prerequisite. Only the parameter's expected format is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single 'syntax' parameter, and the description only implies it is SPSS syntax text without clarifying format (raw string vs. file path) or size constraints. It adds little meaning beyond the parameter name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Validate SPSS syntax') and immediately differentiates itself from the executing sibling with 'without executing it'. An agent can distinguish it from spss_run_syntax, though the contrast is implied rather than naming the sibling directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Without executing it' hints at the pre-flight use case, but the description never explicitly says when to choose this over spss_run_syntax or that it should precede execution. Usage is implied rather than directed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
51 tool updates
v0.1.0- First observed
spss_anova - First observed
spss_chart_area - First observed
spss_chart_bar - First observed
spss_chart_bar_error - First observed
spss_chart_boxplot - First observed
spss_chart_errorbar - First observed
spss_chart_histogram - First observed
spss_chart_histogram_density - First observed
spss_chart_km_curve - First observed
spss_chart_line - First observed
spss_chart_qqplot - First observed
spss_chart_scatter - First observed
spss_check_status - First observed
spss_cluster_hierarchical - First observed
spss_compute_scale_score - First observed
spss_correlations - First observed
spss_cox_regression - First observed
spss_crosstabs - First observed
spss_descriptives - First observed
spss_discriminant - First observed
spss_factor - First observed
spss_file_summary - First observed
spss_frequencies - First observed
spss_genlin - First observed
spss_genlinmixed - First observed
spss_get_method_schema - First observed
spss_get_method_support - First observed
spss_glm_univariate - First observed
spss_import_csv - First observed
spss_kaplan_meier - First observed
spss_list_files - First observed
spss_list_supported_methods - First observed
spss_list_variables - First observed
spss_logistic_regression - First observed
spss_manova - First observed
spss_mediation - First observed
spss_mixed - First observed
spss_moderation - First observed
spss_nonparametric_tests - First observed
spss_normality_outliers - First observed
spss_ordinal_regression - First observed
spss_read_data - First observed
spss_read_metadata - First observed
spss_regression - First observed
spss_reliability_alpha - First observed
spss_repeated_measures_anova - First observed
spss_run_syntax - First observed
spss_structured_result - First observed
spss_t_test - First observed
spss_twostep_cluster - First observed
spss_validate_syntax
TDQS
Scored across 51 tools
Most tools target clearly distinct SPSS procedures, file operations, or chart types. A few pairs overlap in practice (spss_list_variables vs spss_read_metadata, spss_chart_histogram vs spss_chart_histogram_density, spss_chart_bar vs spss_chart_bar_error), but descriptions generally clarify the boundaries.
All tools use a consistent spss_ snake_case prefix, with chart tools grouped under spss_chart_*. The naming is uniform and predictable across file I/O, statistical procedures, and chart exports.
51 tools is excessive for agent selection and tool management. While SPSS is a broad domain, many tools could be consolidated, especially the 12 chart-export tools and the family of regression/GLM procedures.
The surface covers data discovery, metadata inspection, import, syntax execution, a wide range of statistical analyses, and chart exports. Dedicated data-manipulation tools (recode, compute, merge, select cases) are absent, but spss_run_syntax provides a workaround.
Maintenance
Related MCP Connectors
The statistical analyst in your AI chat — validated, citable, re-runnable analysis of your data.
Use your own Mac from ChatGPT, Claude or Codex: files, commands, documents, and a browser.
SEO & marketing toolkit for AI agents: GA4, Search Console, AdSense, GTM, PageSpeed, Trends.
MCP connector that lets ChatGPT list, search, and run your Apple Shortcuts via a local Mac agent
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceConnects AI agents to a local Stata installation, enabling execution of Stata code, data inspection, graph generation, and result verification through natural language interactions.538 PyPI86AGPL 3.0
- AlicenseAqualityDmaintenanceEnables AI agents to control macOS desktop apps via screenshots, mouse clicks, keyboard input, accessibility queries, and AppleScript.1120 npmMIT
- AlicenseAqualityCmaintenanceEnables AI agents to drive IBM SPSS Statistics for statistical analysis and publication-quality chart generation through natural language, returning structured results and ensuring safe execution.515MIT
- AlicenseNot gradedqualityCmaintenanceProvides a virtual statistician for AI agents, offering real statistical methods such as design of experiments, hypothesis testing, regression, and process control. It includes an advisor tool to recommend appropriate analyses and generates plain-language interpretations of results.MIT