Skip to main content
Glama
Ian3738
by Ian3738
README.md
# r-stats-mcp

[![test](https://github.com/Ian3738/r-stats-mcp/actions/workflows/test.yml/badge.svg)](https://github.com/Ian3738/r-stats-mcp/actions/workflows/test.yml)
[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)
[![Python 3.12+](https://img.shields.io/badge/python-3.12+-blue.svg)](https://www.python.org/downloads/)
[![R 4.0+](https://img.shields.io/badge/R-4.0+-276DC3.svg)](https://www.r-project.org/)

An MCP server that gives an LLM the whole of R's statistical toolkit — hypothesis tests, regression, psychometrics, survival analysis, time series and plots — over a **persistent R session**.

繁體中文說明請見 [README.zh-TW.md](README.zh-TW.md)。

---

## What makes it different

**One R session, kept alive.** Load a dataset once and every later tool call sees it. Fitted models are saved back into that session under a name, so you can fit a model, diagnose it, compare it against another, and then run arbitrary R on it — all without re-reading the data.

**Results written for a reader, not a parser.** Tools return formatted markdown tables rather than raw JSON: p-values as `<.001`, effect sizes with magnitude labels, confidence intervals already assembled. When an assumption is violated the output says what to do about it — a significant Levene's test points you at Welch, sparse expected counts bring in Fisher's exact test automatically.

**No dead ends.** 32 structured tools cover the common ground, and `r_run` executes arbitrary R in the same session for everything else.

```
You: Load survey.csv and check whether the two groups differ on score.

  data_load(path="survey.csv", name="df")
  data_inspect(data="df")
  data_transform(data="df", to_factor=["group"])
  check_assumptions(data="df", variables=["score"], group="group")
  test_ttest(data="df", y="score", group="group", nonparametric=true)
```

---

## Requirements

| | Version | Notes |
| --- | --- | --- |
| **R** | ≥ 4.0 | Must be on `PATH`, or set `R_MCP_RSCRIPT` |
| **Python** | ≥ 3.12 | Managed by uv |
| **uv** | any | [Install](https://docs.astral.sh/uv/getting-started/installation/) |

Required R packages: `jsonlite` and `evaluate`. Everything else is needed only by the tools that use it, and each tool tells you exactly what to install when something is missing.

```r
install.packages(c("jsonlite", "evaluate"))
```

For full coverage of every tool:

```r
install.packages(c(
  "ggplot2", "ragg", "corrplot",          # plotting
  "car", "emmeans", "rstatix",            # ANOVA, post-hoc, VIF
  "psych", "GPArotation", "lavaan",       # scales, factor analysis, SEM
  "lme4", "lmerTest",                     # mixed models
  "survival", "forecast", "tseries",      # survival, time series
  "mgcv", "MASS", "nnet",                 # GAM, ordinal, multinomial
  "haven", "readxl", "openxlsx",          # SPSS/Stata/Excel
  "sandwich", "lmtest"                    # robust standard errors
))
```

---

## Install

### Claude Code

```bash
claude mcp add r-stats -- uvx --from git+https://github.com/Ian3738/r-stats-mcp r-stats-mcp
```

### Claude Desktop

Add to `claude_desktop_config.json` (**macOS**: `~/Library/Application Support/Claude/`, **Windows**: `%APPDATA%\Claude\`):

```json
{
  "mcpServers": {
    "r-stats": {
      "command": "uvx",
      "args": ["--from", "git+https://github.com/Ian3738/r-stats-mcp", "r-stats-mcp"]
    }
  }
}
```

If the client cannot find `uvx`, give the absolute path (`which uvx`).

### From a clone

```bash
git clone https://github.com/Ian3738/r-stats-mcp
cd r-stats-mcp
uv sync
claude mcp add r-stats -- "$(pwd)/run-server.sh"
```

### Configuration

| Variable | Default | Purpose |
| --- | --- | --- |
| `R_MCP_RSCRIPT` | `Rscript` from `PATH` | Path to the R executable |
| `R_MCP_WORKDIR` | Startup directory (home if `/`) | Where relative paths in `data_load` resolve |
| `R_MCP_TIMEOUT` | `180` | Default per-call timeout in seconds |

Verify the setup by asking the model to call `r_session_info`.

---

## Tools

### Session and code

| Tool | Purpose |
| --- | --- |
| `r_run` | Execute arbitrary R code; returns console output and plots |
| `r_install_packages` | Install packages from CRAN |
| `r_session_info` | R version, platform, which statistics packages are available |

### Data

| Tool | Purpose |
| --- | --- |
| `data_load` | CSV, TSV, Excel, SPSS, Stata, SAS, RDS, RData, JSON, Parquet |
| `data_builtin` | Datasets bundled with R or an installed package |
| `data_list` | Datasets and fitted models currently in the session |
| `session_clear` | Drop objects between analyses, keeping named ones |
| `data_inspect` | Types, missing counts, distinct counts, first rows |
| `data_transform` | Filter, derive, recode, factor conversion, sort, long/wide reshape |
| `data_export` | Write out as CSV, TSV, Excel or RDS |

### Descriptives and assumptions

| Tool | Purpose |
| --- | --- |
| `describe` | n, mean, SD, SE, CI, median, quartiles, skew, kurtosis — by group, and survey-weighted |
| `frequency_table` | Frequency tables and cross-tabulations, weighted or unweighted |
| `check_assumptions` | Normality, homogeneity of variance, outliers, VIF, residual diagnostics |

### Complex surveys

Large-scale assessments (TIMSS, PIRLS, PISA, ICILS) need two corrections that
ordinary tools skip: replicate weights for clustered sampling, and plausible
values for latent ability. Omitting either understates the standard error —
on TIMSS 2023 Taiwan, ignoring the weights alone shifts the mean by 4.1 points,
more than its own standard error.

| Tool | Purpose |
| --- | --- |
| `survey_mean` | Means with replicate-weight standard errors and plausible-value pooling |
| `survey_regression` | Weighted regression with the same corrections, plus standardised betas |
| `survey_correlation` | Correlations with design-correct standard errors |

JK2 (TIMSS, PIRLS), JK1, BRR and Fay (PISA) are supported. Give either
`jkzone` + `jkrep` to have JK2 replicates built, or `replicate_weights` to use
columns the study already ships. Achievement scores go in as plausible values —
`pv_prefix="BSMMAT"` (TIMSS style) or `pv_pattern="PV{i}MATH"` (PISA style) —
and the output reports how much of the standard error is sampling versus
measurement.

### Hypothesis tests

| Tool | Purpose |
| --- | --- |
| `test_ttest` | One-sample, independent and paired t-tests, with Levene, Cohen's d and rank-based equivalents |
| `test_anova` | One-way, factorial, ANCOVA, repeated-measures and mixed designs, with effect sizes and post-hoc comparisons |
| `test_categorical` | Chi-square independence and goodness-of-fit, Fisher, McNemar, Cramér's V, odds ratios |
| `test_proportion` | One-sample and multi-group proportion tests, including the exact binomial |
| `correlation` | Pearson, Spearman, Kendall; partial correlations; multiple-comparison adjustment |

### Regression and models

| Tool | Purpose |
| --- | --- |
| `regression` | Linear, logistic, Poisson, negative binomial, ordinal, multinomial, mixed-effects, GAM |
| `model_diagnostics` | Residual normality, heteroscedasticity, autocorrelation, VIF, influential cases |
| `model_compare` | AIC, BIC, log-likelihood, ΔAIC, plus nested-model tests |
| `model_predict` | Estimated marginal means, or predictions on new data |

### Scales and questionnaires

| Tool | Purpose |
| --- | --- |
| `reliability` | Cronbach's α, McDonald's ω, item-total correlations, α-if-dropped |
| `factor_analysis` | EFA and PCA with KMO, Bartlett, parallel analysis, rotated loadings, scree plot |
| `sem` | lavaan CFA / SEM / path / growth models with fit indices, CR and AVE |
| `mediation` | Bootstrap confidence intervals for indirect effects, multiple mediators supported |
| `moderation` | Interaction models with ΔR², simple slopes and an interaction plot |

### Survival and time series

| Tool | Purpose |
| --- | --- |
| `survival_analysis` | Kaplan-Meier with log-rank, Cox regression with the proportional-hazards test |
| `time_series` | Stationarity tests, STL decomposition, automatic ARIMA/ETS, forecasts |

### Plotting

| Tool | Purpose |
| --- | --- |
| `plot` | ggplot2 histogram, density, box, violin, scatter, line, bar, Q-Q, correlation heatmap |

---

## Examples

**Regression with model comparison**

```
regression(data="df", dv="score", predictors=["age","sex","group"], save_as="full")
regression(data="df", dv="score", predictors=["age"], save_as="base")
model_compare(models=["base","full"])
model_diagnostics(model="full")
```

**Complex survey analysis (TIMSS)**

```
data_load(path="bsgtwnm8.sav", name="twn")
survey_mean(data="twn", pv_prefix="BSMMAT", weight="TOTWGT",
            jkzone="JKZONE", jkrep="JKREP", by="ITSEX")
survey_regression(data="twn", pv_prefix="BSMMAT", predictors=["BSBGHER","BSBGSCM"],
                  weight="TOTWGT", jkzone="JKZONE", jkrep="JKREP")
```

**Scale validation**

```
reliability(data="df", items=["q1",...,"q10"], reverse=["q3","q7"], scale_max=5)
factor_analysis(data="df", variables=["q1",...,"q10"])
sem(data="df", type="cfa", model="anxiety =~ q1 + q2 + q3\ndepression =~ q4 + q5 + q6")
```

**Repeated measures**

```
data_transform(data="df", reshape={
  "direction": "long", "value_cols": ["t1","t2","t3"],
  "id_cols": ["id"], "names_to": "time", "values_to": "score"
})
test_anova(data="df", dv="score", within=["time"], id="id", posthoc="bonferroni")
```

**Survival**

```
survival_analysis(data="df", time="days", event="died", group="treatment",
                  type="km", times_of_interest=[180, 365])
survival_analysis(data="df", time="days", event="died",
                  covariates=["age","sex","stage"], type="cox")
```

---

## How it works

```
MCP client  ──stdio/JSON-RPC──▶  Python server  ──NDJSON over pipes──▶  Rscript worker
                                 (tool schemas)                          (.GlobalEnv)
```

**Transport.** A long-lived `Rscript` process reads newline-delimited JSON requests and writes sentinel-delimited JSON responses. Plots are captured with the `evaluate` package — the same machinery `knitr` uses — so base graphics, ggplot2 and lattice all work, and are returned as inline PNGs.

**Namespace hygiene.** User objects live in `.GlobalEnv`; the server's own machinery lives in a separate `.rmcp_sys` environment. `data_list` and `r_run` therefore see only your data and models.

**Timeouts.** Two layers: R enforces its own limit with `setTimeLimit()`, and Python applies a hard timeout on the pipe. If the session has to be restarted, the tool says so explicitly rather than silently losing your data.

**Failure messages are actionable.** Passing a three-level factor to a t-test does not produce a stack trace — it tells you the levels it found and points you at `test_anova`.

---

## Development

```bash
uv sync
uv run python -c "
import asyncio
from r_stats_mcp.server import server
print(len(asyncio.run(server.list_tools())), 'tools')
"
```

Adding a statistical tool means two edits: an R helper in `src/r_stats_mcp/R/` returning `list(md=, plots=)`, and a decorated function in `server.py` describing its parameters. The R files are sourced in filename order at worker startup.

```
src/r_stats_mcp/
├── server.py          MCP tool definitions and schemas
├── session.py         persistent R subprocess, transport, timeouts
└── R/
    ├── worker.R           protocol loop, plot capture
    ├── 00-util.R          markdown tables, formatting, session objects
    ├── 10-data.R          loading, inspection, reshaping
    ├── 20-descriptive.R   descriptives, frequencies, assumptions
    ├── 30-htest.R         t-tests, ANOVA, categorical, correlation
    ├── 40-regression.R    regression family, diagnostics, comparison
    ├── 50-psychometrics.R reliability, factor analysis, SEM, mediation
    ├── 60-survival-ts.R   survival analysis, time series
    ├── 70-viz.R           ggplot2 charts
    └── 80-survey.R        replicate weights, plausible values
```

## License

MIT — see [LICENSE](LICENSE).

TDQS

A3.8/5.0

Scored across 33 tools

Disambiguation4/5

Most tools are clearly separated by domain (data_, test_, model_, survey_), and the survey_* tools are well distinguished by their design-corrected standard errors. However, check_assumptions and model_diagnostics overlap heavily when given a fitted model: both report residual normality, Breusch-Pagan, Durbin-Watson, VIF and the diagnostic plots, which could lead an agent to pick the wrong tool.

Naming Consistency4/5

The set uses snake_case throughout and useful prefixes (data_, test_, model_, survey_) give a predictable pattern. It is not perfect: r_session_info vs session_clear splits the r_/session_ prefix, and a few commands are bare nouns (describe, correlation, sem, plot) rather than verb_noun forms.

Tool Count2/5

33 tools is well above the 25+ threshold and will burden an agent with a very large action space. Each tool is individually defensible for a broad statistics server, but the set could be consolidated—for example, survey_mean, survey_regression and survey_correlation share a lot of design specification boilerplate.

Completeness5/5

The surface is remarkably complete: data ingestion/inspection/transform/export, descriptives, common tests, regression with diagnostics/comparison/prediction, survey methods, psychometrics, SEM, mediation/moderation, survival and time series are all covered. The r_run escape hatch also prevents dead ends for any analysis the structured tools do not anticipate.