educator-toolkit-mcp
by TyRobbins
README.md
# educator-toolkit-mcp
A portable MCP toolkit exposing six educator-craft AI tools for reviewing, critiquing, and improving learning artifacts. Wraps structured AI-powered review tools and exposes them three ways:
1. **Hosted on GCP** — call directly over HTTPS (REST) or connect Claude Desktop / Cursor via MCP Streamable HTTP. No install needed.
2. **Local MCP server** — clone and run as a subprocess of Claude Desktop over stdio.
3. **Importable Python kernel** — embed `run_tool` directly in your own application.
---
## Option 1: Use the hosted service (no install)
**Base URL:** `https://educator-toolkit-server-h4ldfaoucq-uw.a.run.app`
> **Prototype, no auth.** The URL is unauthenticated for prototyping. Don't paste it into public places — anyone with it can call Anthropic on the shared API budget. Real auth arrives with the e-commerce milestone.
### REST — `POST /v1/{tool_name}`
```bash
curl -X POST https://educator-toolkit-server-h4ldfaoucq-uw.a.run.app/v1/ai_tool_critique \
-H "Content-Type: application/json" \
-d '{
"question": "Is the input schema asking too much upfront?",
"artifact": "A form with 8 required fields for a first-time user, no examples.",
"artifact_kind": "ai_tool_prompt"
}'
```
Response shape:
```json
{
"status": "complete" | "needs_input" | "out_of_lane" | "refused" | "error",
"assessment": "...",
"key_points": ["..."],
"follow_ups": ["..."] | null,
"suggested_tool": "..." | null,
"caveats": ["..."] | null,
"error": null
}
```
Replace `ai_tool_critique` with any of the six tool IDs listed in **The six tools** section below.
### Claude Desktop (remote MCP)
In `claude_desktop_config.json`:
```json
{
"mcpServers": {
"educator-toolkit-remote": {
"url": "https://educator-toolkit-server-h4ldfaoucq-uw.a.run.app/mcp/",
"transport": "streamable-http"
}
}
}
```
Note: trailing slash on `/mcp/` is required (Starlette mount convention).
Restart Claude Desktop. All six tools will be available in the tool picker.
### Cursor / Cline / other MCP clients
Same shape — point them at `https://educator-toolkit-server-h4ldfaoucq-uw.a.run.app/mcp/` over the Streamable HTTP transport. No headers, no auth.
### Health check
```bash
curl https://educator-toolkit-server-h4ldfaoucq-uw.a.run.app/health
# {"status":"ok"}
```
### OpenAPI / docs
Auto-generated by FastAPI:
- Swagger UI: https://educator-toolkit-server-h4ldfaoucq-uw.a.run.app/docs
- OpenAPI JSON: https://educator-toolkit-server-h4ldfaoucq-uw.a.run.app/openapi.json
---
## Option 2: Run locally as a Claude Desktop subprocess
For when you want to use your own Anthropic API key or run without Cloud Run.
```bash
git clone https://github.com/TyRobbins/educator-toolkit-mcp.git
cd educator-toolkit-mcp
uv sync --all-packages --all-groups
export ANTHROPIC_API_KEY=sk-ant-...
```
In `claude_desktop_config.json` (Windows path; macOS path differs):
```json
{
"mcpServers": {
"educator-toolkit-local": {
"command": "uv",
"args": [
"run",
"--directory", "C:\\Users\\YOU\\path\\to\\educator-toolkit-mcp",
"educator-toolkit-mcp"
],
"env": {
"ANTHROPIC_API_KEY": "sk-ant-..."
}
}
}
}
```
The local server uses MCP over **stdio** (Claude Desktop launches it as a child process). The hosted version uses MCP over **HTTP**.
---
## Option 3: Import the kernel directly
```python
import asyncio
from educator_toolkit_kernel import AnthropicClient, run_tool
from educator_toolkit_kernel.schema import ToolInput
client = AnthropicClient() # reads ANTHROPIC_API_KEY from env
response = asyncio.run(run_tool(
tool_id="ai_tool_critique",
input=ToolInput(
question="Is the input schema asking too much upfront?",
artifact="A form with 8 required fields for a first-time user, no examples.",
artifact_kind="ai_tool_prompt",
),
client=client,
))
print(response.status)
print(response.assessment)
print(response.key_points)
```
---
## The six tools
| Tool ID | What it does | Use when |
|---|---|---|
| `educational_design_review` | Review a learning artifact for instructional-design quality — objectives, scaffolding, Bloom's taxonomy, cognitive load | You want a learning-science review of a lesson, module, or activity |
| `curriculum_coherence_review` | Review a curriculum or module sequence for coherence, redundancy, and gaps across phases | You have 2+ modules and want sequencing feedback |
| `ai_tool_critique` | Critique an AI tool's prompt, input schema, and output usefulness for its intended learner | You have one specific AI tool (prompt + schema) and want a critique |
| `student_experience_simulation` | Simulate what a realistic learner would do, feel, and produce when using an artifact | You want a persona-grounded prediction of learner behavior |
| `assessment_design_review` | Review or design assessments: cases, rubrics, formative checks, capstone scenarios | You have an assessment artifact (rubric, case, exam) or want one designed |
| `sme_review` | Review content for accuracy, terminology consistency, and fidelity to a supplied source | You supply BOTH a content artifact AND source material for comparison |
---
## Worked example
An instructional designer has built an AI tool that helps learners write stakeholder communication plans. The tool has a long intake form. They want to know if the schema is too heavy.
**ToolInput sent:**
```json
{
"question": "Is the input schema asking too much upfront?",
"artifact": "A form with 8 required fields for a first-time user, no examples.",
"artifact_kind": "ai_tool_prompt"
}
```
**ToolResponse returned:**
```json
{
"status": "complete",
"assessment": "Eight required fields with no examples is excessive for first-time users and will cause high abandonment. Core fixes: reduce required fields to 3-4, add placeholder examples to every field, use progressive disclosure for optional context, and sequence by cognitive load.",
"key_points": [
"8 required fields exceeds the friction threshold for first-time users — expect abandonment",
"Zero examples or placeholder text creates blank-canvas paralysis across all fields",
"All-required structure signals rigidity and prevents progressive onboarding",
"Fix: cut to 3-4 required fields, make the rest optional with a clear rationale",
"Fix: add one inline example per field showing what a good answer looks like",
"Fix: consider a two-step form — generate output first, refine with more input second"
],
"follow_ups": [
"Can you share the actual field labels? Some fields may be combinable or eliminable once we see them.",
"Does the underlying AI prompt use all 8 fields, or are some fields only used for logging/routing?"
],
"caveats": [
"Artifact was described, not shown — specific field label critique requires the actual schema",
"Whether the 8 fields are pedagogically correct is out of lane for this review (see educational_design_review)"
],
"error": null
}
```
---
## How it works
Each tool makes **one Sonnet call** — no coordinator, no sub-tool loop. The caller (Claude Desktop, a REST client, or your own app) decides which tool to route to and composes results from multiple tools if needed.
- Stateless: each call is independent. The caller maintains context across follow-up rounds.
- Schema gate: every response is validated against `ToolResponse` before returning. If the model omits the JSON envelope, the prose is wrapped verbatim.
- One model per tool: all six tools currently use `claude-sonnet-4-6`.
---
## Quality
Validated against a 12-fixture golden eval set (2 per tool, LLM-judge rubric, 4 dimensions x 1-5 scale). Current average: **4.90/5.00**. See `evals/`.
CI gate: score must stay >= 3.5.
Schema validity gate: `uv run python -m evals.runners.schema_validity` — all 12 fixtures must return parseable `ToolResponse` JSON.
---
## Architecture
```
┌─────────────────────────────────────────┐
│ Hosts: Claude Desktop, Cursor, scripts │
└────────────┬──────────────┬─────────────┘
│ │
stdio MCP HTTP / MCP
│ │
┌────────────▼──┐ ┌───────▼───────────────┐
│ packages/mcp │ │ packages/server │
│ (local sub- │ │ (Cloud Run) │
│ process) │ │ /v1/{tool_name} │
│ │ │ /mcp/ │
└────────────┬──┘ └───────┬───────────────┘
│ │
└──────┬───────┘
│
┌───────▼────────────────────┐
│ packages/kernel │
│ - schema (ToolInput, │
│ ToolResponse) │
│ - registry (TOOLS) │
│ - runner (run_tool) │
│ - prompts (6 tool prompts) │
└────────────────────────────┘
```
---
## Development
```bash
git clone https://github.com/TyRobbins/educator-toolkit-mcp.git
cd educator-toolkit-mcp
uv sync --all-packages --all-groups
uv run pytest packages/
# Lint + type-check
uv run ruff check .
uv run mypy packages/kernel/src packages/mcp/src packages/server/src
# Evals (require ANTHROPIC_API_KEY)
uv run python -m evals.runners.schema_validity
uv run python -m evals.runners.judge_rubric
# Run the HTTP server locally
ANTHROPIC_API_KEY=sk-ant-... uv run educator-toolkit-server
# -> http://localhost:8080/health
```
---
## Deploy
```bash
gcloud builds submit --config cloudbuild.yaml --project ou-executive-persuasion
```
Builds the Dockerfile, pushes to GCR, deploys to Cloud Run (region `us-west1`, public, no auth). The `ANTHROPIC_API_KEY` is read from Secret Manager (entry `anthropic-api-key`).
---
## License
TBD pending IP review.
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessNo issues