Skip to main content
Glama

education-data-mcp

An MCP server over the Urban Institute's Education Data Portal — harmonized federal education data (CCD, CRDC, IPEDS, EdFacts, SAIPE, College Scorecard, MEPS, PSEO) for schools, school districts, and colleges.

Queries the live EDP API. No API key required. Read-only.

Prerequisites

  • Python 3.11+

  • uv

Related MCP server: data-utah

Installation

git clone https://github.com/UrbanInstitute/education-data-mcp.git
cd education-data-mcp
uv sync

Tools

Six tools, following a discovery-first workflow:

Tool

Purpose

search_datasets

Find available datasets by level, source, topic, or keyword

describe_dataset

Inspect a dataset's variables, filters, and coded value formats

get_data

Fetch raw data records with human-readable labels (by default)

get_summary

Get aggregated statistics (counts, sums, averages) by group

lookup_codes

Translate between codes and labels (e.g., FIPS 6 = California)

resolve_entity

Find a school/district/college ID by name (e.g., "Harvard" → unitid 166027)

Typical workflow: search_datasets → describe_dataset → get_summary or get_data, with lookup_codes and resolve_entity as needed.

Tool reference

search_datasets

Find available datasets.

Parameter

Type

Description

level

string (optional)

"schools", "school-districts", or "college-university"

source

string (optional)

"ccd", "ipeds", "crdc", "edfacts", "saipe", "scorecard", etc.

topic

string (optional)

"enrollment", "directory", "finance", "discipline", etc.

search

string (optional)

Keyword to match against dataset URLs and descriptions

describe_dataset

Inspect a dataset's variables, filters, and coded value formats.

Parameter

Type

Description

path

string

Dataset path, raw template or filled in: "schools/ccd/enrollment/{year}/{grade}/race/" or "schools/ccd/enrollment/2022/grade-99/race/"

filterable_only

boolean (default: false)

Only show variables usable as query filters

get_data

Fetch education data. Returns human-readable labels by default.

Parameter

Type

Description

path

string

A path with every {placeholder} filled in: "schools/ccd/directory/2022/". One year per call

filters

string (optional)

Query filters: "fips=11&charter=1". Also accepts "ordering=-enrollment" to rank server-side

fields

string (optional)

Columns to return: "ncessch,school_name,enrollment"

add_labels

boolean (default: true)

Decode coded values to human-readable labels

preview

boolean (default: false)

Return a small labelled sample instead of a complete result

get_summary

Get aggregated statistics from the Education Data Portal. Use for counts, totals, or averages across groups. Much faster than fetching raw data and computing yourself.

Parameter

Type

Description

path

string

The dataset path WITHOUT {placeholders} — the static head of the template: "schools/ccd/enrollment"

var

string

Variable to aggregate: "enrollment", "teachers_fte", etc.

stat

string

Statistic: "sum", "count", "avg", "min", "max", "median", "stddev", "variance"

by

string

Grouping variables (comma-separated): "fips", "race", "fips,charter"

filters

string (optional)

Query filters: "fips=6" or "fips=6&year=2022". ordering= is not supported here

Results are always grouped by year in addition to the specified groupings.

Example: Total enrollment by state: get_summary(path="schools/ccd/enrollment", var="enrollment", stat="sum", by="fips")

lookup_codes

Look up code-to-label mappings. Use to find filter values (e.g., "California" = FIPS 6) or understand coded results.

Parameter

Type

Description

format_name

string

Format name from describe_dataset output: "fips", "race", "sex", etc.

codes

string (optional)

Comma-separated codes: "1,2,3". Omit to see all values.

Coded values are field-specific. Code meanings come from each variable's API metadata, not a fixed table, and get_data/get_summary decode them to labels automatically. Negative codes often mean -1 = Missing/not reported, -2 = Not applicable, -3 = Suppressed for privacy — but not always: for grade, -1 means Pre-K (0 = Kindergarten). Use lookup_codes to see a variable's full, authoritative code list.

resolve_entity

Resolve a school, district, or college name to its ID for filtering. Use when you need to filter data by a specific entity.

Parameter

Type

Description

name

string

Name to search for (case-insensitive substring match)

entity_type

string

"school", "district", or "college"

fips

integer (optional)

State FIPS code (e.g., 6 for California). Required for schools and districts

max_results

integer (default: 15)

Maximum number of results to return

Returns matching entities with their IDs (ncessch for schools, leaid for districts, unitid for colleges) for use in get_data filters.


Running it

MCP Inspector (interactive testing)

uv run mcp dev src/edp_mcp/server.py

Opens a browser at http://localhost:6274 — connect, open Tools, and run any tool with parameters.

Claude Desktop / Claude Code / VS Code / Copilot CLI

{
  "mcpServers": {
    "edp-mcp": {
      "command": "uv",
      "args": ["run", "--directory", "/absolute/path/to/education-data-mcp", "edp-mcp"]
    }
  }
}

Client

Where it goes

Claude Desktop

claude_desktop_config.json — macOS: ~/Library/Application Support/Claude/; Windows: %APPDATA%\Claude\

Claude Code

.claude/settings.json, or claude mcp add edp-mcp -- uv run --directory /absolute/path/to/education-data-mcp NAME

VS Code (Copilot)

.vscode/settings.json, nested as {"mcp": {"servers": {...}}}

stdio (direct)

uv run edp-mcp

Streamable HTTP (hosted)

MCP_TRANSPORT=streamable-http PORT=8080 uv run edp-mcp

MCP is served at POST /mcp; GET /health is a plain unauthenticated health check for a load balancer or orchestrator.

Variable

Default

Purpose

MCP_TRANSPORT

stdio

stdio or streamable-http

PORT

8080

Listen port

MCP_HOST

0.0.0.0

Bind address

MCP_ALLOWED_HOSTS

(unset)

Comma-separated Host allow-list

MCP_ALLOWED_ORIGINS

(unset)

Comma-separated Origin allow-list

Set MCP_ALLOWED_HOSTS to the public hostname before exposing this beyond a private network. Both are unset by default, which leaves the SDK's DNS-rebinding protection off — setting either turns it on. Note that enabling it with an allow-list that omits the real hostname rejects every request.

Hosted

Also available as a hosted streamable-HTTP server: https://educationdata.urban.org/mcp/edp

Tests

uv run pytest -q

Tests use saved API fixtures (tests/fixtures/) with mocked HTTP, so no network is needed. scripts/smoke_live.py exercises the real API end to end.

Available Tools

6 tools
describe_datasetDescribe datasetA
Read-onlyIdempotent

Inspect one dataset: every variable with its definition, data type, coded-value format, and whether it is filterable. Also returns the years covered, the source description, and equivalent R/Stata code.

This is where the [FILTER] marks come from — filtering on anything else returns UNFILTERED data rather than an error.

Args: path: Dataset path, template or filled in — both resolve to the same dataset: "schools/ccd/enrollment/{year}/{grade}/race/" or "schools/ccd/enrollment/2022/grade-99/race/" filterable_only: Show only the variables usable as filters

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
filterable_onlyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, covering safety. The description adds valuable behavioral details beyond that: the FILTER mark behavior, template path resolution (both template and filled paths resolve to same dataset), and the inclusion of equivalent R/Stata code in the output. No contradictions. This exceeds the baseline and earns a 4 for adding meaningful behavioral context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded: a concise first sentence summarizing the output, followed by a crucial behavioral note about FILTER marks, then a clear 'Args' section. Every sentence earns its place — there is no fluff or repetition. The length is appropriate for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema (has output schema: true), the description doesn't need to detail the return structure, but it still lists the content types. It explains the path parameter flexibility, filter behavior, and the presence of R/Stata code. Annotations cover safety. Nothing an agent needs to call this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must fully explain parameters. It does: 'path' is described as a dataset path, template or filled, with concrete examples, and 'filterable_only' is described as showing only variables usable as filters. This adds meaningful meaning beyond the raw schema, fully compensating for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Inspect') and resource ('one dataset'), then enumerates exactly what it returns: variable definitions, data types, coded-value format, filterability, years covered, source description, and equivalent R/Stata code. It also explicitly ties to the [FILTER] marks, distinguishing it from siblings like get_data or search_datasets by describing its unique role in revealing filterable fields.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use this tool: it is 'where the [FILTER] marks come from' — implying that when the agent needs to know which fields are filterable before querying data, this is the tool. It also warns about behavior ('filtering on anything else returns UNFILTERED data rather than an error'), effectively telling the agent to consult this tool to avoid incorrect filtering. It doesn't explicitly name sibling tools, but the guidance is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dataGet data recordsA
Read-onlyIdempotent

Fetch raw data records from one dataset path. For totals or averages across groups use get_summary instead — it aggregates server-side.

Results are COMPLETE or refused, never truncated, so counting and averaging over the rows returned is valid. A query matching over 10,000 rows is refused with its true size and how to narrow it. When one is too big, aggregate with get_summary, narrow to a state or district, or hand the user the bulk CSV link from the refusal or the R/Stata snippet from describe_dataset — for whole-country or long multi-year analysis those beat paging through calls.

COST: this hits a live API and a broad query can take 30s+. Filter before fetching rather than issuing many wide calls in parallel.

Args: path: A dataset path from search_datasets with every {placeholder} filled in — "schools/ccd/enrollment/2022/grade-99/race/". One year per call; a leading "/api/v1/" is optional. filters: "fips=11&charter=1". Only fields marked [FILTER] in describe_dataset work; others are rejected here rather than silently returning unfiltered data. Also takes "ordering=" to rank server-side: "ordering=-enrollment" largest first, "ordering=enrollment" smallest. Ranking sorts the whole result before paging, so a ranked query answers where an unranked one is refused as too large. fields: Comma-separated columns — "ncessch,school_name,enrollment". ALWAYS PASS THIS. These tables are 50-96 columns wide; requesting only what you need is typically an 8-17x reduction and is usually the difference between an answer and a refusal. add_labels: Decode coded values to labels (default True). Decoding is per variable, so meanings are FIELD-SPECIFIC — trust the decoded label over any assumption about what a raw code means. preview: Return a small labelled SAMPLE rather than a complete result. The rows are the API's first N by ID, NOT a random sample — never count, rank or average over them.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
fieldsNo
filtersNo
previewNo
add_labelsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint, destructiveHint. The description adds critical behavioral details: results are complete or refused, never truncated; queries over 10,000 rows are refused with true size and narrowing suggestions; a cost warning about 30s+ latency; preview rows are not random samples; decoding is field-specific. These go well beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence carries essential information. It front-loads the core purpose and alternative guidance, then behavioral notes, then cost, then parameter details in a structured Args list. While dense, the structure makes it scannable; the length is justified by the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 params, no schema descriptions, output schema present), the description covers all necessary context: when to use, cost, refusal behavior, sampling caveats, and parameter semantics. An agent has everything needed to call it correctly and avoid misuse.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description is the sole source for parameter meaning. It explains path (with placeholder syntax and optional leading slash), filters (with example, only [FILTER] fields work, ordering semantics), fields (with example and strong advice to always pass), add_labels (default and caveat), and preview (non-random sample caveat). Each parameter is thoroughly documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Fetch raw data records from one dataset path' — a specific verb, resource, and scope. It explicitly distinguishes from get_summary ('For totals or averages across groups use get_summary instead'), so an agent can tell it apart without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly directs when to use get_summary for aggregation, and suggests narrowing or using the bulk CSV link from refusal or the R/Stata snippet from describe_dataset for large analyses. It also advises filtering before fetching rather than issuing many wide calls in parallel, providing clear alternatives and context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_summarySummarize dataA
Read-onlyIdempotent

Aggregate a dataset server-side: counts, sums, averages by group. Use this rather than get_data whenever the question is about totals ("how many schools per state", "total enrollment by race").

EVERY year comes back in one call — results are grouped by year on top of by, so a trend needs one call, never one per year. Filter years only to narrow a large result. Results are COMPLETE or refused. The API cannot rank aggregates, and refuses groupings that produce too many groups (by=leaid nationally, for instance).

var and by are drawn from two different pools — describe_dataset lists both for any dataset. Guessing costs a slow round trip; reading them does not.

COST: this hits a live API and a broad aggregation can take 30s+.

NO COUNTY AGGREGATION for schools or districts. Several datasets return county_code as a column, but it is neither filterable nor groupable, so county totals cannot be computed here — fetch the rows with get_data and aggregate them yourself, or say the portal does not support it. Do not retry with county in by.

Args: path: The SUMMARY path — no year, no {placeholders}, e.g. "schools/ccd/enrollment". describe_dataset prints the right one for any dataset. Some differ from the data path ("college-university/ipeds/fall-enrollment/race"). var: The measure to aggregate — a numeric, non-filter variable stat: sum, count, avg, min, max, median, stddev, or variance by: Comma-separated groupings — "fips", "fips,race". May also use variables from the source's directory ("school_level", "sector"). filters: "fips=6" — omit year to get every year; "year=2018,2019,2020" to restrict the range

ParametersJSON Schema
NameRequiredDescriptionDefault
byYes
varYes
pathYes
statYes
filtersNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds substantial behavioral context beyond annotations: 'Results are COMPLETE or refused' (failure mode), 'this hits a live API and a broad aggregation can take 30s+' (performance), 'Guessing costs a slow round trip; reading them does not' (latency cost), and 'The API cannot rank aggregates' (capability limit). It also reveals that 'EVERY year comes back in one call' and that the API refuses groupings producing too many groups. This is rich disclosure that materially helps the agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place. It is front-loaded with purpose and usage, then moves into behavioral constraints, performance cost, and parameter semantics. The Args section is clearly structured. Given the tool's complexity and the need to compensate for the lack of schema descriptions, the length is appropriate — there is no fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters, 0% schema descriptions, and no enums, the description covers everything an agent needs: purpose, usage, exclusions, behavioral constraints, failure modes, performance, and parameter semantics. It also names alternative tools and explains when to use them. Since an output schema exists, it does not need to describe return values. No critical information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% — no parameter descriptions are provided in the input schema. The description fully compensates by explaining every parameter in the Args section with concrete examples and constraints: path ('The SUMMARY path — no year, no {placeholders}, e.g. schools/ccd/enrollment'), var ('the measure to aggregate — a numeric, non-filter variable'), stat (lists the allowed values), by (examples 'fips' and 'fips,race' plus directory variables), and filters (examples with year and fips). This is exactly what the agent needs to construct valid calls.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise statement of what the tool does ('Aggregate a dataset server-side: counts, sums, averages by group') and immediately differentiates it from the sibling get_data by stating 'Use this rather than get_data whenever the question is about totals.' This gives the agent a clear verb, resource, and scope, and explicitly names the alternative it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance ('Use this rather than get_data whenever the question is about totals'), plus concrete exclusions and alternatives: 'NO COUNTY AGGREGATION... fetch the rows with get_data and aggregate them yourself,' and 'Do not retry with county in by.' It also explains when to filter years and warns that 'The API cannot rank aggregates' and 'refuses groupings that produce too many groups.' These are clear usage rules and alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lookup_codesLook up coded valuesA
Read-onlyIdempotent

Look up the code-to-label mapping for a coded variable — mainly to find a filter value for get_data ("California" -> fips=6).

Args: format_name: The "Format" shown for a variable in describe_dataset — "fips", "race", "sex", "school_level", "charter", … codes: Comma-separated codes to look up — "6,48". Omit for all values.

ParametersJSON Schema
NameRequiredDescriptionDefault
codesNo
format_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, destructiveHint=false. The description adds behavioral context beyond that: its role in supporting get_data and that omitting 'codes' returns all values. This is useful supplemental transparency without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded: the main purpose and example appear in the first sentence, followed by clear parameter definitions. No wasted words; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return format is covered there. The description provides everything an agent needs to call the tool correctly: purpose, parameter semantics, and the key behavior of omitting codes. It is complete for a lookup helper with these annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden for parameters. It completely explains format_name ('The "Format" shown for a variable in describe_dataset' with examples) and codes ('Comma-separated codes to look up — "6,48". Omit for all values.'). This fully compensates for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Look up the code-to-label mapping for a coded variable' and immediately ties it to a concrete use case with an example ('"California" -> fips=6'). This clearly differentiates it from sibling tools like describe_dataset or get_data. The purpose is unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says 'mainly to find a filter value for get_data', giving clear primary context and the example reinforces when it applies. It lacks an explicit 'when not to use' or alternatives, but the main scenario is well conveyed. A 4 is appropriate for strong context without explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_entityResolve entity name to IDA
Read-onlyIdempotent

Resolve a school/district/college NAME to the ID used in get_data filters.

Returns CANDIDATES with their city and level, not one authoritative answer — names repeat, so the right row is the one whose city and level fit what the user meant. Type words ("Elementary", "High", "ISD") are ignored when matching, which is what lets natural phrasings reach the directory's abbreviations ("Lakewood Elementary" finds "LAKEWOOD EL").

Schools and districts require fips: without it this asks which state rather than guessing, since "Homestead High" exists in CA, WI, IN and FL.

COST: downloads a whole state's directory, so the first call for a state often takes 10-30s and later ones are cached. Resolve names in the same state together.

Args: name: Name to search for; distinctive words matter most. entity_type: "school", "district", or "college" fips: State FIPS code (6 = California). Required for schools and districts. See lookup_codes(format_name="fips"). max_results: Maximum candidates to return (default 15)

ParametersJSON Schema
NameRequiredDescriptionDefault
fipsNo
nameYes
entity_typeYes
max_resultsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readOnly/idempotent/non-destructive; the description adds substantial non-obvious behavior: candidate-style returns, type-word stripping, state-asking rather than guessing, and 10-30s first-call cost with caching. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence carries operational value, with the core task front-loaded and behavior/cost notes structured before Args. No filler or repetition beyond the necessary Args section.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter resolution tool with an output schema, it covers matching semantics, fips dependency, latency/caching, batching advice, and candidate interpretation. Nothing essential is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the Args section explains every parameter: name's distinctive-word advice, entity_type's allowed values, fips's conditional requirement and lookup pointer, and max_results default. This fully compensates for the empty schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource: 'Resolve a school/district/college NAME to the ID used in get_data filters,' which positions it clearly against get_data and lookup_codes. It also distinguishes itself by stating it returns candidates, not an authoritative answer, because names repeat.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tells the agent this feeds get_data filters and that fips is required for schools/districts, with a pointer to lookup_codes. It provides clear context but stops short of explicit when-not-to-use/alternative selection among sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_datasetsSearch datasetsA
Read-onlyIdempotent

Find datasets in the Education Data Portal. Returns each one's path template, description and years available; pass a template to describe_dataset for its variables and its summary call. All arguments are optional filters.

NOT COVERED by this portal at all: NAEP scores (see the NAEP Data Explorer), teacher salaries (BLS), private K-12 (limited; CRDC covers some), and curriculum data. Say so rather than searching repeatedly.

Args: level: "schools", "school-districts", or "college-university" source: "ccd", "ipeds", "crdc", "edfacts", "saipe", "scorecard", … topic: "enrollment", "directory", "finance", "discipline", … search: Words matched per-token against dataset paths and descriptions, not as a phrase — so drop specific terms like grade numbers ("8th") if a query returns nothing, since datasets rarely spell those out verbatim.

ParametersJSON Schema
NameRequiredDescriptionDefault
levelNo
topicNo
searchNo
sourceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/openWorld/idempotent annotations, the description explains that all arguments are optional filters and that search matches per-token rather than as a phrase. It also warns about likely-empty results for grade-level terms, adding useful behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose and return shape, followed by a compact Args list. The NOT COVERED paragraph is longer but earns its place by preventing wasted searches. No filler or redundant restatement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no lengthy explanation. All four parameters are documented, the key sibling routing is provided, and known coverage gaps are spelled out. An agent has everything needed to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden, and it delivers: each parameter gets concrete allowed values or examples, and 'search' gets a detailed matching semantics explanation. This compensates fully for the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('Find datasets') on a clear resource ('Education Data Portal'), and describes the result: each dataset's path template, description, and years available. It also routes the agent to describe_dataset when variables are needed, which helps distinguish it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to pass a template to describe_dataset for variables and summary calls, and lists domains not covered by the portal so the agent should not search repeatedly. This gives clear when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.0
    • First observeddescribe_dataset
    • First observedget_data
    • First observedget_summary
    • First observedlookup_codes
    • First observedresolve_entity
    • First observedsearch_datasets

TDQS

A4.8/5.0

Scored across 6 tools

Disambiguation5/5

Each tool has a clearly distinct role: discover datasets, inspect schema, fetch raw rows, aggregate, decode codes, and resolve names to IDs. The potentially confusing get_data/get_summary pair is explicitly separated and cross-referenced.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun convention: describe_dataset, search_datasets, get_data, get_summary, lookup_codes, resolve_entity. No mixed casing or vague generic verbs appear.

Tool Count5/5

Six tools form a compact, non-redundant set that covers the read-only data-portal workflow well. Each tool earns its place, and the count is well within the ideal range for an agent to navigate comfortably.

Completeness5/5

The tool surface covers the full query lifecycle: find datasets, inspect variables and filter fields, resolve codes and entity names, fetch raw data, and aggregate server-side. Known unsupported operations are explicitly documented with workarounds, leaving no obvious dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers