cms-datagov-mcp-server
README.md
# CMS Data.gov MCP Server
A Model Context Protocol (MCP) server that provides Claude and other MCP clients with direct access to CMS (Centers for Medicare & Medicaid Services) healthcare data from data.cms.gov.
## Overview
This MCP server enables AI assistants to:
- Search and discover CMS healthcare datasets
- Query dataset records with filters and pagination
- Get dataset statistics and metadata
- Obtain CSV download links for large-scale analysis
Perfect for healthcare analytics workflows, especially when working with LEJR (Lower Extremity Joint Replacement) analyses, provider enrollment data, hospital quality metrics, and other CMS datasets.
## Features
### Five Core Tools
1. **cms_search_datasets** - Find datasets by keyword or theme
2. **cms_get_dataset** - Get detailed dataset information
3. **cms_query_dataset** - Query data with filters (up to 5000 rows)
4. **cms_get_dataset_stats** - Get row counts and column info
5. **cms_get_csv_link** - Get CSV download URL for Athena
### Resource Templates
- `cms://datasets` - Browse all available CMS datasets
- `cms://dataset/{id}` - Access specific dataset metadata
- `cms://csv/{id}` - Get CSV download link
## Installation
### Prerequisites
- Node.js 18 or higher
- Claude Desktop or other MCP-compatible client
### Quick Install
```bash
# Clone or navigate to the project directory
cd cms-datagov-mcp-server
# Install dependencies
npm install
# Build the TypeScript code
npm run build
# Link globally (for Claude Desktop)
npm link
```
### Configure Claude Desktop
Edit your Claude Desktop configuration file:
**macOS:** `~/Library/Application Support/Claude/claude_desktop_config.json`
**Windows:** `%APPDATA%\Claude\claude_desktop_config.json`
**Linux:** `~/.config/Claude/claude_desktop_config.json`
Add this configuration:
```json
{
"mcpServers": {
"cms-datagov": {
"command": "cms-datagov-mcp-server",
"args": [],
"env": {}
}
}
}
```
Restart Claude Desktop to activate the server.
## Usage
### Search for Datasets
```
"Search CMS datasets for TEAM episode"
```
Returns a list of matching datasets with IDs, descriptions, and available formats.
### Get Dataset Details
```
"Get details for dataset 9887a515-7552-4693-bf58-735c77af46d7"
```
Returns comprehensive metadata including API endpoints, CSV links, and column information.
### Query Dataset Records
```
"Query dataset 9887a515-7552-4693-bf58-735c77af46d7 where organization_ccn=100007, return 100 rows"
```
Returns up to 5000 rows with optional filtering, sorting, and column selection.
### Get Dataset Statistics
```
"Get statistics for dataset 9887a515-7552-4693-bf58-735c77af46d7"
```
Returns column names, types, and metadata about the dataset.
### Get CSV Download Link
```
"Get CSV link for dataset 9887a515-7552-4693-bf58-735c77af46d7"
```
Returns direct download URL for creating Athena EXTERNAL TABLEs.
## API Reference
### cms_search_datasets
Search CMS datasets by keyword, theme, or title.
**Parameters:**
- `query` (optional): Search term for titles, descriptions, or keywords
- `theme` (optional): Filter by theme (e.g., 'Medicare', 'Medicaid')
- `limit` (optional): Maximum results to return (default: 10)
**Returns:** Array of matching datasets with metadata
### cms_get_dataset
Get detailed information about a specific dataset.
**Parameters:**
- `dataset_id` (required): UUID of the dataset
**Returns:** Complete dataset metadata including API endpoint and CSV URL
### cms_query_dataset
Query dataset records with filtering and pagination.
**Parameters:**
- `dataset_id` (required): UUID of the dataset
- `filter` (optional): Filter expression (e.g., `[field]=value`)
- `columns` (optional): Comma-separated column list
- `sort` (optional): Column to sort by (prefix with `-` for descending)
- `offset` (optional): Number of rows to skip (default: 0)
- `size` (optional): Number of rows to return (max 5000, default: 100)
**Returns:** JSON array of matching records
### cms_get_dataset_stats
Get statistics about a dataset.
**Parameters:**
- `dataset_id` (required): UUID of the dataset
**Returns:** Column information and dataset metadata
### cms_get_csv_link
Get direct CSV download URL.
**Parameters:**
- `dataset_id` (required): UUID of the dataset
**Returns:** CSV download URL with usage instructions
## Integration with Athena
For large datasets or complex analysis:
1. Use `cms_get_csv_link` to get the download URL
2. Download CSV to your S3 bucket
3. Create an Athena EXTERNAL TABLE:
```sql
CREATE EXTERNAL TABLE cms_team_data (
organization_ccn STRING,
organization_name STRING,
-- ... other columns
)
ROW FORMAT DELIMITED
FIELDS TERMINATED BY ','
STORED AS TEXTFILE
LOCATION 's3://your-bucket/cms-data/'
TBLPROPERTIES ('skip.header.line.count'='1');
```
4. Run complex SQL queries in Athena
## Troubleshooting
### Server Not Appearing in Claude Desktop
1. Verify the server is linked: `npm list -g @clarify/cms-datagov-mcp-server`
2. Check configuration file syntax (valid JSON)
3. Restart Claude Desktop completely
4. Check Claude Desktop logs for errors
### API Errors
- **404 Not Found**: Invalid dataset ID
- **Timeout**: Dataset too large, use CSV download instead
- **Rate Limit**: Wait a moment and retry
### Build Errors
```bash
# Clean and rebuild
rm -rf build node_modules
npm install
npm run build
```
## CMS API Details
- **Base URL:** https://data.cms.gov/data-api/v1
- **Catalog:** https://data.cms.gov/data.json
- **Max Rows:** 5000 per request
- **No Authentication:** Public data, no API key required
## When to Use MCP vs Athena
**Use MCP for:**
- Dataset discovery and exploration
- Quick lookups (< 1000 rows)
- Data validation
- Getting CSV links
**Use Athena for:**
- Large datasets (> 5000 rows)
- Complex joins and aggregations
- GROUP BY operations
- Repeated analysis
## Development
### Build
```bash
npm run build
```
### Watch Mode
```bash
npm run watch
```
### Testing
```bash
node test-validation.mjs
```
## License
MIT License - See LICENSE file for details
## Support
For CMS API questions: OEDAUserResearch@cms.hhs.gov
For MCP protocol documentation: https://modelcontextprotocol.io/
## Related Resources
- [CMS API Documentation](https://data.cms.gov/api-docs)
- [CMS API Guide PDF](https://data.cms.gov/sites/default/files/2022-12/API%20Guide%20Formatted%201_2_0.pdf)
- [Model Context Protocol](https://modelcontextprotocol.io/)
- [Claude Desktop](https://claude.ai/desktop)
TDQS
A3.7/5.0
Scored across 5 tools
Disambiguation4/5
Each tool targets a distinct operation: querying, stats, CSV link, search, and metadata retrieval. cms_get_dataset and cms_get_dataset_stats could be slightly confused, but descriptions clarify their purposes.
Naming Consistency5/5
All tool names follow a consistent snake_case pattern with the 'cms_' prefix and clear verb_noun structure, making them predictable and easy to scan.
Tool Count5/5
Five tools are well-scoped for a data catalog server, covering discovery, metadata, querying, stats, and export without redundancy.
Completeness4/5
The tool set covers key aspects of dataset interaction, but lacks tools for schema exploration (e.g., listing columns) or direct data export beyond CSV links, which might be needed for some workflows.
Maintenance
ActivityInactive
ResponsivenessNo issues