csv-explorer-mcp
# CSV Explorer MCP Server
A Model Context Protocol (MCP) server for exploring and analyzing CSV files. Provides tools for inspection, sampling, schema inference, statistics, filtering, and more.
## Installation
```bash
npm install
npm run build
```
## Usage
Add to your MCP configuration:
```json
{
"mcpServers": {
"csv-explorer": {
"command": "node",
"args": ["path/to/dist/index.js"]
}
}
}
```
## Tools
### csv_inspect
Get an overview of a CSV file including size, row/column count, detected delimiter, and a preview of the data. Large field values are automatically truncated with content-type hints.
```typescript
csv_inspect({ file: "/path/to/data.csv", previewRows: 5 })
```
### csv_sample
Get sample records using various sampling strategies.
```typescript
csv_sample({ file: "/path/to/data.csv", mode: "random", count: 10 })
// modes: "first", "last", "random", "range"
```
### csv_schema
Infer the schema by sampling records. Returns column names, types, and nullability.
```typescript
csv_schema({ file: "/path/to/data.csv", sampleSize: 1000 })
// outputFormat: "inferred", "json-schema", "formatted"
```
### csv_stats
Collect aggregate statistics for fields. Includes min/max, mean, median, stdDev for numeric fields, and top values for categorical fields.
```typescript
csv_stats({ file: "/path/to/data.csv", fields: ["price", "category"] })
```
### csv_search
Search for records where a field matches a regex pattern.
```typescript
csv_search({ file: "/path/to/data.csv", field: "email", pattern: "@example\\.com$" })
```
### csv_filter
Filter records using query expressions. Supports comparisons (`==`, `!=`, `<`, `>`, `<=`, `>=`), text operations (`contains`, `startswith`, `endswith`, `matches`), and compound queries (`AND`, `OR`).
```typescript
csv_filter({ file: "/path/to/data.csv", query: 'status == "active" AND age > 30' })
```
### csv_validate
Validate a CSV file for syntax errors and optionally against a schema.
```typescript
csv_validate({
file: "/path/to/data.csv",
schema: {
columns: [
{ name: "id", type: "integer", required: true },
{ name: "email", type: "string", pattern: "^[^@]+@[^@]+$" }
]
}
})
```
### csv_tail
Read new records appended since a cursor position. Use for monitoring actively-written files.
```typescript
csv_tail({ file: "/path/to/data.csv", cursor: 1024, maxRecords: 100 })
```
### csv_get_cursor
Get the current end-of-file position for use with `csv_tail`.
```typescript
csv_get_cursor({ file: "/path/to/data.csv" })
```
### csv_diff
Compare two CSV files and report differences.
```typescript
csv_diff({ file1: "/path/to/old.csv", file2: "/path/to/new.csv", keyField: "id" })
```
### csv_extract
Extract a specific field value from a CSV record. Use for retrieving large/truncated field data. Can write to file for binary data (e.g., base64 images).
```typescript
// Get field value inline
csv_extract({ file: "/path/to/data.csv", field: "description", line: 5 })
// Decode base64 and write to file
csv_extract({
file: "/path/to/data.csv",
field: "screenshot",
line: 1,
decode: "base64",
outputFile: "/tmp/screenshot.png"
})
```
### csv_large_fields
List fields containing large values (e.g., base64 images, JSON blobs). Helps identify which fields were truncated in `csv_inspect`.
```typescript
csv_large_fields({ file: "/path/to/data.csv", threshold: 1000, sampleRows: 100 })
```
## Features
- **Streaming Architecture**: Memory-efficient processing of large files
- **Auto-Detection**: Automatically detects delimiters (comma, tab, semicolon, pipe) and encoding
- **Smart Truncation**: Large field values are truncated with content-type hints (base64, JSON, HTML)
- **Query Engine**: Filter records with SQL-like expressions supporting AND/OR logic
- **Schema Inference**: Detect column types (string, integer, number, boolean, date, email, url)
- **Online Statistics**: Uses Welford's algorithm for efficient single-pass statistics
## Development
```bash
# Run tests
npm test
# Build
npm run build
# Watch mode
npm run dev
```
## License
MIT
TDQS
Scored across 12 tools
Tools are largely distinct with clear purposes. csv_search and csv_filter both retrieve matching records but differ in query syntax (regex vs expression), which could cause some initial confusion. csv_inspect and csv_schema provide different levels of overview, but descriptions clarify the difference.
All tools share the csv_ prefix and snake_case, creating a recognizable family. However, naming style is inconsistent: some use verbs (csv_filter, csv_validate), some nouns (csv_schema, csv_stats), and some adjective_noun (csv_large_fields). This slight inconsistency prevents a perfect score.
At 12 tools, the count is well-suited to the server's purpose of comprehensive CSV exploration. Each tool addresses a distinct aspect, and the count is neither sparse nor bloated.
The tool set covers the core lifecycle of CSV analysis: inspection, sampling, schema, stats, search/filter, validation, monitoring, diffing, and handling large fields. Minor gaps include lack of sorting or column manipulation, but these are beyond the apparent scope.