Skip to main content
Glama
README.md
# CSV Explorer MCP Server

A Model Context Protocol (MCP) server for exploring and analyzing CSV files. Provides tools for inspection, sampling, schema inference, statistics, filtering, and more.

## Installation

```bash
npm install
npm run build
```

## Usage

Add to your MCP configuration:

```json
{
  "mcpServers": {
    "csv-explorer": {
      "command": "node",
      "args": ["path/to/dist/index.js"]
    }
  }
}
```

## Tools

### csv_inspect

Get an overview of a CSV file including size, row/column count, detected delimiter, and a preview of the data. Large field values are automatically truncated with content-type hints.

```typescript
csv_inspect({ file: "/path/to/data.csv", previewRows: 5 })
```

### csv_sample

Get sample records using various sampling strategies.

```typescript
csv_sample({ file: "/path/to/data.csv", mode: "random", count: 10 })
// modes: "first", "last", "random", "range"
```

### csv_schema

Infer the schema by sampling records. Returns column names, types, and nullability.

```typescript
csv_schema({ file: "/path/to/data.csv", sampleSize: 1000 })
// outputFormat: "inferred", "json-schema", "formatted"
```

### csv_stats

Collect aggregate statistics for fields. Includes min/max, mean, median, stdDev for numeric fields, and top values for categorical fields.

```typescript
csv_stats({ file: "/path/to/data.csv", fields: ["price", "category"] })
```

### csv_search

Search for records where a field matches a regex pattern.

```typescript
csv_search({ file: "/path/to/data.csv", field: "email", pattern: "@example\\.com$" })
```

### csv_filter

Filter records using query expressions. Supports comparisons (`==`, `!=`, `<`, `>`, `<=`, `>=`), text operations (`contains`, `startswith`, `endswith`, `matches`), and compound queries (`AND`, `OR`).

```typescript
csv_filter({ file: "/path/to/data.csv", query: 'status == "active" AND age > 30' })
```

### csv_validate

Validate a CSV file for syntax errors and optionally against a schema.

```typescript
csv_validate({
  file: "/path/to/data.csv",
  schema: {
    columns: [
      { name: "id", type: "integer", required: true },
      { name: "email", type: "string", pattern: "^[^@]+@[^@]+$" }
    ]
  }
})
```

### csv_tail

Read new records appended since a cursor position. Use for monitoring actively-written files.

```typescript
csv_tail({ file: "/path/to/data.csv", cursor: 1024, maxRecords: 100 })
```

### csv_get_cursor

Get the current end-of-file position for use with `csv_tail`.

```typescript
csv_get_cursor({ file: "/path/to/data.csv" })
```

### csv_diff

Compare two CSV files and report differences.

```typescript
csv_diff({ file1: "/path/to/old.csv", file2: "/path/to/new.csv", keyField: "id" })
```

### csv_extract

Extract a specific field value from a CSV record. Use for retrieving large/truncated field data. Can write to file for binary data (e.g., base64 images).

```typescript
// Get field value inline
csv_extract({ file: "/path/to/data.csv", field: "description", line: 5 })

// Decode base64 and write to file
csv_extract({
  file: "/path/to/data.csv",
  field: "screenshot",
  line: 1,
  decode: "base64",
  outputFile: "/tmp/screenshot.png"
})
```

### csv_large_fields

List fields containing large values (e.g., base64 images, JSON blobs). Helps identify which fields were truncated in `csv_inspect`.

```typescript
csv_large_fields({ file: "/path/to/data.csv", threshold: 1000, sampleRows: 100 })
```

## Features

- **Streaming Architecture**: Memory-efficient processing of large files
- **Auto-Detection**: Automatically detects delimiters (comma, tab, semicolon, pipe) and encoding
- **Smart Truncation**: Large field values are truncated with content-type hints (base64, JSON, HTML)
- **Query Engine**: Filter records with SQL-like expressions supporting AND/OR logic
- **Schema Inference**: Detect column types (string, integer, number, boolean, date, email, url)
- **Online Statistics**: Uses Welford's algorithm for efficient single-pass statistics

## Development

```bash
# Run tests
npm test

# Build
npm run build

# Watch mode
npm run dev
```

## License

MIT

TDQS

A3.5/5.0

Scored across 12 tools

Disambiguation4/5

Tools are largely distinct with clear purposes. csv_search and csv_filter both retrieve matching records but differ in query syntax (regex vs expression), which could cause some initial confusion. csv_inspect and csv_schema provide different levels of overview, but descriptions clarify the difference.

Naming Consistency4/5

All tools share the csv_ prefix and snake_case, creating a recognizable family. However, naming style is inconsistent: some use verbs (csv_filter, csv_validate), some nouns (csv_schema, csv_stats), and some adjective_noun (csv_large_fields). This slight inconsistency prevents a perfect score.

Tool Count5/5

At 12 tools, the count is well-suited to the server's purpose of comprehensive CSV exploration. Each tool addresses a distinct aspect, and the count is neither sparse nor bloated.

Completeness4/5

The tool set covers the core lifecycle of CSV analysis: inspection, sampling, schema, stats, search/filter, validation, monitoring, diffing, and handling large fields. Minor gaps include lack of sorting or column manipulation, but these are beyond the apparent scope.

Maintenance

ActivityInactive
ResponsivenessNo issues