CSV Insight & Cleaner MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@CSV Insight & Cleaner MCPInspect this CSV and flag any quality issues"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
CSV Insight & Cleaner MCP ๐๐งน
CSV Insight & Cleaner is a minimalistic Model Context Protocol (MCP) server that lets an AI assistant understand, quality-check, and clean CSV datasets โ without loading raw file paths or guessing at missing data.
Built with the official Python MCP SDK (FastMCP) and pandas, in a small,
readable, single-file codebase.
๐ Key Highlights
โก Minimal & Zero-Bloat โ one
server.py, under 150 lines.๐ฏ All 3 MCP Primitives: Tools (inspect/summarize/clean/preview/report/export), Resources (live dataset + report snapshots), Prompts (analysis & cleaning workflows).
๐ง Facts vs. Explanation split โ Python/pandas computes the facts (row counts, duplicates, missing values); the AI client turns those facts into natural language. The server never invents or silently guesses data.
๐ Safe by design โ missing values are reported, never auto-filled. Only deterministic, reversible cleanup (duplicates, empty rows, whitespace, column names) is automated.
๐ Deploy-friendly โ tools accept raw CSV text content, not local file paths, so the same server works locally and once deployed publicly (e.g. on Glama), where it has no access to your filesystem.
Related MCP server: mcp-csv-analyst
๐ Project Structure
csv-insight-mcp/
โโโ sample_data/
โ โโโ messy_sales.csv # sample dataset for demos
โโโ src/
โ โโโ csvinsight/
โ โโโ __init__.py # package exports
โ โโโ __main__.py # `python -m csvinsight` entrypoint
โ โโโ server.py # core server (Tools, Resources, Prompts)
โโโ tests/
โ โโโ test_server.py # pytest smoke tests
โโโ .vscode/
โ โโโ mcp.json # VS Code Copilot Chat MCP config
โโโ pyproject.toml
โโโ README.md๐ Architecture Overview
+-------------------------------------------------------------------------------+
| MCP CLIENT |
| (VS Code Copilot Chat / Claude Desktop / Cursor IDE / Custom AI) |
+-------------------------------------------------------------------------------+
โฒ
โ JSON-RPC 2.0 (stdio)
โผ
+-------------------------------------------------------------------------------+
| CSV INSIGHT & CLEANER MCP SERVER |
| |
| [TOOLS] [RESOURCES] [PROMPTS] |
| โข inspect_csv โข csv://current โข analyze_dataset |
| โข summarize_csv (dataset snapshot) โข clean_and_report |
| โข preview_csv โข csv://cleaning-report |
| โข clean_csv (before/after report) |
| โข get_cleaning_report |
| โข export_cleaned_csv |
+-------------------------------------------------------------------------------+
โ
โผ
Python / pandas Engine
(in-memory dataframe, per session)๐ ๏ธ MCP Primitives Catalog
1. Tools
Tool Name | Parameters | Description |
|
| Loads raw CSV text and returns rows, columns, dtypes, missing values, duplicates, sample rows. |
| none | Returns the same structural facts for the currently loaded dataset. |
|
| Returns the first N rows of the original or cleaned dataset. |
|
| Deterministic cleanup; missing values are reported, never guessed. |
| none | Returns the before/after report from the last |
| none | Returns the cleaned dataset as raw CSV text, ready to save. |
2. Resources
Resource URI | Description |
| Markdown snapshot of the currently loaded dataset (rows, columns, quality). |
| Markdown before/after report from the most recent cleaning. |
3. Prompts
Prompt Name | Description |
| Workflow: inspect the dataset, explain what it represents, flag quality issues, give recommendations. |
| Workflow: clean the dataset, then explain exactly what changed. |
๐ Quickstart
# Clone and enter the project
git clone https://github.com/your-username/csv-insight-mcp.git
cd csv-insight-mcp
# Install in editable mode
pip install -e .
# Run tests
pytest -v
# Run the server directly over stdio
python -m csvinsight๐ Client Configuration
VS Code Copilot Chat
Already included at .vscode/mcp.json:
{
"servers": {
"csv-insight-cleaner": {
"type": "stdio",
"command": "python",
"args": ["-m", "csvinsight"],
"cwd": "${workspaceFolder}"
}
}
}Open Copilot Chat โ switch to Agent mode.
Click the tools icon โ confirm
csv-insight-cleanertools are listed.Try: "Load sample_data/messy_sales.csv and tell me what's wrong with it." (paste the file contents, or ask Copilot to read the file and pass its text into
inspect_csv.)
Claude Desktop
claude_desktop_config.json:
{
"mcpServers": {
"csv-insight-cleaner": {
"command": "python",
"args": ["-m", "csvinsight"],
"cwd": "/absolute/path/to/csv-insight-mcp"
}
}
}๐ฌ Example Demo Flow
"Analyze this CSV." (paste contents of messy_sales.csv)
โ inspect_csv โ rows, columns, 1 duplicate row, 1 empty row, 2 missing emails
"What problems does it have?"
โ AI explains the data-quality section in plain language
"Clean the safe issues."
โ clean_csv โ duplicates & empty rows removed, columns standardized
"What did you change?"
โ get_cleaning_report โ before/after row counts + list of changes
"Give me the cleaned file."
โ export_cleaned_csv โ ready-to-save CSV textโ๏ธ Deployment (Glama)
The server only ever receives CSV text content through its tool parameters โ never a local file path โ so it is safe to deploy publicly:
Push this repository to GitHub.
Go to glama.ai/mcp/servers โ Add Server.
Authenticate with GitHub and submit the repo URL.
Glama builds and verifies MCP compliance automatically.
(Smithery is also compatible if preferred โ add a smithery.yaml pointing
at the same python -m csvinsight entrypoint.)
๐งช Testing
pytest -v๐ License
MIT License ยฉ 2026 CSV Insight & Cleaner Contributors.
Available Tools
6 toolsclean_csvA
Perform safe, deterministic cleaning (duplicates, empty rows, whitespace, column names). Missing values are reported, never guessed.
| Name | Required | Description | Default |
|---|---|---|---|
| drop_duplicates | No | ||
| drop_empty_rows | No | ||
| trim_whitespace | No | ||
| standardize_columns | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It adds key behavioral traits: 'safe, deterministic' and 'Missing values are reported, never guessed,' which clarify that the tool won't invent data. However, it does not disclose whether it modifies the original CSV or returns a cleaned copy, a notable omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short, well-structured sentences. It front-loads the primary action and then adds the crucial safety and missing-value behavior in a second sentence, with no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema and no annotations, so the description should explain what the tool returns or whether it mutates input. It does not mention output or side effects, leaving a gap for an agent to understand the tool's full I/O contract. Given the simple nature of the tool, it's adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description explicitly lists the operations (duplicates, empty rows, whitespace, column names) that map directly to the four boolean parameters. The parameter names are self-explanatory, and the description effectively compensates for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs CSV cleaning with specific operations (duplicates, empty rows, whitespace, column names), using a specific verb ('clean'). This distinguishes it from siblings like inspect_csv, preview_csv, and summarize_csv, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for cleaning a CSV before export or analysis, and the listed operations give a clear context of when to apply it. However, it does not explicitly state when not to use it or mention alternatives, leaving some implicitness.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_cleaned_csvA
Return the cleaned dataset as raw CSV text, ready to save/download.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden of behavioral transparency. It clearly states the output is raw CSV text, which is helpfulcluster, but it does not disclose any potential side effects (e.g., whether it consumes or deletes the cleaned data) or performance characteristics. Given the tool is likely a read-only export, the absence of side effects is not a concern, but the description could be more explicit about it being non-destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that clearly states the tool's purpose and output. It contains no unnecessary words or repetition, making it highly efficient and appropriately front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters, an output schema is present (which likely describes the CSV structure), and the description is simple, the description is largely sufficient. It lacks a few details such as confirming the output is the result of a prior cleaning step or any encoding specifics, but these are minor given the simplicity. The output schema likely covers return value details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema confirms this with no properties defined. Schema description coverage is 100%, and there is nothing to document for parameters. The description adds no parametric information, but baseline is 4 for zero-parameter tools, and the description appropriately avoids inventing parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns the cleaned dataset as raw CSV text for saving/downloading, which aligns with its name. It distinguishes itself from siblings by focusing on export, whereas siblings like inspect_csv, preview_csv, and clean_csv serve different purposes. However, it does not explicitly contrast with get_cleaning_report, which could also be an output tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly suggests using this tool when you need the final cleaned data in a downloadable format, which is clear enough. However, it does not explicitly state when to use it over siblings like preview_csv or get_cleaning_report, nor does it mention any prerequisites like having cleaned the data first. This leaves some ambiguity for the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_cleaning_reportA
Return the before/after report from the most recent clean_csv call.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It tells us the report is a result of the most recent clean_csv call, but doesn't disclose what the report contains, how long it's retained, or if it errors if no clean_csv has been called.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, clear sentence. No wasted words. Perfectly concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter getter tool, the description does its job, but could mention what the report contains or what happens if no clean_csv call has been made. Since there's no output schema, a bit more detail on the report structure would have been helpful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0 parameters, there is no parameter semantics burden on the description. The description correctly focuses on the return value since there's nothing else to explain.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool returns a before/after report from the most recent clean_csv call, which is specific about what it returns. However, it could be clearer on what 'before/after' encompasses and how it differs from siblings like inspect_csv.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for after a clean_csv call, but it doesn't explicitly state when to use it versus alternatives. However, the phrase 'most recent clean_csv call' provides useful context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_csvA
Load a dataset and return structural facts: row/column counts, dtypes, missing values, duplicates, and a sample.
Provide ONE of:
file_path: local path to a .csv, .tsv, or Excel (.xlsx/.xls) file.
csv_data: raw CSV/TSV text content (for when there's no file access).
| Name | Required | Description | Default |
|---|---|---|---|
| csv_data | No | ||
| filename | No | ||
| file_path | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry full disclosure burden. It mentions the outputs and the two input options, but entirely ignores the 'filename' parameter present in the schema. It also does not state what happens if both file_path and csv_data are provided or if none are provided, leaving significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, with the main purpose stated first and input options listed clearly. It is front-loaded with the most important information and contains no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The purpose and input methods are clear, and the outputs are described. However, the unexplained 'filename' parameter and lack of edge-case handling (e.g., what occurs if both inputs are provided) make the description incomplete for a tool with multiple input modes and no annotations. It could also better differentiate from sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning to file_path (supported formats) and csv_data (raw text content), but provides no explanation for the 'filename' parameter. With schema descriptions entirely absent (0% coverage), this partial coverage is insufficient for complete parameter understanding, though it does help for two of the three parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool 'Load a dataset and return structural facts' with specifics like row/column counts, dtypes, missing values, duplicates, and a sample. This is a specific verb-resource pairing and distinguishes from siblings like summarize_csv or preview_csv by emphasizing structural facts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit input guidance: 'Provide ONE of: file_path or csv_data' and explains when csv_data is appropriate ('when there's no file access'). It does not explicitly name alternatives or state when not to use this tool, but the context is clear enough for an agent to infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_csvB
Preview the first N rows of the 'original' or 'cleaned' dataset.
| Name | Required | Description | Default |
|---|---|---|---|
| rows | No | ||
| which | No | original |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears full responsibility. 'Preview' accurately conveys a read-only, non-destructive operation, which is consistent with the tool name and parameters. However, it never explicitly states that no data is modified or that it's safe to call repeatedly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is perfectly sized for this simple toolโfront-loaded with the verb 'Preview' and containing no filler or redundant content. Every word in the description earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with defaults and no output schema, the description is mostly adequate. However, it doesn't explain the return value format (e.g., does it return a table, a string?) or behavior with edge cases like rows > available rows. The sibling tools suggest a cleaning workflow where the original/cleaned split is meaningful, which the description alludes to but doesn't explain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite 0% schema description coverage, the description effectively clarifies both parameters: 'N' maps to the rows integer (default 5) and '"original" or "cleaned"' maps to the which string (default "original"). The defaults are implied but not explicitly repeated in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Preview the first N rows of the original or cleaned dataset.' It identifies both the operation (preview) and the target (first N rows), and even hints at the 'which' parameter by naming the two datasets ('original' or 'cleaned'). A single point is lost because it doesn't explicitly state why previewing is useful.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus its siblings like inspect_csv or summarize_csv. It doesn't explain the typical workflow position (e.g., 'quickly verify data before cleaning') or any prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
summarize_csvC
Return the same structural facts as inspect_csv for the currently loaded dataset, for the AI to turn into a natural-language summary.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since no annotations are provided, the description carries the full burden of behavioral disclosure. It mentions 'structural facts' and that it returns data for the AI to summarize, but it does not clarify any side effects, prerequisites (e.g., a dataset must be loaded), or what happens if no dataset is loaded. The behavior is not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, concise and front-loaded with the key purpose. It is not overly verbose, but it could be slightly improved by clarifying the tool's role relative to inspect_csv. However, it is appropriately sized for a parameterless tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters and no output schema, the description must explain what it returns. It mentions 'structural facts' but does not specify what those facts are (e.g., column names, data types, row count) or how they differ from inspect_csv. Without annotations or an output schema, the description is incomplete for an agent to understand the tool's full behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is nothing for the schema to describe. The description correctly indicates that it operates on the currently loaded dataset, which is the implicit context. This baseline of 4 is appropriate per the rubric for tools with no parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool returns the same structural facts as inspect_csv for a summary, which is clear about its purpose, but it does not explicitly differentiate from sibling tools such as inspect_csv or preview_csv beyond implying it's for summary generation. It lacks a specific verb like 'generate' or 'produce', but the intent is understandable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool versus alternatives. It references inspect_csv but does not state when to choose one over the other, nor does it mention any context like 'after loading a dataset' or 'when a natural-language summary is needed'. This is minimal guidance that relies on inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v1.0.0- First observed
clean_csv - First observed
export_cleaned_csv - First observed
get_cleaning_report - First observed
inspect_csv - First observed
preview_csv - First observed
summarize_csv
TDQS
Scored across 6 tools
inspect_csv and summarize_csv both return structural facts about the dataset, differing only in whether the data is freshly loaded or already in memory. preview_csv also overlaps with the sample output provided by inspect_csv, creating potential ambiguity for an agent.
All six tools follow the consistent verb_noun pattern in snake_case (e.g., inspect_csv, export_cleaned_csv), making the naming predictable and easy to read.
Six tools is a well-scoped count for a CSV cleanup and inspection server, covering core workflows without redundancy or bloat.
The tool surface covers loading, inspection, cleaning, reporting, preview, and export, which satisfies the primary use cases. Minor gaps exist, such as no way to customize cleaning parameters or explicitly unload data, but these are not critical.
Maintenance
Related MCP Connectors
MCP server that lets AI assistants use all OneSchema features exposed via the public API.
Nifty's MCP server โ exposes tasks, projects, messages, and files as tools for AI agents.
Query, join, profile, clean and convert CSV/JSON/Parquet with server-side DuckDB over MCP.
An MCP server that gives your AI access to the source code and docs of all public github repos
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAn MCP server that provides AI assistants with structured, type-safe access to tabular datasets from CSV files. It enables users to list, describe, and query data using filters and projections with support for hot reloading.10 npmMIT
- AlicenseAqualityNot gradedmaintenanceAn MCP server that enables AI assistants to load, query, and analyze local CSV files using tools for filtering, aggregation, and grouping. It provides capabilities to describe schemas, calculate statistics, and sample data directly from CSV files.6-
- FlicenseNot gradedqualityDmaintenanceA local MCP server for analyzing CSV files from your filesystem, particularly suited for chatbot conversation logs. Allows listing, reading, filtering, merging, and statistical analysis of CSV data via natural language.-
- FlicenseAqualityDmaintenanceAn MCP server for dataset exploration and analysis, enabling LLM clients to perform summary, correlation, distribution, missing value analysis, data cleaning, and statistical tests directly on CSV files.3-