A powerful MCP tool for parsing and manipulating MIDI files that allows users to read, analyze, and modify MIDI files through natural language commands, supporting operations like reading file information, modifying tracks, adding notes, and setting tempo.
An AI-powered file management system that integrates with Claude Desktop to provide intelligent file analysis, multimedia processing, and automatic organization.
Enables file operations such as counting, listing, compressing images, creating/extracting archives, copying/moving files, and merging/splitting PDFs via natural language.
Enables interactive chat with AI models via Anthropic API, with controlled file system access to specified directories and video conversion capabilities.
Enables playing and inspecting local audio files in an MCP host with an in-conversation UI showing waveform, spectrogram, and loudness metrics, while exposing playback state and metadata to the model.
Converts various document formats to desired output formats, currently supporting PDF to image conversion. No access keys required for basic file format conversion operations.
Enables Claude AI to control Ableton Live and Max for Live, allowing music production tasks like track management, MIDI editing, and pattern generation directly from conversation.
Wallet-funded MCP client for six paid Utilia tools: Solana priority fees, transaction diagnosis and simulation, token-risk checks, PDF-to-Markdown, and audio normalization over x402. MIT-licensed and installable from npm.
Parse any file or URL into structured text. Extract text from PDF, DOCX, YouTube, web pages, images, and 25+ formats via one API. Tools: parse_url, parse_file, get_youtube_transcript.
MCP server for controlling the mpv media player, enabling playback control, music library browsing, YouTube streaming and downloading, and metadata editing from within an MCP client.
A Model Context Protocol server that converts diverse file types, including PDFs, images, audio, and Office documents, into Markdown format. It also transforms web content like YouTube transcripts and Bing search results into readable text for model consumption.
Enables advanced audio transcription, text-to-speech generation, and audio processing using OpenAI's Whisper and GPT-4o models with support for multiple audio formats, file management, and parallel processing.