Tika MCP Server
Extracts content and metadata from files using Apache Tika, supporting various file formats such as PDF, DOCX, and images with OCR.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Tika MCP Serverextract text from the uploaded report.pdf"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Tika MCP Server
This project provides a Model Context Protocol (MCP) server for extracting content and metadata from files using Apache Tika.
Overview
The Tika MCP server allows AI assistants to extract text and metadata from various file formats (PDF, DOCX, images with OCR, etc.) using Apache Tika. This enables AI assistants to understand and work with the content of files that users upload.
Related MCP server: md-converter-mcp
Features
Extract text content from various file formats
Extract metadata (author, creation date, etc.) from files
Support for PDF, DOCX, images, and many other formats
Simple JSON-RPC API following the Model Context Protocol
Requirements
Python 3.6+
Apache Tika server running (default: http://localhost:9998)
MCP-compatible client
Installation
Clone this repository
Install dependencies:
pip install -r requirements.txtRegister the MCP server:
python -m app.register_mcp_server
Usage
The Tika MCP server provides a single tool:
extract_file
Extracts content and metadata from a file using Apache Tika.
Parameters:
file_path: Path to the file to extract content fromtika_url: URL of the running Tika server (default: http://localhost:9998)
Returns:
metadata: Dictionary of metadata extracted from the filecontent: Array of content blocks extracted from the file
Testing
Several test scripts are provided to verify the functionality:
app/test_tika_simple.py: Tests the Tika client directlyapp/test_simple_mcp.py: Tests the MCP server using the JSON-RPC protocol
Project Structure
app/: Main application codesimple_mcp_server.py: MCP server implementationtika_client.py: Client for Apache Tikamodel.py: Data models and business logicregister_mcp_server.py: Script to register the MCP server
examples/: Example files for testingrequirements.txt: Python dependencies
Setup
Get a venv using either:
uv venvor
python3 -m venv .venvActivate the virtual environment and install dependencies:
source .venv/bin/activate
pip install -r requirements.txtRunning the MCP Server
Start the Apache Tika server (if not already running):
docker run -d -p 9998:9998 apache/tikaRegister and run the MCP server:
python -m app.register_mcp_serverLicense
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Document conversion and OCR for AI agents: PDF, Office docs, images to text.
Convert PDF, Word, Excel and scanned documents to Markdown, tables and RAG chunks. OCR.
Parse PDF/Word/PPT/HTML to Markdown; tables as JSON, image extraction, RAG chunking, page ranges.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to read and extract text content from PDF, Excel, and Word documents via the Model Context Protocol.2MIT
- FlicenseAqualityDmaintenanceConverts files (PDF, DOCX, PPTX, XLSX, images via OCR) and URLs to Markdown, enabling AI clients to read them via a single MCP tool.1-
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to parse and search documents including PDF, Word, Excel, PowerPoint, and images via OCR, with support for semantic search and batch processing.2MIT
- AlicenseNot gradedqualityAmaintenanceConverts documents (PDF, DOCX, XLSX, EPUB, etc.) to clean, structured Markdown, and retrieves document info, for use with AI agents.MIT