content-core
Content Core
統一された非同期Python API、CLI、またはMCPサーバーを通じて、URL、ファイル、テキストからコンテンツを抽出、処理、要約します。
サポートされているフォーマット
カテゴリ | フォーマット |
Web | URL、HTMLページ、YouTube動画、Reddit投稿 |
ドキュメント | PDF、DOCX、PPTX、XLSX、EPUB、Markdown、プレーンテキスト |
メディア | MP3、WAV、M4A、FLAC、OGG (音声); MP4、AVI、MOV、MKV (動画) |
Related MCP server: FreeCrawl MCP Server
クイックスタート
pip install content-coreimport content_core
result = await content_core.extract_content(url="https://example.com")
print(result.content)またはインストール不要で実行:
uvx content-core extract "https://example.com"CLIの使用方法
Content Coreは、抽出、要約、MCPサーバー用のサブコマンドを備えた統一された content-core コマンドを提供します。
抽出
# From a URL
content-core extract "https://example.com"
# From a file
content-core extract document.pdf
# With JSON output
content-core extract document.pdf --format json
# With a specific engine
content-core extract "https://example.com" --engine firecrawl
# From stdin
echo "some text" | content-core extract要約
# Summarize text
content-core summarize "Long article text here..."
# With context
content-core summarize "Long text" --context "bullet points"
# From stdin
cat article.txt | content-core summarize --context "explain to a child"MCPサーバー
content-core mcp設定
# Set persistent config
content-core config set llm_provider anthropic
content-core config set llm_model claude-sonnet-4-20250514
# List current config
content-core config list
# Delete a config value
content-core config delete llm_provider設定は ~/.content-core/config.toml に保存されます。優先順位: コマンドフラグ > 環境変数 > 設定ファイル > デフォルト値。
uvxによるインストール不要の実行
すべてのコマンドは uvx を使用してインストールなしで動作します:
uvx content-core extract "https://example.com"
uvx content-core summarize "text" --context "one sentence"
uvx content-core mcpPython API
抽出
import content_core
# From a URL
result = await content_core.extract_content(url="https://example.com")
# From a file
result = await content_core.extract_content(file_path="document.pdf")
# From text
result = await content_core.extract_content(content="some text")
# With engine override
from content_core import ContentCoreConfig
config = ContentCoreConfig(url_engine="firecrawl")
result = await content_core.extract_content(url="https://example.com", config=config)要約
import content_core
summary = await content_core.summarize("long article text", context="bullet points")設定
from content_core import ContentCoreConfig
config = ContentCoreConfig(
url_engine="firecrawl",
document_engine="docling",
audio_concurrency=5,
)
result = await content_core.extract_content(url="https://example.com", config=config)MCP統合
Content Coreには、Claude Desktopやその他のMCP互換アプリケーションで使用するためのModel Context Protocol (MCP) サーバーが含まれています。
claude_desktop_config.json に以下を追加してください:
{
"mcpServers": {
"content-core": {
"command": "uvx",
"args": ["content-core", "mcp"],
"env": {
"OPENAI_API_KEY": "sk-..."
}
}
}
}MCPサーバーは extract_content と summarize_content の2つのツールを公開します。どちらもプレーンテキストを返します。
詳細なセットアップについては、MCPドキュメントを参照してください。
Claude Codeスキル
Content Coreには、AIエージェントが外部ソースからコンテンツを抽出する方法を教える SKILL.md が含まれています。Claude Codeプロジェクトで利用できるようにするには、スキルディレクトリにコピーしてください:
# Download the skill
curl -o .claude/skills/content-core/SKILL.md --create-dirs \
https://raw.githubusercontent.com/lfnovo/content-core/main/SKILL.mdインストールが完了すると、Claude CodeはCLI (uvx content-core) または設定済みの場合はMCPを介して、content-coreを使用してURL、ドキュメント、メディアファイルからコンテンツを抽出できるようになります。
AIプロバイダー
Content Coreは Esperanto を使用して、複数のLLMおよびSTTプロバイダーをサポートしています。設定を変更するだけでプロバイダーを切り替えることができ、コードの変更は不要です:
# Use Anthropic for summarization
content-core config set llm_provider anthropic
content-core config set llm_model claude-sonnet-4-20250514
# Use Groq for transcription
content-core config set stt_provider groq
content-core config set stt_model whisper-large-v3サポートされているプロバイダーには、OpenAI、Anthropic、Google、Groq、DeepSeek、Ollamaなどがあります。全リストについては Esperantoドキュメント を参照してください。
設定
Content Coreは pydantic-settings を利用した ContentCoreConfig を使用します。設定は優先順位の高い順に解決されます: コンストラクタ引数 > 環境変数 (CCORE_*) > 設定ファイル (~/.content-core/config.toml) > デフォルト値。
環境変数
変数 | 説明 | デフォルト |
| URL抽出エンジン ( |
|
| ドキュメント抽出エンジン ( |
|
| 同時音声文字起こし数 (1-10) |
|
| Crawl4AI Docker API URL (ローカルブラウザモードの場合は省略) | - |
| セルフホストインスタンス用のカスタムFirecrawl API URL | - |
| Firecrawlプロキシモード ( |
|
| 抽出前の待機時間 (ms) |
|
| 要約用LLMプロバイダー | - |
| 要約用LLMモデル | - |
| 音声認識(STT)プロバイダー | - |
| 音声認識(STT)モデル | - |
| 音声認識(STT)タイムアウト (秒) | - |
| YouTube文字起こしの優先言語 | - |
外部サービスのAPIキーは、標準の環境変数 (例: OPENAI_API_KEY, FIRECRAWL_API_KEY, JINA_API_KEY) を介して設定されます。
プロキシ設定
Content Coreは、標準の HTTP_PROXY / HTTPS_PROXY / NO_PROXY 環境変数を自動的に読み取ります。追加の設定は不要です。
オプションの依存関係
# Docling for advanced document parsing (PDF, DOCX, PPTX, XLSX)
pip install content-core[docling]
# Crawl4AI for local browser-based URL extraction
pip install content-core[crawl4ai]
python -m playwright install --with-deps
# LangChain tool wrappers
pip install content-core[langchain]
# All optional features
pip install content-core[docling,crawl4ai,langchain]LangChainでの使用
langchain エクストラをインストールすると、Content CoreはLangChain互換のツールラッパーを提供します:
from content_core.tools import extract_content_tool, summarize_content_tool
tools = [extract_content_tool, summarize_content_tool]ドキュメント
開発
git clone https://github.com/lfnovo/content-core
cd content-core
uv sync --group dev
# Run tests
make test
# Lint
make ruffライセンス
このプロジェクトは MITライセンス の下でライセンスされています。
貢献
貢献を歓迎します!詳細については 貢献ガイド を参照してください。
Available Tools
2 toolsextract_contentB
Extract content from a URL or file. Does not require an API key for most sources (web pages, PDFs, documents, YouTube transcripts). API key is only needed for audio/video transcription.
Args:
url: URL to extract content from (web page, YouTube video, PDF link, etc.)
file_path: Local file path to extract content from
engine: Optional extraction engine override, routed by input type.
With url: auto, simple, firecrawl, jina, crawl4ai.
With file_path: auto, simple, docling — docling requires
pip install "content-core[docling]" and fails with a
configuration error when the extra is missing, in which case use
auto or simple.
Any other value is rejected with an error naming the accepted ones.
formulas: Enable formula extraction via Docling (requires engine=docling)
pictures: Enable image description + chart data extraction via Docling (requires engine=docling)
no_ocr: Disable OCR in Docling (requires engine=docling)
Returns: Extracted text content
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| engine | No | ||
| no_ocr | No | ||
| formulas | No | ||
| pictures | No | ||
| file_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It discloses API key requirements, engine behavior, and the docling extra failure mode, but it does not explicitly state that the operation is read-only or describe the return format beyond 'Extracted text content'. The engine error message is useful but other behavioral aspects remain implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized with an Args/Returns structure and front-loads the purpose. It is detailed but each line earns its place, covering engine specifics and error conditions without excessive verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with six optional parameters and no required ones, the description omits guidance on whether at least one of url or file_path must be provided. The return description is minimal, and error handling for missing inputs is not covered. While engine behavior is well documented, these input-requirement gaps reduce completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains url, file_path, engine with valid values and routing, and clarifies that formulas, pictures, and no_ocr require engine=docling. This adds substantial meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool extracts content from a URL or file and lists common source types. It does not explicitly contrast with the sibling summarize_content, but the verb 'extract' and the scope are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus summarize_content. While it explains engine selection and API key conditions, it never addresses tool-level choice, which is a gap given the sibling exists.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
summarize_contentA
Summarize content using LLM with optional context. Requires OPENAI_API_KEY (or another LLM provider key) to be configured.
Args: content: The text content to summarize context: Optional context to guide summarization (e.g., "summarize as bullet points")
Returns: Summarized text
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | ||
| context | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description bears the full burden of behavioral disclosure. It notes the requirement for an LLM provider API key and indicates that the tool uses an LLM for summarization. However, it does not disclose potential rate limits, costs, or failure modes, leaving some behavioral aspects opaque.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description uses a docstring format with sections (Args, Returns), making it structured but slightly verbose. It front-loads the core purpose but adds extra formatting that could be trimmed. It is not overly long but could be more concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low schema coverage and absence of annotations, the description provides the essential parameter meanings and return type. It also mentions the critical API key dependency. However, it lacks constraints like maximum content length or edge cases, and the output schema existence lightens the burden but doesn't fully compensate for missing details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (no descriptions in input schema), but the description adds meaningful explanations for both parameters: 'content' is the text to summarize, and 'context' is optional guidance with an example ('summarize as bullet points'). This compensates well for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Summarize content using LLM with optional context.' It specifies a specific verb-resource relationship and distinguishes from the sibling tool 'extract_content' which serves a different purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions a prerequisite (API key configuration) but provides no guidance on when to use this tool versus alternatives, such as the sibling 'extract_content'. No explicit when-to-use or when-not-to-use guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v2.0.4- Changed
extract_content9 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / engineAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null +} - added
Input schema / properties / formulasAdded value: +{ + "default": false, + "type": "boolean" +} - added
Input schema / properties / no_ocrAdded value: +{ + "default": false, + "type": "boolean" +} - added
Input schema / properties / picturesAdded value: +{ + "default": false, + "type": "boolean" +} - removed
Output schema / additionalPropertiesRemoved value: -true - added
Output schema / propertiesAdded value: +{ + "result": { + "type": "string" + } +} - added
Output schema / requiredAdded value: +[ + "result" +] - added
Output schema / x-fastmcp-wrap-resultAdded value: +true
- Added
summarize_content
1 tool update
v1.0.0- Changed
extract_content2 fields changed- removed
Input schema / properties / file_path / titleRemoved value: -"File Path" - removed
Input schema / properties / url / titleRemoved value: -"Url"
1 tool update
- First observed
extract_content
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes—extraction vs. summarization—with no functional overlap. Each tool's parameters are also well-differentiated, avoiding ambiguity.
Both tool names follow the same verb_noun pattern (extract_content, summarize_content), providing a predictable and consistent naming convention.
With only 2 tools, the set is borderline thin for a content-processing server. While both are useful, the count is minimal and could be expanded with additional content operations.
The tools cover the core content workflow of extraction and summarization, but lack other common operations (e.g., translation, keyword extraction) that would round out a comprehensive content toolkit. Minor gaps exist.
Maintenance
Related MCP Connectors
Extract structured insights from videos, podcasts, articles, and PDFs with multi-model AI
Extract and parse web pages into clean HTML, links, or Markdown. Handle dynamic, complex, or block…
Parse PDF/Word/PPT/HTML to Markdown; tables as JSON, image extraction, RAG chunking, page ranges.
Turns any URL into SEO metadata, contacts, tech stack, and AI-ready Markdown, in one call.
Related MCP Servers
- AlicenseCqualityCmaintenanceA powerful tool for fetching and extracting text content from web pages and APIs, supporting web scraping, REST API requests, and Google Custom Search integration.510MIT
- AlicenseNot gradedqualityDmaintenanceEnables web scraping and document processing with JavaScript execution, anti-detection measures, batch processing, and structured data extraction. Supports multiple formats including markdown, HTML, screenshots, and handles PDFs with OCR capabilities.4MIT
- AlicenseBqualityCmaintenanceParse any file or URL into structured text. Extract text from PDF, DOCX, YouTube, web pages, images, and 25+ formats via one API. Tools: parse_url, parse_file, get_youtube_transcript.38 npmMIT
- AlicenseNot gradedqualityCmaintenanceConverts any URL into clean, LLM-ready Markdown, text, or HTML with production-grade features like SSRF protection, rate limiting, retries, caching, and structured error handling.MIT