Skip to main content
Glama

smart-data-extractor

Server Details

smart-data-extractor MCP server on Cloudflare Workers · REST + MCP JSON-RPC · free tier

Status
Healthy
Last Tested
Transport
Streamable HTTP
URL
Repository
lazymac2x/smart-data-extractor-worker
GitHub Stars
0

Glama MCP Gateway

Connect through Glama MCP Gateway for full control over tool access and complete visibility into every call.

MCP client
Glama
MCP server

Full call logging

Every tool call is logged with complete inputs and outputs, so you can debug issues and audit what your agents are doing.

Tool access control

Enable or disable individual tools per connector, so you decide what your agents can and cannot do.

Managed credentials

Glama handles OAuth flows, token storage, and automatic rotation, so credentials never expire on your clients.

Usage analytics

See which tools your agents call, how often, and when, so you can understand usage patterns and catch anomalies.

100% free. Your data is private.
Tool DescriptionsA

Average 3.9/5 across 4 of 4 tools scored.

Server CoherenceB
Disambiguation5/5

Each tool has a clear, distinct purpose: auto_schema_learn for schema inference, batch_extract for multi-source extraction, extract_from_api for API responses, and extract_from_url for URLs. No two tools overlap in their primary function.

Naming Consistency2/5

Tool names use snake_case but with inconsistent patterns: 'auto_schema_learn' is adjective_noun_verb, while 'batch_extract' is noun_verb, and 'extract_from_api'/'extract_from_url' are verb_preposition_noun. This mismatch could confuse an agent.

Tool Count3/5

Four tools is on the low side for a data extraction server, covering basic sources but lacking something like direct file extraction. The count is reasonable for a minimal server but feels slightly incomplete.

Completeness3/5

Covers schema learning and extraction from URLs, APIs, and batch sources, but missing extraction from local files, databases, or direct text input. No schema management or listing tools exist, leaving notable gaps.

Available Tools

4 tools
auto_schema_learnInspect

Idempotent · 30s timeout · Automatically infer JSON Schema from sample data without extraction. Pass idempotency_key to deduplicate within 5 minutes.

ParametersJSON Schema
NameRequiredDescriptionDefault
sample_dataYesRepresentative sample data as JSON string (array of objects or single object). Schema is inferred from structure; use first 1-10 rows for array samples. Max 200KB.
idempotency_keyNoOptional cache key (UUID/string) for 5-minute deduplication. Repeat calls with same key return cached inferred schema instantly.
batch_extractInspect

Idempotent · 30s timeout · Extract data from multiple sources (JSON/JSONL/text) with a single consistent schema. Pass idempotency_key to deduplicate within 5 minutes.

ParametersJSON Schema
NameRequiredDescriptionDefault
schemaNoOptional target JSON Schema (draft-07) applied to all sources. If omitted, inferred from first source and reused across remaining sources. Enables consistent field extraction from diverse formats.
sourcesYesArray of 1-100 data sources to extract from. Each source includes type (format) and content (raw data).
idempotency_keyNoOptional deduplication key (UUID/string) for 5-minute cache. Identical batch calls (same sources + schema + key) return cached results instantly.
extract_from_apiInspect

Idempotent · 30s timeout · Extract structured data from API response JSON with schema adaptation. Pass idempotency_key to deduplicate within 5 minutes.

ParametersJSON Schema
NameRequiredDescriptionDefault
schemaNoOptional target JSON Schema (draft-07) for field extraction. If omitted, inferred from content structure. Enforces consistent field extraction across multiple API responses.
contentYesAPI response body as raw JSON string (max 200KB). Can be single object, array of objects, or array of primitives. Automatically parsed and validated.
idempotency_keyNoOptional deduplication key (UUID or unique string) for 5-minute cache. Identical calls return cached result instantly.
extract_from_urlInspect

Idempotent · 30s timeout · Extract structured data from URL content with auto schema learning. Pass idempotency_key to deduplicate identical calls within 5 minutes.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesHTTP(S) URL to fetch (will auto-download and parse), or raw content string (up to 200KB). Max 200KB after fetch.
schemaNoOptional pre-defined JSON Schema (draft-07). If omitted, schema is auto-inferred from content. Provide to enforce strict field extraction and type coercion.
idempotency_keyNoOptional UUID or unique identifier for 5-minute deduplication cache. Same key + tool = cached result in <5ms, zero re-fetching.

Discussions

No comments yet. Be the first to start the discussion!

Try in Browser

Your Connectors

Sign in to create a connector for this server.