MCP Insights Proxy
Executes OpenSearch DSL queries and returns synthesized output (table, list, summary, compact JSON) instead of raw JSON, supporting all query types including aggregations, nested queries, and has_parent queries.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Insights ProxyShow my top 10 most engaged posts from last week"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Insights Proxy v2
Context-efficient MCP server that executes OpenSearch queries and returns synthesized output instead of raw JSON.
Philosophy
Full query flexibility, synthesized output only.
You retain 100% of OpenSearch DSL capabilities. The proxy only transforms the output.
┌──────────────────────────────────────────────────────────────────────────┐
│ BEFORE (direct OpenSearch MCP) │
│ Query → OpenSearch → 500+ lines raw JSON → Context window 💥 │
├──────────────────────────────────────────────────────────────────────────┤
│ AFTER (this proxy) │
│ Query → Proxy → OpenSearch → Proxy formats → 20-50 lines → Context ✅ │
└──────────────────────────────────────────────────────────────────────────┘Related MCP server: MCP of MCPs
Key Difference from v1
v1 (Limited) | v2 (Full Flexibility) |
5-6 predefined tools | 1 main tool accepting ANY DSL |
Hardcoded query patterns | You build the query |
Limited aggregations | ALL aggregations supported |
No nested/has_parent | Full query DSL support |
Quick Start
1. Install
cd mcp-insights-proxy
npm install
npm run build2. Configure MCP Client
Claude Desktop (~/Library/Application Support/Claude/claude_desktop_config.json):
{
"mcpServers": {
"insights-proxy": {
"command": "node",
"args": ["/absolute/path/to/mcp-insights-proxy/dist/index.js"],
"env": {
"OPENSEARCH_URL": "https://your-cluster:9200",
"OPENSEARCH_INDEX": "channel_posts",
"OPENSEARCH_USER": "username",
"OPENSEARCH_PASS": "password"
}
}
}
}3. Restart Claude
Tools
opensearch_query (Main Tool)
Execute any OpenSearch DSL query. Full flexibility.
Parameters:
Parameter | Type | Description |
| object | Required. Full OpenSearch DSL query object |
| string | Index name (default: |
| enum |
|
| number | Max hits to show (default: 20, max: 50) |
Example - Complex aggregation with nested sub-aggs:
{
"query": {
"size": 0,
"query": {
"bool": {
"must": [
{"term": {"join_field": "post"}},
{"term": {"channel.type": "ig"}}
],
"filter": [
{"range": {"published_at": {"gte": "now-30d"}}}
]
}
},
"aggs": {
"by_hashtag": {
"terms": {"field": "hashtags", "size": 10},
"aggs": {
"avg_engagement": {"avg": {"field": "engagement"}},
"top_creators": {"terms": {"field": "channel.name", "size": 3}}
}
}
}
}
}Example - has_parent query:
{
"query": {
"query": {
"bool": {
"must": [
{"term": {"join_field": "post"}},
{
"has_parent": {
"parent_type": "channel",
"query": {
"bool": {
"must": [
{"term": {"channel.geo.country.code": "IT"}},
{"range": {"channel.followers": {"gte": 100000}}}
]
}
}
}
}
]
}
},
"sort": [{"engagement": "desc"}],
"size": 10,
"_source": ["channel.name", "engagement", "published_at"]
}
}opensearch_count
Quick count without full query overhead.
{
"query": {
"bool": {
"must": [
{"term": {"join_field": "post"}},
{"term": {"hashtags": "skincare"}}
]
}
}
}opensearch_mapping
Get field list for an index.
Output Formats
auto (default)
Automatically detects:
Aggregation-only → summary format
Few hits (≤5) → list format
Many hits → table format
table
| # | name | type | engagement | likes | published_at |
|---|------|------|------------|-------|--------------|
| 1 | creator1 | IG | 156K | 142K | 2025-01-15 |
| 2 | creator2 | TT | 98K | 89K | 2025-01-18 |list
**1.** @creator1 (IG) • eng: 156K • likes: 142K • views: 2.3M • 2025-01-15
**2.** @creator2 (TT) • eng: 98K • likes: 89K • views: 1.8M • 2025-01-18summary
Just counts and aggregation results, no individual hits.
compact_json
Minimal JSON with only _source (no _id, _score, _index metadata).
Aggregation Output Examples
Terms aggregation
**by_platform:**
• ig: 45.2K
• tt: 32.1K
• yt: 12.8KStats aggregation
**engagement_stats:** count=89.1K avg=2.3K min=0 max=1.2M sum=205MNested sub-aggregations
**by_hashtag:**
• **skincare** (12.4K)
**avg_engagement:** 3.2K
**top_creators:**
• creator1: 342
• creator2: 287
• **beauty** (8.7K)
**avg_engagement:** 2.8K
...Context Savings
Query Type | Raw JSON | Proxy Output | Savings |
Top 10 posts | ~3000 tokens | ~300 tokens | 90% |
Aggregation (5 buckets) | ~1500 tokens | ~150 tokens | 90% |
Complex nested agg | ~5000 tokens | ~400 tokens | 92% |
Environment Variables
Variable | Required | Default | Description |
| Yes |
| Cluster URL |
| No |
| Default index |
| No | - | Basic auth username |
| No | - | Basic auth password |
Development
npm run dev # Development with auto-reload
npm run build # Build for production
npm start # Run production buildArchitecture
┌─────────────┐ ┌────────────────────────────────────────────┐ ┌────────────┐
│ │ │ MCP Insights Proxy │ │ │
│ Claude │────▶│ 1. Receive DSL query (any complexity) │────▶│ OpenSearch │
│ (builds │ │ 2. Execute against cluster │ │ │
│ full │◀────│ 3. Parse response │◀────│ │
│ DSL) │ │ 4. Format: table/list/summary │ │ │
│ │ │ 5. Return compact markdown │ │ │
└─────────────┘ └────────────────────────────────────────────┘ └────────────┘
│ │
│ Returns:
│ "*45.2K hits • took 23ms*
│
│ ## Aggregations
│ **by_platform:**
│ • ig: 45.2K
│ • tt: 32.1K
│
│ ## Results
│ | # | name | engagement |..."
│
└─── ~300 tokens instead of ~3000License
MIT
Available Tools
3 toolsopensearch_countA
Quick document count with optional query filter. Lighter than full query.
| Name | Required | Description | Default |
|---|---|---|---|
| index | No | Index name | |
| query | No | Query filter (optional, counts all if omitted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of disclosure, but it only says 'Quick' and 'Lighter than full query'. It doesn't explain the return value, potential errors, performance ceilings, or how it handles missing indexes. The behavioral trait hint is minimal but present, so a 2 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that conveys the core action, optional filter, and comparative value. It is highly concise with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description should explain the return value and any caveats. It does not mention that the result is a count or what happens when no index is specified. The completeness is lacking, especially given that the schema has zero required parameters, creating ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description repeats the 'optional query filter' concept already in the schema, adding no new semantics. Thus it doesn't elevate beyond baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs a 'document count' with an optional query filter, and it distinguishes itself from the sibling 'opensearch_query' by being 'lighter than full query'. This gives a specific verb-resource pair and scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says 'Lighter than full query', which clearly implies using this when only a count is needed and not the actual documents. It provides clear context for when to choose this over a full query, though it doesn't explicitly name the alternative tool or mention exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
opensearch_mappingA
Get index mapping to understand available fields. Returns compact field list.
| Name | Required | Description | Default |
|---|---|---|---|
| index | No | Index name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the tool returns a 'compact field list,' which is useful behavioral context. However, it does not mention read-only nature, potential errors, or other operational details. For a simple read operation, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, front-loaded with the primary purpose. It is concise with no filler or redundancy, earning a perfect score for efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema, the description provides the essential context: what it does and what it returns. It is complete enough for its simplicity, though it could have briefly differentiated from siblings to improve context further.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 100% coverage for the single parameter 'index' with description 'Index name.' The tool description adds no additional parameter semantics beyond what the schema already provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Get index mapping to understand available fields.' It uses a specific verb and resource, and the purpose is distinct from sibling tools like opensearch_count and opensearch_query, which focus on counts and queries rather than schema discovery.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when one needs to understand available fields in an index, but it does not explicitly state when to use this tool over alternatives or provide exclusions. There is no mention of context such as 'use this before querying' or 'for field discovery, not data retrieval,' so guidance is minimal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
opensearch_queryA
Execute ANY OpenSearch DSL query and return synthesized results.
You have FULL FLEXIBILITY to build any query:
Complex bool queries with must/should/filter/must_not
has_parent, has_child, nested queries
Any aggregations (terms, stats, date_histogram, scripted_metric, etc.)
Script sorting, script fields, script filters
Geo queries, range queries, wildcards, match, term, etc.
The proxy executes your query and returns COMPACT formatted output instead of raw JSON.
OUTPUT FORMATS:
auto: Detects best format (default)
table: Tabular format for multiple hits
list: One-line-per-hit format
summary: Just counts and aggregations
compact_json: Minimal JSON (only _source, no metadata)
IMPORTANT: Build the full DSL query object as you normally would.
| Name | Required | Description | Default |
|---|---|---|---|
| index | No | Index name (optional, defaults to channel_posts) | |
| query | Yes | Complete OpenSearch DSL query object (same as you'd send to _search) | |
| max_display | No | Max hits to display (default: 20, max: 50) | |
| output_format | No | Output format (default: auto) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosure. It does add value by explaining the proxy behavior ('returns COMPACT formatted output instead of raw JSON'), detailing output formats, and emphasizing full query flexibility. Yet it lacks information on potential side effects, permissions, error handling, or performance implications, which prevents a higher score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: it opens with a clear purpose, then uses bulleted lists for capabilities and output formats, making it easy to scan. Length is justified by the tool's flexibility, though some redundancy exists between 'ANY' and the enumerated query types.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and lack of output schema, the description covers essential aspects: what it does, how to build queries, and what output formats are available. It does not provide explicit examples of synthesized output or discuss handling of large result sets, but the guidance is sufficient for a competent agent to infer expected behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Since schema coverage is 100%, baseline is 3. The description enhances the output_format parameter by explaining each mode (auto, table, list, summary, compact_json) and enriches the query parameter by providing concrete examples of DSL constructs (bool, has_parent, aggregations, etc.), adding real meaning beyond the schema's brief descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Execute ANY OpenSearch DSL query and return synthesized results,' with an explicit list of supported query types and aggregations. This specific verb+resource+scope differentiates it from sibling tools like opensearch_count and opensearch_mapping, which are narrower in function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context through phrases like 'FULL FLEXIBILITY' and lists a wide range of query capabilities, effectively positioning this as the go-to tool for complex OpenSearch queries. However, it does not explicitly contrast with siblings or state when not to use it, so it falls short of full alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
opensearch_count - First observed
opensearch_mapping - First observed
opensearch_query
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: count for lightweight counts, query for arbitrary DSL queries, and mapping for schema inspection. No overlap or ambiguity between them.
All tool names follow the same pattern: 'opensearch_' prefix followed by a single descriptive verb/noun (count, query, mapping). This is perfectly consistent and predictable.
With 3 tools, the set is compact and focused on core OpenSearch operations (count, query, mapping). It feels slightly minimal but appropriate for a read-only insights proxy, and each tool earns its place.
The domain is OpenSearch insights/querying, and the set covers count, query, and schema discovery, which are the essential operations. Minor gaps like direct document fetch or index listing can be handled via queries or mapping, so no critical dead ends.
Maintenance
Related MCP Connectors
Hosted MCP server for LLM cost estimation, model comparison, and budget-aware routing.
MCP server for building and testing AI agents with multi-model experimentation and insights.
A paid remote MCP for OpenAI Codex context compressor, built to return verdicts, receipts, usage log
Related MCP Servers
- AlicenseBqualityDmaintenanceA Model Context Protocol server implementation that enables natural language interactions with OpenSearch clusters, allowing users to search documents, analyze indices, and manage clusters through simple conversational commands.1411Apache 2.0
- AlicenseNot gradedqualityDmaintenanceA meta-server that aggregates multiple MCP servers into a single interface, reducing token usage by 98%+ through progressive tool discovery and direct code execution that processes data between tools without consuming context window space.16 npm10Apache 2.0
- AlicenseAqualityAmaintenanceMCP server for OpenSearch that enables AI assistants to interact with OpenSearch clusters through a standardized interface for search, index management, and cluster operations.977,891 PyPI151Apache 2.0
- AlicenseNot gradedqualityCmaintenanceMCP server that cuts cloud LLM costs 36-42% by indexing context locally and giving agents precision retrieval tools instead of raw context dumps.7 npm7MIT