Skip to main content
Glama

Metricairn

Free, self-hosted product analytics for your dashboard and AI agent.

See what changed, where the change is concentrated, and what to check next. Every investigation exposes its query plan, comparison windows, collection coverage, and caveats. Save the evidence; record your decision separately.

Metricairn investigation with simulated data

Start

git clone https://github.com/rnepal2/metricairn.git
cd metricairn
docker compose up --build

Open http://localhost:8000, create a project, and save the three keys shown once. The container serves the dashboard, API, and tracker; SQLite data persists in a named volume. No account, cloud service, or AI key required. For source development, see CONTRIBUTING.

Install the tracker on your site:

<script defer src="https://YOUR-HOST/static/metricairn.js"
  data-api="https://YOUR-HOST/api/v1/ingest" data-key="alw_YOUR_TRACKING_KEY"></script>

Track a meaningful action with window.metricairn.event('signup'), then define it as a goal in the dashboard. HTTPS and a reachable deployment are required for a real site; the local container binds to loopback. See deployment.

Related MCP server: posthog-mcp

What you get

  • Traffic, realtime activity, custom events, ordered funnels, conversion goals, and weekly retention.

  • Filtered analytics queries and change investigations with additive segment counts, coverage checks, timeline context, and evidence export.

  • Immutable saved reports with separate review status and notes.

  • 20 read-only MCP tools, plus three optional management tools. Deterministic workflows run without an LLM; natural-language SQL is optional.

  • A small browser tracker with SPA navigation, bounded retries, deduplication, query cleanup, and DNT/GPC controls.

  • Scoped keys, rotation, project-data deletion, optional alerts and digests. Existing recorded-revenue support remains optional; payment and billing expansion is deferred.

Example: Maya at Billwise

Maya Chen, the fictional solo founder of an invoicing app, investigates traffic, campaign signups, conversion, and returning visitors. The demo shows copyable MCP calls beside live results from simulated events; a companion script verifies them through an actual MCP session. Walkthrough and demo setup.

Connect an agent

After make setup, configure your MCP client:

{
  "mcpServers": {
    "metricairn": {
      "command": "/ABSOLUTE/REPO/.venv/bin/python",
      "args": ["-m", "metricairn_mcp"],
      "env": {
        "METRICAIRN_API_URL": "http://localhost:8000",
        "METRICAIRN_READ_KEY": "alr_YOUR_READ_KEY"
      }
    }
  }
}

Try: “Investigate signup changes over the last week. Check collection coverage and show the largest changed segments before suggesting a cause.” MCP reference.

Scope

Metricairn is an early open-source release for small teams operating their own analytics. Its focus is an inspectable investigation workflow—not feature parity with PostHog. There are no hosted user accounts, OAuth grants, session replay, feature flags, experiments, or high-volume rollups. Cookieless tracking still uses browser-local identifiers. Metric definitions · Security.

API · Architecture · Roadmap · Product direction

MIT licensed. All implemented features are free; there is no paid tier or license server. Self-hosting and optional external providers have their own costs. Contributions are welcome through issues and pull requests.

Available Tools

20 tools
askC
Read-only

Ask a natural-language question about your analytics, e.g. 'why did revenue dip last Tuesday?'

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
questionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds no behavioral context beyond that — nothing about cost, latency, scope, or how the answer is produced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with an inline example, front-loaded and free of waste. Nothing is extraneous and it is appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations cover safety. However, the description leaves the 'days' parameter and the relationship to sibling query/investigation tools unaddressed, which matters given the crowded toolset.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate but does not. The example question loosely maps to the 'question' parameter, but the 'days' parameter (default 30) is never explained, leaving a key scoping input undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (Ask) and resource (natural-language question about analytics), plus a concrete example. It does not distinguish itself from siblings like investigate_change or query_metrics, so it falls short of the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus structured alternatives such as query_metrics or run_query, and no when-not conditions. The example implies a diagnostic use case but does not route the agent explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

breakdownB
Read-only

Top values for a dimension: path, referrer, utm_source, utm_medium, utm_campaign, device, browser, os, country, event.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
limitNo
dimensionNopath

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only, non-destructive safety profile. The description adds no further behavioral context such as aggregation method, time-range behavior, sorting, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the dimension list is compact and directly useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema and annotations cover return format and safety, and the dimension list helps invocation. However, the definition lacks usage guidance and does not clarify time-range or limit semantics, leaving avoidable ambiguity for a 3-parameter analytics tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema coverage, the description compensates for the key dimension parameter by enumerating valid values (path, referrer, etc.). It leaves days and limit undocumented, though their titles and defaults make them somewhat inferable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Describes the operation as returning top values for a dimension and lists the valid dimensions, making the core resource clear. It does not distinguish itself from sibling list_dimension_values, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus list_dimension_values, query_metrics, or other analytics siblings. The description only states output content, leaving selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compareA
Read-only

Compare an equal-length previous period: traffic, custom events and revenue. Changes are descriptive. Null percent change means the previous baseline was zero.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnlyHint, destructiveHint=false, closed world), and the description adds genuinely useful behavioral detail beyond them: the comparison is descriptive, and a null percent change signals a zero baseline. It does not state permissions or limits, but for a read-only comparison this is largely covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and followed by the edge-case semantics. Nothing is wasted or repeated from structured fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be documented, and the null-percent-change note covers the main interpretive edge case. The only real gap is the undocumented 'days' parameter, which an agent would have to infer from the period framing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single 'days' parameter (default 30) is never mentioned; only the abstract notion of an 'equal-length previous period' hints at a duration window. The description adds conceptual framing but does not document the parameter's meaning, units, or default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (compare) and the resource/scope: traffic, custom events, and revenue over an equal-length previous period. It clearly states what is produced without restating the tool name, though it does not explicitly differentiate itself from siblings like breakdown or revenue_attribution.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Compare an equal-length previous period" implies the context of use (comparing current vs prior period), but there is no explicit when-to-use/when-not guidance and no sibling alternatives are named, leaving routing up to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

detect_anomaliesB
Read-only

Unusual spikes/dips (screening signals, not significance tests) in pageviews and revenue (robust historical scores).

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds interpretive context ('screening signals, not significance tests', 'robust historical scores'), but it does not disclose time-window behavior, sensitivity, or output characteristics beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence with no wasted words, and it front-loads the core concept: unusual spikes/dips in pageviews and revenue. It is appropriately sized, though the parenthetical packing slightly reduces readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering safety, the description need not explain return values. Still, for an anomaly detection tool among many analytics siblings, it leaves out the days parameter and does not explain when this is preferable to related investigation tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter, days, with a default of 30, but schema description coverage is 0% and the description never mentions the time window. For a tool whose only parameter controls the historical lookback, this is a meaningful omission.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific phenomenon (unusual spikes/dips) in specific metrics (pageviews and revenue), and clarifies these are screening signals rather than significance tests. It is clear enough, but it does not explicitly name an alternative sibling tool or fully differentiate from investigate_change.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical 'screening signals, not significance tests' provides important usage context and implies when this tool is appropriate. However, it does not state when to use this versus siblings like investigate_change, query_metrics, or compare.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

funnel_reportA
Read-only

Conversion report for a funnel (match by name or id). Steps must be completed in order. Pass segment_by (e.g. 'device', 'utm_source') to compare conversion per segment — a visitor's segment is the dimension value on their entry-step event.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
funnelNo
segment_byNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false, destructiveHint=false, so the safety profile is covered. The description usefully adds domain behavior — ordered steps and how segment membership is assigned. It omits what happens on an unmatched funnel name and how the default lookback window behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core purpose and the ordering constraint before the optional segmentation detail. No filler, though the segment explanation is slightly dense for its position.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and annotations cover the safety profile. The description supplies the domain semantics an agent needs; only the unexplained days parameter leaves a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must carry the load. It does well for two of three params, giving segment_by examples and a precise definition of segment membership, but the days parameter (default 30) is never explained in either the schema or the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: a conversion report for a funnel, with the lookup key (name or id) named. It is clearly distinguishable from sibling reports like goal_report and retention_report by scope, though it never explicitly contrasts itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implied usage is present (analyse funnel conversion, optionally segmented), and it warns that steps must be completed in order. However it gives no when-to-use/when-not guidance relative to the many sibling reporting tools, leaving the agent to infer selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_investigationB
Read-only

Read an immutable saved investigation and its separate human review note.

ParametersJSON Schema
NameRequiredDescriptionDefault
investigation_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds that the investigation is immutable and that the human review note is separate, which is useful behavioral context, but it does not cover authorization, rate limits, or retrieval failures.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded sentence with no wasted words. Every clause earns its place by specifying the read action, immutability, and the separate review note.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and has an output schema, so return-value details are unnecessary. However, given the 0% schema description coverage and no usage guidance, the definition is only minimally complete for an agent deciding when and how to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the undocumented investigation_id parameter. It does not explain the ID format, source, or whether the ID must come from list_investigations, leaving only the self-explanatory parameter name to guide the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb, resource, and scope: reading an immutable saved investigation plus its separate human review note. It clearly distinguishes this from creating or listing investigations, but it does not explicitly name a sibling alternative or state retrieval by ID.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus list_investigations, investigate_change, or other related tools. The intended workflow, such as retrieving an investigation after listing or investigating, is left entirely implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_realtimeB
Read-only

Live activity: visitors, pageviews and top pages in the last 30 minutes.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered by structured data. The description adds the 30-minute recency window, which is meaningful behavioral context, but says nothing about refresh cadence, caching, or sample-size caveats that often matter for real-time endpoints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single fragment sentence that front-loads the key scope ('Live activity') and lists the payload in order. Nothing is wasted, though as a fragment it reads more like a label than a full instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the return shape needn't be described, and annotations cover the safety profile; the description supplies the remaining essentials: purpose and time window. The only real gap is sibling differentiation, which is a usage concern rather than a completeness one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing to disambiguate and the baseline for a parameterless tool applies. The description correctly implies no filtering options are available, matching the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (live activity) and the concrete measures returned (visitors, pageviews, top pages), so an agent immediately understands this reports current real-time traffic. It is clearly distinct from analytics siblings like query_metrics or goal_report, though it doesn't explicitly contrast itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to choose this tool over list_metrics, query_metrics, or breakdown, which all touch similar traffic data. The 30-minute window is stated as a fact rather than as a usage condition, so the agent must infer applicability on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

goal_reportB
Read-only

Period conversion: unique visitors with a pageview then this goal. Reports denominator, repeated occurrences and missing identity. Use date_from/date_to together for an exact window; otherwise days determines it.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
date_toNo
goal_idYes
date_fromNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, non-destructive, and closed-world, so the safety profile is covered. The description adds only a hint of report contents (denominator, repeated occurrences, missing identity), and with an output schema present that content is largely redundant, so it contributes little beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, followed by the date-window rule, in two compact sentences with no filler. The phrasing is clipped ('repeated occurrences and missing identity') but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Annotations and an output schema are available, so return values need no explanation. The gap is the required goal_id, which is unexplained anywhere, and the absence of any sibling differentiation for a report tool among many report siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the full burden and it does explain the date_from/date_to vs days precedence clearly. However, goal_id is required yet never described, leaving the single mandatory parameter undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (goal conversion) and the measurement basis (unique visitors with a pageview then this goal), which is enough for an agent to distinguish it from generic query tools. It falls short of 5 because it never names or contrasts with the close sibling funnel_report, so the boundary between the two is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The date guidance ('use date_from/date_to together... otherwise days determines it') is real usage direction, but it is purely about parameter mechanics. There is no when-to-use-this-vs-alternatives guidance despite several report siblings (funnel_report, retention_report, breakdown), leaving tool selection implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

integration_healthA
Read-only

Is the instrumentation actually flowing? Checklist: tracker pageviews, revenue events, custom events, funnels, recency. Run this first when answers look empty or suspicious — most 'wrong' answers are missing data, not wrong analysis.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, non-destructive, and closed-world. The description adds a checklist of what it inspects and a rationale ('most wrong answers are missing data'), providing useful behavioral context beyond the safety annotations. It does not describe output format, but an output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with a diagnostic question, then a concise checklist, then usage guidance. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters, an output schema, and safety annotations, the description supplies the necessary context: what it checks and when to use it. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so there is nothing to document. Baseline for zero params is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description conveys that this is a diagnostic tool checking whether instrumentation (pageviews, revenue events, custom events, funnels, recency) is flowing. It distinguishes itself as a pre-check by advising to 'run this first when answers look empty or suspicious,' though it does not name specific sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: when answers look empty or suspicious. It does not list when not to use or name alternative tools, but the context strongly implies it is a diagnostic step before other analysis.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

investigate_changeA
Read-only

Investigate changes in pageviews, events, or event_count (requires event_name). Returns equal-duration comparison, additive source/device/browser/path changes, coverage, timeline notes, caveats and next checks. This describes associations, not causes. Use date_from/date_to together to fix an exact window; otherwise days determines it.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
metricNopageviews
date_toNo
filtersNo
date_fromNo
event_nameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds meaningful context beyond annotations: the equal-duration comparison model, the additive breakdown dimensions, and the explicit "associations, not causes" caveat that frames how results should be interpreted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with purpose and return scope, then the windowing rule. Efficient with little waste, though the return-content sentence partially overlaps the existing output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only investigative tool with an output schema already covering return values, the description supplies the key framing (what it compares, causal limitation, window logic). Missing only explicit sibling routing, which is not strictly required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so parameters are undocumented structurally, and the description only compensates partially — it clarifies event_name requirement for event_count and the date_from/date_to vs days interaction. Metric values, filter shapes, and date formats remain unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Investigate") and resource ("changes in pageviews, events, or event_count"), with a conditional qualifier for event_count. This distinguishes it from siblings like detect_anomalies and compare without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the date_from/date_to vs days window logic, which is real usage guidance, but never states when to choose this tool over detect_anomalies, compare, or query_metrics, nor any exclusions. Usage is implied rather than routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_dimension_valuesB
Read-only

Discover which values a dimension actually has (e.g. real page paths, real UTM sources).

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
dimensionNopath

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, covering the safety profile. The description adds that it discovers values that 'actually' exist in the data, implying a data-derived enumeration rather than a static list, which is useful context. It does not disclose sampling, time-window behavior, or result shape beyond this.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. The core purpose and illustrative examples are delivered immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and annotations cover the safety hints. However, for a tool with two undocumented parameters and many sibling analytics tools, the description omits how `days` and `dimension` work and provides no guidance for selecting this tool over alternatives. It is too sparse to be complete for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters (`days` and `dimension`), so the description carries the full burden of explaining them. It mentions the concept of a dimension and gives examples but never explains the `dimension` parameter's accepted values or the `days` parameter's role in limiting the time window. This is a significant gap for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Discover which values a dimension actually has,' with examples of real page paths and UTM sources. It is clear enough to distinguish from metric-oriented siblings like list_metrics or query_metrics, though it does not explicitly name an alternative tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the phrase 'actually has (e.g. real page paths, real UTM sources),' suggesting it is for enumerating existing dimension values rather than querying metrics. However, there is no explicit guidance on when to use this versus breakdown, query_metrics, or other sibling tools, and no exclusions or prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_goalsA
Read-only

Discover named custom-event conversion goals and their IDs.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is fully covered. The description adds only that goals are 'named' and that IDs are returned, which is minimal extra context for a zero-parameter, side-effect-free listing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the resource and the returned identifier are stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-value detail is unnecessary, and a no-argument read-only list needs little explanation. The only real gap is the absence of any routing guidance relative to goal_report.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema carries no semantics the description needs to compensate for. Baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Discover') and a precise resource ('named custom-event conversion goals'), plus the useful detail that IDs are returned. It does not explicitly distinguish itself from the sibling goal_report, which presumably consumes those IDs, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives such as goal_report or list_metrics. The natural use case (fetching goal IDs to feed a report) must be inferred rather than read.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_investigationsC
Read-only

Recent saved evidence records, their IDs and human review status.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and destructiveHint, so the safety profile is covered. The description only restates return content, which is also captured by the output schema, and adds no behavioral context such as pagination, sorting, or recency limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise fragment, but it is structured as a noun phrase without a clear action verb. It is appropriately sized but not front-loaded with a purpose statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple schema and existing annotations/output schema, the description should at minimum clarify that it lists investigations and how it relates to get_investigation. It fails to do so, leaving the core purpose ambiguous for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description does not need to explain parameter semantics. Baseline score of 4 is appropriate for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the returned content ('recent saved evidence records, their IDs and human review status') but never states a verb or mentions 'investigations'. This creates ambiguity with the sibling get_investigation and fails to clearly differentiate this tool as a listing operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like get_investigation or other list_* tools. There is no context, prerequisite, or exclusion mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_metricsA
Read-only

Catalog of available metrics and dimensions. Start here to discover what you can query.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false, and destructiveHint=false, so the safety profile is fully covered externally. The description adds the useful framing of what the catalog contains (metrics and dimensions) but discloses nothing beyond that, such as pagination or result volume behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero waste, and the front-loaded statement of contents is followed immediately by the actionable 'start here' directive. Nothing could be removed without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain return values, and with no parameters there is no input surface to document. It is nearly complete for this simple discovery tool, missing only a hint about how the catalog is used downstream.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate, and it correctly does not pad with irrelevant parameter detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource (available metrics and dimensions) as a catalog, which implicitly distinguishes it from the query/reporting siblings like query_metrics. However, it never states the verb explicitly (list/retrieve) and does not name a sibling it is not, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Start here to discover what you can query' gives clear context: this is the discovery entry point before querying. It provides positive usage context but no explicit exclusions or named alternatives, which keeps it at a 4 rather than a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_notesA
Read-only

Timeline annotations: launches, deploys, campaigns the founder (or an agent) logged. Newest first.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds one genuinely new behavioral fact — results come back newest-first — which is useful for pagination-style consumption, but nothing about volume, limits, or empty results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short fragments, zero filler, with the resource identity and sort order front-loaded. Nothing here could be removed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and an output schema present, the description need not explain return fields; it covers resource, content examples, and sort order. The only real shortfall is the absence of any when-to-use signal relative to the many sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. The description introduces no misleading parameter notions, and there is nothing for the schema to document.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (timeline annotations) and enumerates concrete examples (launches, deploys, campaigns), plus the ordering (newest first). It is clear what the tool returns, though it never explicitly says 'list' as a verb or differentiates itself from siblings like list_investigations or list_metrics beyond the resource name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus the many sibling list/query tools, no prerequisites, and no exclusions. The parenthetical '(or an agent) logged' hints at provenance but does not tell the agent when to reach for this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mcp_usageB
Read-only

How AI agents are using this MCP server: tool-call counts, error rates, and recent questions.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, non-destructive, and closed-world behavior, so the safety profile is covered. The description adds what data is reported (counts, error rates, recent questions), which is useful context beyond annotations, but it does not disclose aggregation level, permissions required, or how 'recent questions' are defined.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, well-structured sentence that front-loads the purpose and lists the key metrics. No filler or redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description need not explain return values, and the low complexity (one optional parameter, no nesting) keeps the description from being inadequate. However, the missing semantics for 'days' and the absence of any usage context leave clear gaps for an agent deciding whether and how to call the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has a single parameter 'days' with 0% description coverage, and the tool description does not mention it at all. The description therefore fails to compensate for the missing parameter documentation, leaving the agent without any indication of what the time window controls or how its default of 30 is used.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the resource (MCP server usage) and the specific metrics reported (tool-call counts, error rates, recent questions). It is readily distinguishable from sibling tools, which deal with product analytics rather than the MCP server's own usage. However, it lacks an explicit verb and does not name any sibling alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when the tool is useful—when an agent wants to see how it and others are using the server—but gives no explicit when-to-use or when-not-to-use guidance. No alternative tools are mentioned, nor are prerequisites or edge cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_metricsC
Read-only

Time series for a metric. metric: visitors|pageviews|sessions|events|revenue. interval: day|hour.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
metricNovisitors
intervalNoday

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds nothing behavioral beyond that - no note on scoping, rate limits, or what the series represents (raw vs aggregated).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely terse and front-loaded - the resource and the two key parameter vocabularies come first with no filler. It is telegraphic to the point of being cryptic, but there is no wasted sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. Still, a 3-param metric tool with 0% schema description coverage leaves the days parameter and the tool's relationship to siblings unexplained, which is a real gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description partially compensates by listing valid values for metric (visitors|pageviews|sessions|events|revenue) and interval (day|hour). However, the days parameter is never mentioned, and the accepted-value lists are not framed as enum constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Time series for a metric' identifies the resource and the shape of the result, and it enumerates the valid metric values. But there is no verb and no differentiation from close siblings like list_metrics, get_realtime, breakdown, or compare, so an agent cannot tell from the description alone which of these to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no exclusions, and no mention of the many sibling tools that also surface metric data (get_realtime, breakdown, compare, detect_anomalies). The agent must infer the context from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retention_reportA
Read-only

Weekly first-observed visitor cohorts, weeks 0..12. Incomplete cells are null; identities and acquisition history affect accuracy. event_name optionally restricts returning activity, not the first-observed cohort definition. Use date_from/date_to together for an exact window; otherwise days determines it.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
date_toNo
date_fromNo
event_nameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this a safe read-only operation, so the description is free to add domain behavior — and it does: incomplete cohort cells are null, weeks are bounded to 0..12, and identities/acquisition history affect accuracy. That caveat about accuracy limits is genuine context an agent could not infer from readOnlyHint/destructiveHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the output shape before the parameter caveats, with no filler. It is terse to the point that the date-window sentence requires a second read, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return structure needn't be re-explained, and the description still covers cohort definition, cell nulls, and parameter precedence. The only substantive omission is the accepted date string format for date_from/date_to, which matters because those parameters are untyped strings documented nowhere else.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the load, and it does for three of four parameters: event_name's scope ('restricts returning activity, not the first-observed cohort definition') and the days vs date_from/date_to precedence rule are both meaningful. It still never gives the expected date string format, which is the one remaining gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific resource and shape: 'Weekly first-observed visitor cohorts, weeks 0..12', which tells an agent exactly what kind of report this produces. It is distinct from siblings like funnel_report or breakdown by virtue of the cohort/retention framing, though it never names an alternative to differentiate explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives real invocation guidance for parameters (use date_from/date_to together, otherwise days determines the window; event_name restricts returning activity), but never says when to reach for retention_report versus funnel_report, compare, or breakdown. Usage is implied by the report type rather than stated as a selection rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revenue_attributionC
Read-only

Revenue total, transactions, revenue-per-visitor, and breakdown by traffic source.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, destructiveHint=false, and closed-world, so the safety profile is covered. The description adds no behavioral context beyond that: no mention of the time window applied, whether results are aggregated or per-source rows, or any limits on the breakdown. It only hints at shape by naming the metrics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, which is appropriately sized. However, it is a trailing fragment without a leading verb or purpose statement, so it is not front-loaded with what an agent most needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out — yet the description spends its whole budget naming outputs anyway. It omits purpose, usage context, and the semantics of its single input, leaving the definition thin for even a one-parameter analytics tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter (`days`) with 0% schema description coverage, and the description never mentions a time window or its default. With only one parameter the burden for documenting it falls to the description, and it fails to do so.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a noun list of outputs (revenue total, transactions, revenue-per-visitor, traffic-source breakdown) rather than a verb+resource statement of what the tool does. It is inferable that this returns revenue analytics, but it does not differentiate itself from siblings like breakdown or query_metrics, which also produce breakdowns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus breakdown, query_metrics, retention_report, or goal_report. The only implicit signal is that it is revenue-themed. An agent must guess which sibling to call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_queryA
Read-only

Execute a typed analytics plan without AI or SQL. Fields: metric (pageviews|visitors|sessions|events|event_count), mode (total|timeseries|breakdown), event_name (required for event_count), dimension (for breakdown), interval (day|hour), date_from/date_to (ISO UTC), limit (1..100), filters ([{field, values}]). Filter fields: path, utm_source, utm_medium, utm_campaign, device, browser, os, country, event. Filters are ANDed; values within each filter are ORed. Returns resolved plan and caveats. Example: {"metric":"event_count","event_name":"signup","mode":"breakdown","dimension":"device"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
planYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safe read-only profile, yet the description adds real behavioral context: the AND-across-filters / OR-within-values semantics and the fact that it returns a resolved plan plus caveats. It omits auth/rate-limit context, but the added filter semantics are valuable beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the one-line purpose, then proceeds through fields, filter semantics, and an example in dense, ordered form. It runs long but nearly every clause carries invocation-relevant information, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers conditional field requirements, formats, filter combination logic, and return contents, and with an output schema present it needn't detail the response. A concrete example seals invocation understanding; only explicit sibling routing is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the burden, and it does: conditional requirements (event_name required for event_count, dimension for breakdown), date format (ISO UTC), limit range, full filter field list, and a worked example. The enum value lists partly restate the schema, but the conditional logic and format details are net-new meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('Execute a typed analytics plan') and differentiates by channel ('without AI or SQL'), which distinguishes it from an ask-style sibling. It does not, however, name the specific siblings it competes with (query_metrics, breakdown, compare).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without AI or SQL' implies this is the deterministic path versus AI-driven tools, giving implicit usage guidance. But there is no explicit when-to-use/when-not framing and no named alternatives among the many analytical siblings (breakdown, funnel_report, compare).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 20 tool updatesv0.1.0
    • First observedask
    • First observedbreakdown
    • First observedcompare
    • First observeddetect_anomalies
    • First observedfunnel_report
    • First observedget_investigation
    • First observedget_realtime
    • First observedgoal_report
    • First observedintegration_health
    • First observedinvestigate_change
    • First observedlist_dimension_values
    • First observedlist_goals
    • First observedlist_investigations
    • First observedlist_metrics
    • First observedlist_notes
    • First observedmcp_usage
    • First observedquery_metrics
    • First observedretention_report
    • First observedrevenue_attribution
    • First observedrun_query

TDQS

B3.1/5.0

Scored across 20 tools

Disambiguation3/5

Several tools overlap heavily: run_query subsumes query_metrics and breakdown (total/timeseries/breakdown modes), while compare and investigate_change both perform equal-duration period comparisons, and ask overlaps with the structured query tools. Funnel/goal/retention reports and the note/investigation lifecycle are clearly distinct, but the querying layer has fuzzy boundaries that invite misselection.

Naming Consistency3/5

There is a recognizable verb_noun pattern for many tools (list_metrics, query_metrics, run_query, list_goals, get_realtime, detect_anomalies, investigate_change, get_investigation), but it is broken by bare nouns and ambiguous names (breakdown, compare, ask, funnel_report, revenue_attribution, mcp_usage). The mix is readable but not a predictable convention.

Tool Count3/5

At 20 tools this is on the heavy side for an analytics server, and the count is inflated by overlapping query paths that could be consolidated (e.g. query_metrics/breakdown folded into run_query). It is not egregious—each tool maps to a plausible analytics task—but it sits at the borderline of too many.

Completeness4/5

The surface covers the analytics lifecycle well: discovery (list_metrics, list_dimension_values), querying, funnel/goal/retention/attribution reporting, anomaly detection, realtime, data-health checks, and investigation save/review. Minor gaps exist (e.g. no explicit breakdown-by-metric filter builder output, no scheduling), but core workflows are covered without dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Enables AI assistants to interact with Metabase for database operations, SQL queries, dashboard management, and analytics automation.
    28
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    Enables AI agents to query PostHog analytics data directly via tool calls, including insights, events, feature flags, trends, and persons.
    5
    498 npm
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    First-party web analytics MCP server for AI agents, providing 42 tools to query traffic, events, funnels, conversions, sources, and performance data.
    40
    78 npm
    MIT