Skip to main content
Glama
paulsham

Wiki Analytics Specification MCP Server

by paulsham

Wiki Analytics Specification MCP Server

A system that maintains analytics event specifications in a Wiki, then transforms them into formats that AI coding tools to query efficiently via Model Context Protocol (MCP).

Overview

Analytics specs often live in scattered documentation that's hard for developers to consume, or in technical formats that PMs and data scientists can't easily maintain. This project bridges that gap: non-technical stakeholders author specs in familiar Wiki markdown tables, while developers get structured, queryable data through AI coding tools.

This project enables a Wiki-based workflow for managing analytics specifications:

  1. Author in Wiki - Define events, properties, and property groups using markdown tables

  2. Build automatically - Convert Wiki markdown → CSV → JavaScript modules

  3. Query with Claude - MCP server provides tools for Claude to search and validate specs

Note: This project uses GitHub/GitLab wiki conventions, where wikis are stored as markdown files in a separate git repository (e.g., repo.wiki.git). This allows the wiki content to be cloned and processed programmatically.

Related MCP server: Amplitude MCP Server

Features

  • Wiki-based authoring - Human-friendly markdown tables with version control

  • Property reuse - Define properties once, reference everywhere via property groups

  • Compact responses - MCP tools return structured JSON, reducing token usage by ~66%

  • Validation support - Validate tracking implementations against specs

  • Local execution - Runs locally with Claude Desktop, no cloud hosting required

Installation

Note:

  • This project is currently optimized for GitHub (template repositories, GitHub Actions, GitHub wikis). The concepts translate to other platforms like GitLab, but implementation details differ.

  • In Github, adding a wiki to a private repo requires a paid plan.

This project is designed as a template for your own analytics specifications.

On GitHub:

  1. Click "Use this template" → "Create a new repository" on GitHub

  2. Clone your new repository locally and install dependencies:

    git clone https://github.com/yourusername/your-repo-name.git
    cd your-repo-name
    npm install  # Automatically sets up git hooks
  3. Set up your wiki with example content:

    • Go to your repository's Wiki tab on GitHub

    • Create pages: Events.md, Property-Groups.md, Properties.md

    • Copy content from this project's wiki-examples/ directory

  4. Trigger the build workflow:

    • Go to Actions tab → "Transform Wiki to Specs" → "Run workflow"

  5. Pull the generated specs:

    git pull
  6. Configure with your AI tool (see "Configure with AI Coding Tools" below)

Optional enhancements:

  • Enable automated sync by uncommenting the cron job in .github/workflows/transform-wiki.yml

  • Add GitHub branch protection rules for additional server-side protection (Husky hooks already block local commits)

Requirements:

  • Node.js 20+

  • Git

  • GitHub account (for template and CI/CD workflow)

Quick Test (Using Example Data)

To test the MCP server without setting up a wiki:

# Clone the repository
git clone https://github.com/username/wiki-mcp-analytics.git
cd wiki-mcp-analytics

# Install dependencies
npm install

# Build from example data
npm run build:example

# Start the MCP server
npm start

This uses the wiki-examples/ directory to generate test specs.

Usage

Build Specs from Wiki

# Run full build pipeline (Wiki markdown → CSV → JavaScript)
npm run build

# Run individual steps
npm run build:csv      # Wiki markdown → CSV only
npm run build:js       # CSV → JavaScript only
npm run build:example  # Use wiki-examples/ for testing

Note: Developers typically don't need to run build commands. The CI/CD workflow automatically generates and commits specs when the wiki changes. Just git pull to get the latest.

Start the MCP Server

npm start

Configure with AI Coding Tools

Claude Code

claude mcp add wiki-analytics node /path/to/wiki-mcp-analytics/src/mcp-server/index.js

Other MCP-compatible tools

Add to your MCP configuration file:

{
  "mcpServers": {
    "wiki-analytics": {
      "command": "node",
      "args": ["/path/to/wiki-mcp-analytics/src/mcp-server/index.js"]
    }
  }
}

Wiki Format

The Wiki uses three markdown pages with tables where each row represents one item.

Events.md

Event Name

Event Table

Event Description

Property Groups

Additional Properties

Notes

user_registered

Registration

User completed registration

user_contextdevice_info

registration_methodreferral_code

Fire after successful registration

Property-Groups.md

Group Name

Description

Properties

user_context

Common user identification properties

user_idemailaccount_created_at

Properties.md

Property Name

Type

Constraints

Description

Usage

user_id

string

regex: ^[0-9a-f-]{36}$

Unique user identifier

Include in all authenticated events

Key conventions:

  • Use <br> for line breaks in multi-value cells

  • All properties must be defined in Properties.md

  • Events and property groups reference properties by name only

Project Structure

wiki-mcp-analytics/
├── src/
│   ├── builder/             # Build pipeline (Wiki → CSV → JS)
│   │   ├── index.js         # Pipeline orchestration
│   │   ├── wiki-to-csv.js   # Parse markdown → CSV
│   │   └── csv-to-javascript.js  # Generate JS modules
│   └── mcp-server/          # MCP server implementation
│       └── index.js
├── specs/                   # Generated specs (committed by CI/CD)
│   ├── csv/                 # CSV format for tools
│   │   ├── .gitkeep
│   │   └── *.csv (generated)
│   └── javascript/          # JS modules for runtime
│       ├── .gitkeep
│       └── */ (generated)
├── .husky/                  # Git hooks (pre-commit protection)
│   └── pre-commit
├── wiki-examples/           # Example wiki content for testing
│   ├── Events.md
│   ├── Property-Groups.md
│   └── Properties.md
└── package.json

MCP Tools

The server provides developer-focused tools for implementation and validation:

get_event_implementation

Get complete event specification with all properties expanded.

// Returns structured JSON with property groups, constraints, and notes
get_event_implementation("user_registered")

validate_event_payload

Validate a tracking implementation against the spec.

// Returns errors, warnings, and valid fields
validate_event_payload("user_registered", { user_id: "123", ... })

search_events

Find events by criteria.

// Search by name, table, or property usage
search_events({ query: "registration", has_property: "user_id" })

get_property_details

Get property definition and usage across events.

// Returns type, constraints, description, and where it's used
get_property_details("user_id")

Find events in the same flow/table.

// Returns related events for funnel analysis
get_related_events("user_registered")

Architecture

Wiki Repo (separate git repository)
    ↓ (sync via CI/CD)
Main Repo: wiki-mcp-analytics
    ↓ (build pipeline)
specs/csv/ + specs/javascript/
    ↓ (read by)
MCP Server (runs locally)
    ↓ (stdio)
Claude Desktop / Claude Code

Note: GitHub/GitLab wikis are separate repositories with a .wiki suffix. This project syncs from the wiki repo and builds the specs.

Development

Automated Workflow

When you update your wiki, the GitHub Action automatically:

  1. Detects wiki changes

  2. Builds fresh specs (CSV + JavaScript)

  3. Commits to your repo as github-actions[bot]

  4. Developers pull the updated specs

Optional: Enable daily sync by uncommenting the cron schedule in .github/workflows/transform-wiki.yml

Local Development

# Test the builder with example data (no wiki setup needed)
npm run build:example

# Build from your wiki (requires wiki/ directory cloned locally)
git clone https://github.com/yourname/wiki-mcp-analytics.wiki.git wiki
npm run build

# Run the MCP server
npm start

Protection Against Stale Commits

The project includes a pre-commit hook (via Husky) that blocks manual commits to specs/. This ensures only CI/CD commits generated specs.

To bypass (not recommended): git commit --no-verify

For additional protection, consider setting up branch protection rules to restrict specs/ changes.

License

MIT License - see LICENSE for details.

Available Tools

5 tools
get_event_implementationB

Get complete event specification with all properties expanded. Use when implementing tracking code.

ParametersJSON Schema
NameRequiredDescriptionDefault
event_nameYesName of the event to retrieve

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions retrieving 'complete event specification with all properties expanded,' which implies a read-only operation, but doesn't address potential behavioral traits like error handling, rate limits, authentication needs, or what 'complete' entails. This leaves significant gaps for a tool with no structured safety hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and front-loaded, consisting of just two sentences that directly state the purpose and usage without any wasted words. Every sentence earns its place by providing essential information efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete for a tool that retrieves event specifications. It doesn't explain what 'complete event specification' includes, the return format, or any behavioral aspects like pagination or errors. For a read operation with no structured output, more context is needed to fully guide the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the parameter 'event_name' clearly documented as 'Name of the event to retrieve.' The description doesn't add any additional meaning or context beyond this, such as format examples or constraints, so it meets the baseline score when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Get') and resource ('complete event specification with all properties expanded'), making it easy to understand what it does. However, it doesn't explicitly distinguish this from sibling tools like 'get_property_details' or 'get_related_events', which might also retrieve event-related information, so it doesn't reach the highest score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool ('Use when implementing tracking code'), which helps guide the agent. However, it doesn't specify when not to use it or name alternatives among the sibling tools, such as when to choose 'search_events' instead, so it lacks explicit exclusions or comparisons.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_property_detailsB

Get property definition and see where it is used across events and property groups.

ParametersJSON Schema
NameRequiredDescriptionDefault
property_nameYesName of the property to retrieve

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It describes a read operation ('Get'), implying it's likely non-destructive, but doesn't specify permissions, rate limits, error conditions, or response format. For a tool with zero annotation coverage, this leaves significant gaps in understanding its behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core purpose ('Get property definition') and adds useful context ('see where it is used across events and property groups'). There is no wasted language, and it's appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (single parameter, no output schema, no annotations), the description is minimally adequate. It explains what the tool does but lacks details on behavioral traits, usage context, or output. Without annotations or an output schema, the agent has incomplete information about how to interpret results or handle errors.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the single parameter 'property_name' documented as 'Name of the property to retrieve'. The description adds no additional meaning beyond this, such as format examples or constraints. With high schema coverage, the baseline score of 3 is appropriate, as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Get') and resource ('property definition'), and specifies what information is retrieved ('see where it is used across events and property groups'). However, it doesn't explicitly differentiate this tool from its sibling tools like 'get_event_implementation' or 'search_events', which might also involve property-related queries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any prerequisites, exclusions, or specific contexts for usage, nor does it reference sibling tools. The agent must infer usage based solely on the purpose statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_eventsC

Search for events by name, description, table, or property usage.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNoSearch term for event name or description
tableNoFilter by event table
has_propertyNoFilter events that include this property

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the search functionality but doesn't describe key behavioral traits like whether this is a read-only operation, potential rate limits, authentication needs, or what the output looks like (e.g., list format, pagination). For a search tool with zero annotation coverage, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It's appropriately sized and front-loaded, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a search tool with 3 parameters, no annotations, and no output schema, the description is incomplete. It doesn't explain return values, error handling, or behavioral constraints, leaving the agent with insufficient context for effective tool invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds minimal value beyond the input schema, which has 100% coverage with clear parameter descriptions. It mentions searchable attributes ('name, description, table, or property usage') that map to the parameters, but doesn't provide additional syntax, format details, or usage examples. This meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs ('Search for events') and resources ('events'), and identifies searchable attributes ('by name, description, table, or property usage'). However, it doesn't explicitly differentiate from sibling tools like 'get_related_events' or 'validate_event_payload', which might also involve event queries, so it doesn't reach the highest clarity level.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools or contexts where other tools might be more appropriate, such as using 'get_event_implementation' for detailed event data or 'validate_event_payload' for validation tasks. This leaves the agent without explicit usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_event_payloadB

Validate a tracking implementation payload against the event spec. Returns errors, warnings, and valid fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
event_nameYesName of the event to validate against
payloadYesThe payload object to validate

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool validates a payload and returns errors, warnings, and valid fields, which covers the basic operation and output. However, it lacks details on error handling, validation rules, performance implications, or any side effects, leaving gaps in transparency for a validation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded, consisting of a single sentence that directly states the tool's function and output. There is no wasted text, but it could be slightly improved by structuring usage hints or examples without sacrificing brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (validation with two parameters, no output schema, and no annotations), the description is minimally adequate. It explains what the tool does and its output, but lacks details on error types, validation scope, or integration with sibling tools. Without an output schema, more information on return values would be beneficial, but it meets the basic threshold.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with clear documentation for both parameters ('event_name' and 'payload'). The description adds no additional semantic details beyond what the schema provides, such as format examples or validation criteria. According to the rules, with high schema coverage, the baseline is 3, which is appropriate here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Validate a tracking implementation payload against the event spec.' It specifies the verb (validate) and resource (payload against event spec), making the function unambiguous. However, it doesn't explicitly differentiate from sibling tools like 'get_event_implementation' or 'search_events', which might also involve event-related operations, so it doesn't reach the highest score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites, context, or exclusions, nor does it reference sibling tools like 'get_event_implementation' or 'search_events' for comparison. This leaves the agent without clear usage instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updates
    • First observedget_event_implementation
    • First observedget_property_details
    • First observedget_related_events
    • First observedsearch_events
    • First observedvalidate_event_payload

TDQS

A3.5/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a clearly distinct purpose with no overlap: get_event_implementation retrieves full event specs, get_property_details focuses on property definitions, get_related_events finds contextual events, search_events performs searches, and validate_event_payload validates payloads. The descriptions explicitly state when to use each tool, eliminating ambiguity.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern with snake_case (e.g., get_event_implementation, search_events, validate_event_payload). The verbs are descriptive and aligned with the actions (get, search, validate), making the set predictable and readable.

Tool Count5/5

With 5 tools, this server is well-scoped for wiki analytics specification management. Each tool serves a specific role in event and property handling, from retrieval to validation, without being overly sparse or bloated. The count aligns perfectly with the domain's needs.

Completeness4/5

The tool set covers core workflows comprehensively: retrieving event specs, property details, related events, searching, and payload validation. A minor gap exists in update or creation tools for modifying specifications, but agents can work around this for read-only and validation tasks in the analytics domain.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers