pm-copilot
Incorporates acquisition metrics and traffic trends from Google Analytics to provide broader business context for synthesizing customer feedback and generating product plans.
Integrates quantitative business metrics such as churn and revenue data from Metabase into the product planning process to refine feature prioritization.
PM Copilot
An MCP server that triangulates customer support tickets, feature requests, and AI support agent conversations to help PMs decide what to build next.
Real results: Analyzed 3,353 signals in one 30-day window: 1,678 support tickets, 276 feature requests, and 1,399 AI support agent conversations across 4 products.
The chats are signal a ticket-only analysis never sees.
Read the full story: I built an MCP server that changed how I prioritize products
What makes this different
Signal triangulation. Matches support tickets against feature requests to find convergent themes, and gives convergent themes a 2x priority boost.
The deflection blind spot. An AI support agent answers questions that never become tickets, so ticket-based prioritization undercounts every theme the bot handles. Chatbase conversations come in as a third signal class, with a per-theme
self_serve_failure_rate.Composability. Pass churn or traffic data from other MCP servers into
generate_product_planviakpi_context, and the methodology adjusts priorities.
Related MCP server: MindBacklog
Architecture
graph TD
A[Claude Desktop / Code] -->|stdio| B[pm-copilot]
A -->|stdio| C[Metabase MCP]
A -->|stdio| D[Google Analytics MCP]
B -->|Reactive| E[HelpScout: tickets]
B -->|Proactive| F[ProductLift: feature requests]
B -->|Deflected| I[Chatbase: AI agent chats]
C -->|Quantitative| G[Conversion, Churn, Revenue]
D -->|Acquisition| H[Traffic, Channels, Trends]
B -.->|kpi_context| AQuick start
Requires Node 20+. Not published to npm, so install from source:
git clone https://github.com/dkships/pm-copilot.git
cd pm-copilot
npm install
cp .env.example .env # Edit with your credentials
npm run build.env lives in the repo root. The server loads it from there regardless of the working directory it's launched from.
Credentials
Configure at least one source; the analysis adapts to whichever you set up. Without HelpScout there are no support tickets, so themes get no severity score and no convergence boost; ProductLift and Chatbase still rank themes by frequency and votes.
Variable | Required | Description |
| No | OAuth app ID from https://secure.helpscout.net/apps/custom/ (set both or neither) |
| No | OAuth app secret |
| No | Multi-portal: |
| No | Single portal URL |
| No | Single portal Bearer token |
| No | Portal display name (default: |
| No | Account-wide secret key from Chatbase → Settings → API keys |
| No | Multi-agent: |
| No | Single agent id |
| No | Single agent display name (default: |
Chatbase API access needs a Standard plan or higher; on a lower plan the deflection signal becomes a warning. One agent per product gives product-level attribution a shared mailbox doesn't.
Claude Desktop
Add to ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"pm-copilot": {
"command": "node",
"args": ["/absolute/path/to/pm-copilot/dist/index.js"]
}
}
}Claude Code
claude mcp add pm-copilot -- node /absolute/path/to/pm-copilot/dist/index.jsOr open Claude Code in the repo: it prompts you to approve the project .mcp.json.
Verify
Restart the client and ask it to run list_sources. It should list the sources you configured: HelpScout mailboxes, ProductLift portals and Chatbase agents.
Tools
Common filters
Shared by synthesize_feedback and generate_product_plan.
Parameter | Type | Default | Description |
| number | 30 | Days to look back (1-90) |
| number | 50 | Top-voted requests per portal (1-200). Recent requests in the timeframe are always included on top |
| string | — | HelpScout mailbox ID |
| string | — | HelpScout mailbox name (case-insensitive), resolved to an ID |
| string | — | ProductLift portal |
| string | — | Chatbase agent |
| string | — | Chatbase conversation source, comma-separated for multiple (e.g. |
| boolean | false | Also fetch customer comment text on feature requests (scrubbed; names and admin replies dropped) for theme matching and quotes. A modest gain for one extra call per request with comments; can take close to a minute on large portals |
| string |
|
|
Run list_sources to see valid mailbox, portal, agent and source names.
synthesize_feedback
Returns themes sorted by priority score, each with per-class counts, a convergence flag, an evidence summary and representative quotes. Roughly 15KB at summary, several hundred KB at full. Common filters only.
generate_product_plan
Builds a prioritized plan with evidence and customer quotes. Takes the common filters plus:
Parameter | Type | Default | Description |
| string | — | Business metrics from other MCP servers, passed through verbatim |
| number | 5 | Number of priorities to return (1-10) |
| boolean | false | Audit mode: show what data would be sent, without fetching it |
| string |
|
|
get_theme_evidence
Drill into one theme: the individual tickets, feature requests and chats behind it, newest first, with ticket numbers, request URLs, votes, channels and dates. Pass the same common filters within a few minutes of the analysis call and it reuses the cached data, so it makes no new API calls. Returns identifiers, metadata and scrubbed titles (for a chat, its opening customer message, truncated to 200 characters), not full conversations.
Parameter | Type | Default | Description |
| string | — | The |
| string |
|
|
| number | 25 | Records per source (1-200), so a busy source can't crowd out the others |
Plus the common filters.
get_feature_requests
Raw ProductLift access. Each request includes its public url.
Parameter | Type | Default | Description |
| string | — | Filter to one portal |
| boolean | true | Include comments on each request |
| string | — | Filter by status (case-insensitive), e.g. |
| number | — | Requests to return per portal (1-500), after the status filter and sort. Comments are fetched only for what's kept |
| string | — |
|
list_sources
Lists configured mailboxes, portals and agents, plus chatbase_conversation_sources (the values source_filter accepts) when Chatbase is set up. Never returns keys or customer data. No parameters.
Signal classes
Class | Source | What it means | Feeds |
Reactive | HelpScout tickets | Something is broken | Frequency, severity, convergence |
Proactive | ProductLift requests | Something is wanted | Frequency, vote momentum, convergence |
Deflected | Chatbase conversations | Something was asked, and self-serve either handled it or didn't | Frequency only |
Deflected signals never affect severity, vote momentum or the convergence boost. Each theme carries:
deflected_count: conversations matching the themeself_serve_failure_rate: share of those where the agent's lowest answer confidence fell below 0.5mean_answer_confidence: mean of that same score
Chatbase doesn't document what its min_score measures, so these are evidence for the LLM to weigh, not part of the score. The analysis also counts conversations per channel (chatbase_sources). An unrecognized source_filter value is passed through with a warning, not rejected. Without Chatbase, the deflection fields are absent.
Example output
A trimmed synthesize_feedback response at summary detail. Values are illustrative. Note the scrubbed email in the first quote.
{
"timeframe_days": 30,
"detail_level": "summary",
"pii_scrubbing_applied": true,
"pii_categories_redacted": ["email", "phone", "credit_card"],
"analysis": {
"total_data_points": 924,
"reactive_count": 548,
"proactive_count": 64,
"deflected_count": 312,
"themes": [
{
"theme_id": "booking-scheduling",
"label": "Booking & Scheduling",
"priority_score": 78.4,
"convergent": true,
"reactive_count": 211,
"proactive_count": 19,
"deflected_count": 96,
"self_serve_failure_rate": 0.41,
"representative_quotes": [
"[Support ticket] \"Double-booked slots again after the timezone change — reach me at [EMAIL REDACTED]\"",
"[Feature request, 47 votes] \"Let me block buffer time between meetings\"",
"[AI chat, answer confidence 0.31] \"how do i stop people booking on weekends\""
]
}
],
"emerging_themes": [{ "pattern": "csv export", "frequency": 12 }],
"unmatched_count": 38
}
}Composability
Ask Claude to pull churn and conversion data from your other MCP servers and pass it as kpi_context:
Product A: booking completion rate dropped from 74% to 66% over last
30 days. Monthly churn increased from 3.1% to 4.2%. Organic traffic
up 22% MoM. Product B: document completion rate steady at 81%.
Churn flat at 2.8%.The methodology says churn overrides the formula, so a theme tied to Product A's falling completion rate can jump to #1 even when another theme scores higher. The server ranks the signal; the KPI context supplies the judgment.
Methodology
The pm-copilot://methodology resource is my product planning framework from 7 years of launching 9 products to 1M+ users. The core rules:
The 5% rule. You complete about 5% of what customers ask for each month. The framework picks which 5%.
Convergent signals win. A theme in both tickets and feature requests is the highest-confidence signal.
Reactive > proactive. Broken stuff drives churn. You can survive a missing feature; you can't survive errors.
Business metrics override the formula. Rising churn or dropping conversion changes everything.
It's versioned (v2.2). Every generate_product_plan response links to it, and tells Claude to apply it when kpi_context is set. Whether it gets read depends on the client surfacing resources.
Evaluation
Themes are matched with keyword lists, not embeddings or an LLM classifier. Customer text never leaves the server, and the same input always produces the same themes, so a ranking can be audited. The cost is recall.
npm run eval measures it. On the committed 86-example fixture, config v3 scores micro precision 96.1%, recall 99.0%, F1 97.5%, with a 1.3% miss rate. That number is in-sample (the config was tuned against it), so it's a regression gate. On held-out real chat data, a third of conversations still match no theme.
Full results, what the first run found, and known limits: docs/evaluation.md.
Security
All customer text is scrubbed before it enters the analysis or leaves the server:
SSNs, credit cards (Luhn-validated), email addresses, and phone numbers (US formats and
+-prefixed international) are redacted, and scrubbed from feature-request URLs as well. The customer email field is always[REDACTED].Agent/admin replies, internal notes, attachments, voter identities, commenter names, and Chatbase assistant turns, lead forms, user IDs and country are excluded entirely.
preview_only: trueongenerate_product_planshows what would be sent without fetching data.Every response includes
pii_scrubbing_appliedandpii_categories_redacted.
Details, known limitations and the reporting process: SECURITY.md.
Theme configuration
themes.config.json in the repo root defines the themes. It's read at runtime, so edits don't need a rebuild. It ships with 18 themes across 12 categories; add your own to the themes array. Unmatched data points are mined for emerging patterns with bigram/trigram frequency.
Single-word keywords match on a word boundary with an optional regular plural. Multi-word keywords also match on word boundaries. After editing, run npm run eval to catch keywords that fire on the wrong theme.
Scoring formula
priority = (frequency × 0.35 + severity × 0.35 + vote_momentum × 0.30) × convergence_boostFrequency (0.35): data point count, normalized across themes. Includes deflected signals.
Severity (0.35): reactive signals only. Thread count, recency (7-day half-life decay), and a boost from the highest-severity matching tag.
Vote momentum (0.30): proactive signals only. 80% votes, 20% comments.
Convergence (2x): applied when a theme has both reactive and proactive signals. Deflected signals don't trigger it.
Frequency and vote momentum are normalized against the top theme in the same call, so scores are relative to one analysis window. Compare rankings across calls, not raw scores.
Troubleshooting
No data sources configured. Check that.envexists in the repo root and sets at least one source.HELPSCOUT_APP_SECRET is missing(or_ID). Set both HelpScout values, or remove both to run without HelpScout.HelpScout auth expired or invalid (403)on every call. If the token request succeeds but API calls get 403, the HelpScout user who owns the OAuth app was deactivated or lost access. Rotating the secret won't help; create a new app from an active user's profile and update bothHELPSCOUT_*values.Changes aren't taking effect. The client runs the compiled
dist/. Runnpm run buildand restart the client.No HelpScout mailbox named "…". Runlist_sourcesfor exact names, or passmailbox_id.No portal found with name "…"/No ProductLift portal named "…". The portal must be inPRODUCTLIFT_PORTALS(or the single-portal vars). Runlist_sources.Chatbase warning:
API access needs a Chatbase Standard plan or higher. The rest of the analysis still runs; only the deflection fields are missing.chatbase_agentsis empty inlist_sources. Set bothCHATBASE_API_KEYand one ofCHATBASE_AGENTS/CHATBASE_AGENT_ID. A key alone configures nothing.
Contributing
See CONTRIBUTING.md.
License
This server cannot be deployed
Maintenance
Related MCP Connectors
- SquadOAuthai.meetsquad
Decision intelligence for product teams. Turn scattered feedback into signal you can act on.
Turn raw customer feedback into evidence-cited specs (free, no key) plus 16 PM tools.
Analyze customer feedback at scale — reviews, surveys, calls. AI-powered themes and sentiment.
Synthesize GitHub Issues, HN and App Store reviews into ranked pain clusters. Pay-per-call x402.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceTransforms scattered customer feedback from sources like Slack, Zoom, and JIRA into actionable product insights and AI-generated PRDs. It features over 50 tools for semantic clustering, sentiment analysis, and VOC-based prioritization to streamline product management workflows.1MIT
- AlicenseNot gradedqualityDmaintenanceConnect Claude directly to your MindBacklog workspace to read PRDs, user stories, and real-time customer signals.3 npmMIT
- AlicenseNot gradedqualityDmaintenanceEnables product management (add, fetch) through natural language using Claude Desktop and GitHub Copilot.MIT
- FlicenseNot gradedqualityDmaintenanceProvides Claude with tools to make Ship/Delay/Kill recommendations on feature requests using product metrics, roadmap, and OKRs.-