Skip to main content
Glama
praveenkumarkunchala2005

Founder Intelligence Engine

Founder Intelligence Engine — MCP Server

A production-grade Model Context Protocol (MCP) server that transforms founder profiles into actionable strategic intelligence.


Architecture

┌───────────────────────────────────────────────────────────┐
│                     MCP Client (Claude, etc.)             │
│                          ▲ stdio                          │
│               ┌──────────┴──────────┐                     │
│               │   MCP Server (Node) │                     │
│               │   3 registered tools│                     │
│               └──────┬──────────────┘                     │
│          ┌───────────┬┼──────────────┐                    │
│          ▼           ▼▼              ▼                    │
│  ┌──────────┐  ┌───────────┐  ┌──────────────┐           │
│  │  Apify   │  │   Groq    │  │  Embeddings  │           │
│  │  Scraping│  │   LLM     │  │  API         │           │
│  └────┬─────┘  └─────┬─────┘  └──────┬───────┘           │
│       └──────────────┬┘──────────────┘                    │
│                      ▼                                    │
│            ┌─────────────────┐                            │
│            │  Supabase       │                            │
│            │  (Postgres +    │                            │
│            │   pgvector)     │                            │
│            └─────────────────┘                            │
└───────────────────────────────────────────────────────────┘

Data Flow

  1. collect_profile — Scrapes LinkedIn + Twitter via Apify → merges data → generates embedding → stores in Supabase

  2. analyze_profile — Fetches stored profile → calls Groq LLM for strategic analysis → caches result

  3. fetch_personalized_news — Checks cache freshness → if stale: generates search queries → scrapes Google News → embeds articles → ranks by cosine similarity → summarizes with Groq → stores; if fresh: returns cached articles

Caching & Cost Optimization

Operation

Cost

When It Runs

LinkedIn/Twitter scraping

High

Only on profile creation

Groq profile analysis

Medium

Once per profile (cached)

Google News + embeddings

High

Only when news > 24h stale

Read cached articles

Free

Every subsequent request

The fetch_history table tracks last_profile_scrape and last_news_fetch timestamps. The staleCheck.js module compares these against configurable thresholds.


Related MCP server: FundzWatch MCP Server

Setup

1. Prerequisites

  • Node.js 20+

  • Supabase project (with pgvector enabled)

  • API keys: Apify, Groq, OpenAI-compatible Embeddings

2. Install

cd /Users/praveenkumar/Desktop/mcp
cp .env.example .env
# Edit .env with your real keys
npm install

3. Database

Run the migration against your Supabase SQL Editor:

-- Paste contents of migrations/001_init.sql

Or via psql:

psql $DATABASE_URL < migrations/001_init.sql

4. Run MCP Server

node src/index.js

5. Configure MCP Client

Add to your MCP client config (e.g., Claude Desktop claude_desktop_config.json):

{
  "mcpServers": {
    "founder-intelligence": {
      "command": "node",
      "args": ["/Users/praveenkumar/Desktop/mcp/src/index.js"],
      "env": {
        "SUPABASE_URL": "...",
        "SUPABASE_SERVICE_KEY": "...",
        "APIFY_API_TOKEN": "...",
        "GROQ_API_KEY": "...",
        "EMBEDDING_API_URL": "...",
        "EMBEDDING_API_KEY": "..."
      }
    }
  }
}

6. Background Worker (Optional)

# Single run (for cron)
node src/backgroundWorker.js

# Daemon mode
BACKGROUND_LOOP=true node src/backgroundWorker.js

Cron example (every 6 hours):

0 */6 * * * cd /app && node src/backgroundWorker.js >> /var/log/worker.log 2>&1

Project Structure

/Users/praveenkumar/Desktop/mcp/
├── migrations/
│   └── 001_init.sql
├── src/
│   ├── db/
│   │   └── supabaseClient.js
│   ├── services/
│   │   ├── apifyService.js
│   │   ├── embeddingService.js
│   │   └── llmService.js
│   ├── tools/
│   │   ├── collectProfile.js
│   │   ├── analyzeProfile.js
│   │   └── fetchPersonalizedNews.js
│   ├── utils/
│   │   ├── similarity.js
│   │   └── staleCheck.js
│   ├── backgroundWorker.js
│   └── index.js
├── .env.example
├── .gitignore
├── .dockerignore
├── Dockerfile
├── package.json
└── README.md

Docker Deployment

Build & Run

docker build -t founder-intelligence-mcp .
docker run --env-file .env founder-intelligence-mcp

Background Worker Container

docker run --env-file .env founder-intelligence-mcp node src/backgroundWorker.js

Docker Compose (production)

version: '3.8'
services:
  mcp-server:
    build: .
    env_file: .env
    stdin_open: true
    restart: unless-stopped

  worker:
    build: .
    env_file: .env
    command: ["node", "src/backgroundWorker.js"]
    environment:
      - BACKGROUND_LOOP=true
    restart: unless-stopped

Scaling Strategy

Component

Strategy

MCP Server

One instance per client (stdio-based)

Background Worker

Single instance or Cloud Run Job on schedule

Supabase

Connection pooling via Supavisor; read replicas for scale

Apify

Concurrent actor runs (up to account limit)

Embeddings

Batch requests (20 per call) to reduce round trips

Groq

Rate-limit aware with retry-after header handling

For high-profile-count deployments:

  • Move background worker to a Cloud Run Job triggered by Cloud Scheduler

  • Use Supabase Edge Functions for scheduled refresh

  • Add a Redis cache layer for hot profile lookups


Security Best Practices

  1. Service-role key only on server side — never expose to clients

  2. All secrets via environment variables — no hardcoded keys

  3. Non-root Docker user — mcp user in container

  4. Input validation — Zod schemas on all tool inputs

  5. Row Level Security — enable RLS on Supabase tables for multi-tenant

  6. API token rotation — rotate Apify, Groq, and embedding keys periodically

  7. Rate limiting — built-in retry logic with exponential backoff

  8. No PII logging — profile data stays in Supabase, not console


Cost Optimization

Service

Cost Driver

Mitigation

Apify

Actor compute units

Scrape only on creation; cache results

Groq

Token usage

Analyze once (cached); batch news summaries

Embeddings

API calls

Batch 20 at a time; embed once per article

Supabase

Row count + storage

Deduplicate articles by URL; prune old articles

Expected cost per profile lifecycle:

  • Initial setup: ~$0.05–0.15 (scrape + embed + analyze)

  • Daily news refresh: ~$0.02–0.08 (scrape + embed + summarize top 10)

  • Cached reads: $0.00


Future Improvement Roadmap

  1. HTTP/SSE transport — support remote MCP clients over HTTP

  2. Multi-tenant profiles — user-scoped access with RLS

  3. Real-time alerts — push notifications when high-relevance news drops

  4. Competitor tracking — dedicated tool to monitor named competitors

  5. Founder network graph — map connections between analyzed founders

  6. Custom embedding models — fine-tuned models for startup/VC domain

  7. Article full-text extraction — deep content scraping for richer embeddings

  8. A/B prompt testing — experiment with different Groq prompts for analysis quality

  9. Dashboard UI — web interface for browsing intelligence feeds

  10. Webhook integrations — push intelligence to Slack, email, or CRM

Available Tools

3 tools
analyze_profileA

Analyze a founder's profile using Groq LLM to extract strategic intelligence: interests, industry tags, market positioning, competitors, pain points, and growth opportunities. Returns cached analysis if available.

ParametersJSON Schema
NameRequiredDescriptionDefault
profile_idYesUUID of the profile to analyze (from collect_profile)

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key behaviors: using Groq LLM for analysis, extracting specific intelligence types, and returning cached results if available. This covers the tool's operational approach and performance optimization, though it omits details like error handling or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is highly concise and well-structured in two sentences: the first states the core action and outputs, and the second adds crucial behavioral detail about caching. Every sentence earns its place with no redundant information, making it easy to parse and front-loaded with essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (analysis with LLM, caching), no annotations, and no output schema, the description is largely complete. It covers the purpose, method, outputs, and caching behavior. However, it lacks details on output format or potential errors, leaving minor gaps for an agent to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, with the single parameter 'profile_id' fully documented in the schema. The description adds minimal value beyond the schema by noting the profile_id comes 'from collect_profile', providing slight contextual meaning. This meets the baseline score of 3 for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs ('analyze', 'extract') and resources ('founder's profile', 'Groq LLM'), listing concrete outputs like interests, industry tags, and pain points. It distinguishes from sibling tools by focusing on analysis rather than collection (collect_profile) or news fetching (fetch_personalized_news).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when strategic intelligence is needed from a founder's profile, but provides no explicit guidance on when to use this tool versus alternatives. It mentions a prerequisite ('profile_id from collect_profile'), which offers some context, but lacks clear when-not-to-use scenarios or comparisons to other analysis methods.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

collect_profileA

Collect a founder's profile by scraping LinkedIn and Twitter via Apify, merge all data, generate an embedding, and store it in Supabase. Returns the profile_id for subsequent analysis.

ParametersJSON Schema
NameRequiredDescriptionDefault
linkedin_urlNoFull LinkedIn profile URL (e.g., https://linkedin.com/in/username)
twitter_urlNoFull Twitter/X profile URL (e.g., https://x.com/username)
business_ideaNoDescription of the founder's business idea or venture

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It describes the multi-step process (scraping, merging, embedding, storing) and the return value (profile_id), which adds useful context. However, it lacks details on potential issues like rate limits, authentication needs, error handling, or data privacy implications, which are important for a tool involving web scraping and data storage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that efficiently outlines the entire workflow from input to output. It front-loads the core purpose and includes all necessary steps without unnecessary details, making it highly concise and easy to understand.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool (involving multiple external services and data processing), no annotations, and no output schema, the description does a good job of explaining the process and return value. However, it could be more complete by addressing potential behavioral aspects like error cases or performance considerations, which would help an agent use it more effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters (linkedin_url, twitter_url, business_idea) with their types and descriptions. The description doesn't add any additional meaning or usage context for these parameters beyond what the schema provides, meeting the baseline for high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('collect a founder's profile'), resources involved (LinkedIn and Twitter via Apify), and the full workflow including merging data, generating embeddings, and storing in Supabase. It distinguishes itself from sibling tools like 'analyze_profile' and 'fetch_personalized_news' by focusing on data collection and storage rather than analysis or content retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context by specifying it's for collecting founder profiles and mentions returning a profile_id 'for subsequent analysis,' suggesting this tool is for initial data gathering. However, it doesn't explicitly state when not to use it or provide direct alternatives to sibling tools, leaving some guidance gaps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_personalized_newsA

Fetch personalized, strategically-relevant news for a founder. Checks cache freshness first — returns stored articles if <24h old. Otherwise scrapes Google News via Apify, ranks by cosine similarity, summarizes with Groq, and stores results.

ParametersJSON Schema
NameRequiredDescriptionDefault
profile_idYesUUID of the profile to fetch news for (must have run analyze_profile first)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden and does an excellent job disclosing behavioral traits: it explains the caching logic, external service dependencies (Google News via Apify), ranking methodology (cosine similarity), summarization process (Groq), and storage behavior. The only minor gap is lack of explicit error handling or rate limit information.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is perfectly front-loaded with the core purpose, followed by efficient explanation of the multi-step process. Every sentence earns its place by adding distinct value about the tool's behavior and implementation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with no annotations and no output schema, the description provides substantial context about the multi-step process, caching behavior, and external dependencies. The main gap is lack of information about return format or error conditions, which would be helpful given the absence of output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents the single parameter. The description adds context about the prerequisite (must have run analyze_profile first) which provides useful semantic meaning beyond the schema's technical specification.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('fetch personalized, strategically-relevant news') and target resource ('for a founder'), distinguishing it from sibling tools like analyze_profile and collect_profile which handle profile analysis rather than news retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context about when to use this tool (after running analyze_profile first, as implied by the profile_id requirement) and mentions the caching behavior (<24h freshness). However, it doesn't explicitly state when NOT to use it or name specific alternatives among sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.0
    • First observedanalyze_profile
    • First observedcollect_profile
    • First observedfetch_personalized_news

TDQS

A4.2/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: analyze_profile extracts intelligence from a profile, collect_profile scrapes and stores profile data, and fetch_personalized_news retrieves and processes news. There is no overlap in functionality, making it easy for an agent to select the right tool.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (analyze_profile, collect_profile, fetch_personalized_news), using snake_case throughout. The verbs (analyze, collect, fetch) are distinct and appropriately descriptive for their actions.

Tool Count4/5

With 3 tools, the server is well-scoped for its purpose of founder intelligence, covering profile collection, analysis, and news personalization. However, it feels slightly thin as it lacks tools for updating or deleting data, which might be needed for a complete lifecycle.

Completeness4/5

The tools cover the core workflows of collecting, analyzing, and fetching news for founder profiles, with no dead ends. A minor gap exists in lacking update or deletion operations for profiles or cached data, but agents can work around this for most use cases.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables detection and analysis of pre-public product launches through web search, content extraction, AI-powered scoring, and automated alerting. Provides comprehensive tools for surfacing stealth startup signals before they trend publicly.
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Provides real-time business event intelligence and AI-scored sales leads to help users track funding rounds, acquisitions, and executive hires. It enables AI agents to generate strategic market briefs and manage company watchlists for predictive business insights.
    7
    59 npm
    3
    MIT