Founder Intelligence Engine
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Founder Intelligence Engineanalyze Sam Altman's profile and fetch the latest relevant news"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Founder Intelligence Engine — MCP Server
A production-grade Model Context Protocol (MCP) server that transforms founder profiles into actionable strategic intelligence.
Architecture
┌───────────────────────────────────────────────────────────┐
│ MCP Client (Claude, etc.) │
│ ▲ stdio │
│ ┌──────────┴──────────┐ │
│ │ MCP Server (Node) │ │
│ │ 3 registered tools│ │
│ └──────┬──────────────┘ │
│ ┌───────────┬┼──────────────┐ │
│ ▼ ▼▼ ▼ │
│ ┌──────────┐ ┌───────────┐ ┌──────────────┐ │
│ │ Apify │ │ Groq │ │ Embeddings │ │
│ │ Scraping│ │ LLM │ │ API │ │
│ └────┬─────┘ └─────┬─────┘ └──────┬───────┘ │
│ └──────────────┬┘──────────────┘ │
│ ▼ │
│ ┌─────────────────┐ │
│ │ Supabase │ │
│ │ (Postgres + │ │
│ │ pgvector) │ │
│ └─────────────────┘ │
└───────────────────────────────────────────────────────────┘Data Flow
collect_profile — Scrapes LinkedIn + Twitter via Apify → merges data → generates embedding → stores in Supabase
analyze_profile — Fetches stored profile → calls Groq LLM for strategic analysis → caches result
fetch_personalized_news — Checks cache freshness → if stale: generates search queries → scrapes Google News → embeds articles → ranks by cosine similarity → summarizes with Groq → stores; if fresh: returns cached articles
Caching & Cost Optimization
Operation | Cost | When It Runs |
LinkedIn/Twitter scraping | High | Only on profile creation |
Groq profile analysis | Medium | Once per profile (cached) |
Google News + embeddings | High | Only when news > 24h stale |
Read cached articles | Free | Every subsequent request |
The fetch_history table tracks last_profile_scrape and last_news_fetch timestamps. The staleCheck.js module compares these against configurable thresholds.
Related MCP server: FundzWatch MCP Server
Setup
1. Prerequisites
Node.js 20+
Supabase project (with pgvector enabled)
API keys: Apify, Groq, OpenAI-compatible Embeddings
2. Install
cd /Users/praveenkumar/Desktop/mcp
cp .env.example .env
# Edit .env with your real keys
npm install3. Database
Run the migration against your Supabase SQL Editor:
-- Paste contents of migrations/001_init.sqlOr via psql:
psql $DATABASE_URL < migrations/001_init.sql4. Run MCP Server
node src/index.js5. Configure MCP Client
Add to your MCP client config (e.g., Claude Desktop claude_desktop_config.json):
{
"mcpServers": {
"founder-intelligence": {
"command": "node",
"args": ["/Users/praveenkumar/Desktop/mcp/src/index.js"],
"env": {
"SUPABASE_URL": "...",
"SUPABASE_SERVICE_KEY": "...",
"APIFY_API_TOKEN": "...",
"GROQ_API_KEY": "...",
"EMBEDDING_API_URL": "...",
"EMBEDDING_API_KEY": "..."
}
}
}
}6. Background Worker (Optional)
# Single run (for cron)
node src/backgroundWorker.js
# Daemon mode
BACKGROUND_LOOP=true node src/backgroundWorker.jsCron example (every 6 hours):
0 */6 * * * cd /app && node src/backgroundWorker.js >> /var/log/worker.log 2>&1Project Structure
/Users/praveenkumar/Desktop/mcp/
├── migrations/
│ └── 001_init.sql
├── src/
│ ├── db/
│ │ └── supabaseClient.js
│ ├── services/
│ │ ├── apifyService.js
│ │ ├── embeddingService.js
│ │ └── llmService.js
│ ├── tools/
│ │ ├── collectProfile.js
│ │ ├── analyzeProfile.js
│ │ └── fetchPersonalizedNews.js
│ ├── utils/
│ │ ├── similarity.js
│ │ └── staleCheck.js
│ ├── backgroundWorker.js
│ └── index.js
├── .env.example
├── .gitignore
├── .dockerignore
├── Dockerfile
├── package.json
└── README.mdDocker Deployment
Build & Run
docker build -t founder-intelligence-mcp .
docker run --env-file .env founder-intelligence-mcpBackground Worker Container
docker run --env-file .env founder-intelligence-mcp node src/backgroundWorker.jsDocker Compose (production)
version: '3.8'
services:
mcp-server:
build: .
env_file: .env
stdin_open: true
restart: unless-stopped
worker:
build: .
env_file: .env
command: ["node", "src/backgroundWorker.js"]
environment:
- BACKGROUND_LOOP=true
restart: unless-stoppedScaling Strategy
Component | Strategy |
MCP Server | One instance per client (stdio-based) |
Background Worker | Single instance or Cloud Run Job on schedule |
Supabase | Connection pooling via Supavisor; read replicas for scale |
Apify | Concurrent actor runs (up to account limit) |
Embeddings | Batch requests (20 per call) to reduce round trips |
Groq | Rate-limit aware with retry-after header handling |
For high-profile-count deployments:
Move background worker to a Cloud Run Job triggered by Cloud Scheduler
Use Supabase Edge Functions for scheduled refresh
Add a Redis cache layer for hot profile lookups
Security Best Practices
Service-role key only on server side — never expose to clients
All secrets via environment variables — no hardcoded keys
Non-root Docker user —
mcpuser in containerInput validation — Zod schemas on all tool inputs
Row Level Security — enable RLS on Supabase tables for multi-tenant
API token rotation — rotate Apify, Groq, and embedding keys periodically
Rate limiting — built-in retry logic with exponential backoff
No PII logging — profile data stays in Supabase, not console
Cost Optimization
Service | Cost Driver | Mitigation |
Apify | Actor compute units | Scrape only on creation; cache results |
Groq | Token usage | Analyze once (cached); batch news summaries |
Embeddings | API calls | Batch 20 at a time; embed once per article |
Supabase | Row count + storage | Deduplicate articles by URL; prune old articles |
Expected cost per profile lifecycle:
Initial setup: ~$0.05–0.15 (scrape + embed + analyze)
Daily news refresh: ~$0.02–0.08 (scrape + embed + summarize top 10)
Cached reads: $0.00
Future Improvement Roadmap
HTTP/SSE transport — support remote MCP clients over HTTP
Multi-tenant profiles — user-scoped access with RLS
Real-time alerts — push notifications when high-relevance news drops
Competitor tracking — dedicated tool to monitor named competitors
Founder network graph — map connections between analyzed founders
Custom embedding models — fine-tuned models for startup/VC domain
Article full-text extraction — deep content scraping for richer embeddings
A/B prompt testing — experiment with different Groq prompts for analysis quality
Dashboard UI — web interface for browsing intelligence feeds
Webhook integrations — push intelligence to Slack, email, or CRM
Available Tools
3 toolsanalyze_profileA
Analyze a founder's profile using Groq LLM to extract strategic intelligence: interests, industry tags, market positioning, competitors, pain points, and growth opportunities. Returns cached analysis if available.
| Name | Required | Description | Default |
|---|---|---|---|
| profile_id | Yes | UUID of the profile to analyze (from collect_profile) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key behaviors: using Groq LLM for analysis, extracting specific intelligence types, and returning cached results if available. This covers the tool's operational approach and performance optimization, though it omits details like error handling or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is highly concise and well-structured in two sentences: the first states the core action and outputs, and the second adds crucial behavioral detail about caching. Every sentence earns its place with no redundant information, making it easy to parse and front-loaded with essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (analysis with LLM, caching), no annotations, and no output schema, the description is largely complete. It covers the purpose, method, outputs, and caching behavior. However, it lacks details on output format or potential errors, leaving minor gaps for an agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, with the single parameter 'profile_id' fully documented in the schema. The description adds minimal value beyond the schema by noting the profile_id comes 'from collect_profile', providing slight contextual meaning. This meets the baseline score of 3 for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('analyze', 'extract') and resources ('founder's profile', 'Groq LLM'), listing concrete outputs like interests, industry tags, and pain points. It distinguishes from sibling tools by focusing on analysis rather than collection (collect_profile) or news fetching (fetch_personalized_news).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when strategic intelligence is needed from a founder's profile, but provides no explicit guidance on when to use this tool versus alternatives. It mentions a prerequisite ('profile_id from collect_profile'), which offers some context, but lacks clear when-not-to-use scenarios or comparisons to other analysis methods.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collect_profileA
Collect a founder's profile by scraping LinkedIn and Twitter via Apify, merge all data, generate an embedding, and store it in Supabase. Returns the profile_id for subsequent analysis.
| Name | Required | Description | Default |
|---|---|---|---|
| linkedin_url | No | Full LinkedIn profile URL (e.g., https://linkedin.com/in/username) | |
| twitter_url | No | Full Twitter/X profile URL (e.g., https://x.com/username) | |
| business_idea | No | Description of the founder's business idea or venture |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It describes the multi-step process (scraping, merging, embedding, storing) and the return value (profile_id), which adds useful context. However, it lacks details on potential issues like rate limits, authentication needs, error handling, or data privacy implications, which are important for a tool involving web scraping and data storage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that efficiently outlines the entire workflow from input to output. It front-loads the core purpose and includes all necessary steps without unnecessary details, making it highly concise and easy to understand.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the tool (involving multiple external services and data processing), no annotations, and no output schema, the description does a good job of explaining the process and return value. However, it could be more complete by addressing potential behavioral aspects like error cases or performance considerations, which would help an agent use it more effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters (linkedin_url, twitter_url, business_idea) with their types and descriptions. The description doesn't add any additional meaning or usage context for these parameters beyond what the schema provides, meeting the baseline for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('collect a founder's profile'), resources involved (LinkedIn and Twitter via Apify), and the full workflow including merging data, generating embeddings, and storing in Supabase. It distinguishes itself from sibling tools like 'analyze_profile' and 'fetch_personalized_news' by focusing on data collection and storage rather than analysis or content retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by specifying it's for collecting founder profiles and mentions returning a profile_id 'for subsequent analysis,' suggesting this tool is for initial data gathering. However, it doesn't explicitly state when not to use it or provide direct alternatives to sibling tools, leaving some guidance gaps.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_personalized_newsA
Fetch personalized, strategically-relevant news for a founder. Checks cache freshness first — returns stored articles if <24h old. Otherwise scrapes Google News via Apify, ranks by cosine similarity, summarizes with Groq, and stores results.
| Name | Required | Description | Default |
|---|---|---|---|
| profile_id | Yes | UUID of the profile to fetch news for (must have run analyze_profile first) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden and does an excellent job disclosing behavioral traits: it explains the caching logic, external service dependencies (Google News via Apify), ranking methodology (cosine similarity), summarization process (Groq), and storage behavior. The only minor gap is lack of explicit error handling or rate limit information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly front-loaded with the core purpose, followed by efficient explanation of the multi-step process. Every sentence earns its place by adding distinct value about the tool's behavior and implementation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with no annotations and no output schema, the description provides substantial context about the multi-step process, caching behavior, and external dependencies. The main gap is lack of information about return format or error conditions, which would be helpful given the absence of output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents the single parameter. The description adds context about the prerequisite (must have run analyze_profile first) which provides useful semantic meaning beyond the schema's technical specification.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('fetch personalized, strategically-relevant news') and target resource ('for a founder'), distinguishing it from sibling tools like analyze_profile and collect_profile which handle profile analysis rather than news retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context about when to use this tool (after running analyze_profile first, as implied by the profile_id requirement) and mentions the caching behavior (<24h freshness). However, it doesn't explicitly state when NOT to use it or name specific alternatives among sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
analyze_profile - First observed
collect_profile - First observed
fetch_personalized_news
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: analyze_profile extracts intelligence from a profile, collect_profile scrapes and stores profile data, and fetch_personalized_news retrieves and processes news. There is no overlap in functionality, making it easy for an agent to select the right tool.
All tool names follow a consistent verb_noun pattern (analyze_profile, collect_profile, fetch_personalized_news), using snake_case throughout. The verbs (analyze, collect, fetch) are distinct and appropriately descriptive for their actions.
With 3 tools, the server is well-scoped for its purpose of founder intelligence, covering profile collection, analysis, and news personalization. However, it feels slightly thin as it lacks tools for updating or deleting data, which might be needed for a complete lifecycle.
The tools cover the core workflows of collecting, analyzing, and fetching news for founder profiles, with no dead ends. A minor gap exists in lacking update or deletion operations for profiles or cached data, but agents can work around this for most use cases.
Maintenance
Related MCP Connectors
Pre-diligence AI for founders, investors, and firms — multi-agent pitch analysis and deal flow.
Company and market intelligence, news, enrichment, and agentic workflows for dealmakers.
Investment research superagent: podcasts, SEC filings, and no-code research pipelines.
Search startups and founders, match companies to products, and manage outreach lists.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables detection and analysis of pre-public product launches through web search, content extraction, AI-powered scoring, and automated alerting. Provides comprehensive tools for surfacing stealth startup signals before they trend publicly.MIT
- AlicenseAqualityCmaintenanceProvides real-time business event intelligence and AI-scored sales leads to help users track funding rounds, acquisitions, and executive hires. It enables AI agents to generate strategic market briefs and manage company watchlists for predictive business insights.759 npm3MIT
- AlicenseBqualityCmaintenanceAgentic pipeline that transforms ideas to revenue — for solo founders and bootstrappers.4513 npm4MIT
- AlicenseAqualityDmaintenanceProvides Hacker News tools with founder-research features for analyzing Show HN launches, Ask HN discussions, and extracting startup insights.14MIT