Spark
Ingests Asana tasks and project activity via Composio triggers into Spark's knowledge base, enabling work items and project context to be searched, asked about, and cited.
Ingests GitHub App webhook events into Spark's signal pipeline, allowing commits, pull requests, and related repository activity to be turned into cited organizational knowledge pages.
Ingests Slack signals via Composio triggers into Spark's knowledge base, making team conversations and decisions queryable and citable through the MCP tools.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@SparkWhat did we decide about checkout payments, and has anything changed since?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Spark is headless. This repository is the product core: the Convex backend, the intelligence pipeline, and an MCP server that makes your org queryable from any AI client. There is no web app in this repo; admin works from the command line (Headless admin). You self-host it on your own Convex, Clerk, Anthropic, OpenAI and Composio accounts β plan on about an hour of vendor setup. See what ships and what doesn't.
π‘ What is Spark
A company is an organism. Every commit, message, ticket and decision encodes how it thinks β but that intelligence is scattered across dozens of tools and fades as fast as it forms. Spark wires into the systems your team already uses and turns the raw signal stream into a living, cited, self-updating understanding of the organization. No adoption required. No behavior change.
The loop is Sense β Think β Reflect β Act:
Sense β GitHub, Slack, Asana, Gorgias and any API push land in one org-scoped
eventstable.Think β an extractor clusters events into work sessions and writes knowledge pages: Compiled Truth (rewritable AI synthesis) on top, an append-only Evidence Timeline underneath.
Reflect β daily, weekly and monthly synthesis, plus an overnight Dream Cycle that reads the whole knowledge base and writes findings: contradictions, stale decisions, blind spots.
Act β you ask, search and push from any MCP client; agents you describe in plain English run on the same knowledge.
Findings feed back into the knowledge base, so each cycle starts richer than the last. That is the compounding. The long version of why lives in the essay Why Spark exists.
Related MCP server: knowledge-base
π Connect your AI
This is the point of Spark: once your deployment is running, your company is on the other end of an MCP connection.
You βΊ What did we decide about checkout payments, and has anything changed since?
Spark βΊ Checkout payments go through Stripe Connect [Checkout Payments Decision].
The decision dates from February; the March migration PRs moved the
webhook handlers to the extensions API [Checkout Extensions Migration].
No later evidence reverses it.Illustrative exchange with a fictional org. spark_ask answers only from your org's knowledge pages, cites them inline as [title], and says what is missing instead of guessing.
1. Get an org API key (after SETUP β prints the key once):
npx convex run debug:yellow2aCreateApiKey '{"orgId":"<orgId>","label":"my-laptop"}'2. Point your client at https://<deployment>.convex.site/mcp (Streamable HTTP, bearer auth):
claude mcp add --transport http spark https://<deployment>.convex.site/mcp \
--header "Authorization: Bearer <org-api-key>"{
"mcpServers": {
"spark": {
"url": "https://<deployment>.convex.site/mcp",
"headers": { "Authorization": "Bearer <org-api-key>" }
}
}
}{
"mcpServers": {
"spark": {
"command": "npx",
"args": ["-y", "mcp-remote", "https://<deployment>.convex.site/mcp",
"--header", "Authorization:${SPARK_AUTH}"],
"env": { "SPARK_AUTH": "Bearer <org-api-key>" }
}
}
}The same endpoint also accepts a simple {tool, input} JSON body, which is handy for scripts and smoke tests:
curl -s https://<deployment>.convex.site/mcp \
-H "Authorization: Bearer <org-api-key>" -H "content-type: application/json" \
-d '{"tool":"spark_search","input":{"query":"checkout payments"}}'3. Use the four tools:
Tool | Input | What you get |
|
| Top knowledge pages from hybrid search (BM25 + vector + rank fusion + optional rerank), with Compiled Truth and highlights |
|
| A cited answer grounded only in your knowledge pages |
|
| A person or entity profile: Compiled Truth plus recent evidence |
|
| Writes a signal into the same pipeline as every other source |
Keys are stored as SHA-256 hashes, scoped to one org, revocable, and rate-limited per key. Browser origins other than localhost are denied unless you list them in MCP_ALLOWED_ORIGINS; desktop and CLI clients are not affected.
β What ships today
Everything below is in main and pinned by the proof gate:
Pillar | What it does | Where it lives |
π₯ Ingestion | One |
|
π Knowledge | Sessions and pasted/URL sources β markdown-aware chunking β AI context prefixes β single-call extraction β 9 page types with versions, attributions and typed cross-references. Live BM25 search plus hybrid deep search. |
|
π Observatory | The intelligence that comes to you. Daily β weekly β monthly synthesis crons; the overnight Dream Cycle, a Managed Agent that writes up to 7 findings per run across 7 types (contradiction, strategic gap, resource drift, emerging capability, stale decision, unanswered question, blind spot); and Pulse, a four-question briefing (what changed / why / so what / now what) cached for 30 minutes. |
|
π€ Agents | Plain-English agents with a full lifecycle, exactly-once scheduler claims, monthly budget caps enforced before a session starts, starter templates and versioned configs with rollback. |
|
π¬ Spark chat | Streaming chat agent with 11 tools (hybrid search, typed live-data queries, approval-gated writes, |
|
π MCP server |
|
|
π Usage accounting | One row per AI call across every surface β the base for budget caps. |
|
π§ Not built yet
Honest edges, so you know what you are adopting:
Not in this repo | Status |
A web UI | An optional official UI is planned as a separate repo. Until then: MCP clients plus |
Full admin and operation over MCP | Reading and asking work over MCP today. Pulse, finding status and agent management do not have MCP tools yet. |
One-command or agent-driven setup | Planned (setup manifest + admin CLI). Today, setup is SETUP.md by hand. Vendor accounts and OAuth consent stay human steps. |
Metered billing | The usage table records every AI call; billing on top of it is future work. |
New Composio toolkits by config alone | Slack, Asana and Gorgias are wired. Each new toolkit is a ~half-day code checklist (SETUP Β§5). |
A credentialed end-to-end proof in CI | CI runs the offline proof gate only. Run the smoke test on your own deployment to prove the live loop. |
Dream Cycle auto-resolve and push notifications stay conservative on purpose until they earn trust.
β‘οΈ Quick start
Verify the code first β about 2 minutes, no accounts, no keys:
git clone https://github.com/joshdayorg/spark.git
cd spark
npm install
npm test # proof gate: 72 verification scripts + behavioral tests
npm run typecheck:offline # strict TypeScript over convex/Then run it on your own accounts β about 45β60 minutes, mostly in vendor dashboards:
npx convex dev --configure # creates your Convex project; Convex IS the runtimeFollow SETUP.md end to end: Clerk (auth + orgs) β Composio (one project, one webhook subscription) β GitHub App β model keys β MCP. Every environment variable the backend reads is listed in .env.example; a proof script keeps that file matched to the code.
You will need Node 20+ (the repo pins 24 in .node-version) and accounts with Convex, Clerk, Anthropic, OpenAI and Composio. Cohere is optional. Customer agents and the Dream Cycle also need the Anthropic Managed Agents beta on your key; chat, extraction and synthesis do not.
The smoke test in SETUP Β§9 proves the whole loop on your deployment: push a signal β it becomes an event β a session β a knowledge page β a search hit. The first extraction runs within 5 minutes.
ποΈ Architecture
The bird's-eye map at the top of this page is the canonical view: LISTEN (signal pipes into one ingest gateway), THINK + REFLECT (extractor, synthesizer, Dream Cycle around the knowledge base), SPEAK (chat, pulse, knowledge, agents, findings, MCP). The system diagram below shows the same machine with its runtimes β it renders live from docs/diagrams/system-diagram.mmd.
Both diagrams include the web surfaces (Chat, Pulse, Knowledge, Agents, Findings, Settings). Theirbackends ship here; the web UI does not. Headless, an MCP client and npx convex run fill those roles.
flowchart TB
subgraph LEGEND["Legend"]
LDIR["Direct Anthropic API call"]
LOUT["AI-derived artifact\n(created by AI, stored/served here)"]
LSYS["Non-AI system / plumbing"]
end
subgraph UI["User Surfaces (Vercel + Convex subscriptions)"]
CHAT["Chat"]
PULSE["Pulse"]
KNOW["Knowledge"]
AGENTS_UI["Agents"]
FVIEW["Findings"]
SETTINGS["Settings"]
end
subgraph CTRL["Spark Control Plane (Convex + Clerk)"]
INGEST["Unified Ingestion Gateway\nConvex HTTP Action"]
EVENTS[("events table\nnormalized + org scoped")]
SESS[("sessions / live context")]
KB[("Knowledge Base\nCompiled Truth + Evidence Timeline\nBM25 + Vector + RRF")]
FIND[("Findings records")]
ORG[("Org / Team / Connection config")]
CRON["Smart fan-out scheduler"]
CHAT_ENGINE["Spark Chat runtime\nConvex action + tools"]
EXTRACT["Extractor job"]
SYNTH["Synthesizer job"]
CUST_WRAP["Customer agent wrapper"]
DREAM_WRAP["Dream Cycle wrapper"]
MCP["MCP Server\nspark_search Β· spark_ask Β· spark_entity Β· spark_push"]
end
subgraph ANTH["Anthropic APIs (execution runtime)"]
MSG["Messages API\nChat + Extractor + Synthesizer"]
MA["Managed Agents API\nCustomer Agents + Dream Cycle"]
end
subgraph EXT["External Systems"]
WH["Universal Webhooks"]
COMP["Composio Integration Layer\n1000+ Connectors"]
PUSH["Direct API Push"]
XAGENT["External Agents / Apps"]
end
WH --> INGEST
COMP --> INGEST
PUSH --> INGEST
INGEST --> EVENTS --> SESS
EVENTS --> CRON --> EXTRACT --> MSG --> EXTRACT --> KB
KB --> SYNTH --> MSG --> SYNTH --> KB
CHAT --> CHAT_ENGINE
PULSE --> CHAT_ENGINE
CHAT_ENGINE --> MSG --> CHAT_ENGINE
CHAT_ENGINE --> KB
SESS --> CHAT_ENGINE
AGENTS_UI --> CUST_WRAP --> MA --> CUST_WRAP --> FIND --> KB
KB --> DREAM_WRAP --> MA --> DREAM_WRAP --> FIND
CUST_WRAP --> COMP
DREAM_WRAP --> COMP
KB --> KNOW
FIND --> FVIEW
COMP --> ORG --> SETTINGS
XAGENT --> MCP
MCP --> KB
MCP --> INGEST
MCP -. "spark_ask" .-> MSG
classDef directAI fill:#d8b4fe,stroke:#7e22ce,stroke-width:2px,color:#2b1147;
classDef aiOutput fill:#f3e8ff,stroke:#7e22ce,stroke-width:2px,stroke-dasharray:6 4,color:#2b1147;
classDef system fill:#e8f1ff,stroke:#2563eb,stroke-width:1px;
classDef ui fill:#e8fff0,stroke:#16a34a,stroke-width:1px;
classDef integration fill:#fff7e8,stroke:#f59e0b,stroke-width:1px;
class LDIR directAI;
class LOUT aiOutput;
class LSYS system;
class CHAT_ENGINE,EXTRACT,SYNTH,CUST_WRAP,DREAM_WRAP,MSG,MA,MCP directAI;
class KB,FIND aiOutput;
class INGEST,EVENTS,SESS,ORG,CRON system;
class CHAT,PULSE,KNOW,AGENTS_UI,FVIEW,SETTINGS ui;
class WH,COMP,PUSH,XAGENT integration;The stack: Convex (database, real-time subscriptions, crons, agent runtime, vector RAG) Β· Anthropic (all reasoning β Messages API in the hot path, Managed Agents isolated off it) Β· Composio (OAuth + triggers in, tool delivery out) Β· Clerk (auth, orgs, multi-tenant JWT) Β· OpenAI embeddings Β· Cohere reranking (optional).
Tenancy iron rule: orgId is never a client argument. Public functions derive it from the Clerk JWT; AI tools derive it from the agent thread; every read goes through an org-prefixed index.
flowchart TB
subgraph EXT["External Sources"]
WH["Universal Webhooks\n(native, e.g. GitHub org webhook)"]
COMP["Composio Triggers\n(metadata.user_id = orgId)"]
API["Direct API Push\n(Authorization: Bearer api_key)"]
MCPP["MCP spark_push\n(external agent write)"]
end
INGEST["Unified Ingestion Gateway\nPOST /ingest (Convex HTTP Action)"]
subgraph NORM["Source Detection + Verification + Normalization"]
DETECT["Detect source\n(headers/body)"]
VERIFY["Verify auth\nGitHub sig / Composio HMAC / API key"]
MAP["Map to normalized event shape\norgId, source, via, type, actor, summary, timestamp, raw, processed=false"]
end
EVENTS[("events table\nsource-agnostic")]
subgraph FANOUT["Smart Fan-out Scheduler"]
CRON["Cron every 5 min"]
QUERY["Query orgIds with\nprocessed=false"]
EXTRACT["scheduler.runAfter(0, extractForOrg, {orgId})"]
end
WH --> INGEST
COMP --> INGEST
API --> INGEST
MCPP --> INGEST
INGEST --> DETECT --> VERIFY --> MAP --> EVENTS
EVENTS --> QUERY
CRON --> QUERY --> EXTRACT
classDef edge fill:#fff7e8,stroke:#f59e0b,stroke-width:1px;
classDef data fill:#e8f1ff,stroke:#3b82f6,stroke-width:1px;
classDef ai fill:#f4e8ff,stroke:#a855f7,stroke-width:1px;
class WH,COMP,API,MCPP,INGEST edge;
class EVENTS,MAP data;
class CRON,QUERY,EXTRACT,DETECT,VERIFY ai;Want the full internals story β the knowledge page model, extraction prompt design, dedup/idempotency contracts, search architecture? Read the deep dive.
π Integrations
Source | Pipe | Self-serve? |
GitHub | GitHub App β native webhooks to | β
Admins generate a signed install link; the post-install callback ( |
Slack | Composio trigger | β
Connect Link ( |
Asana | Composio trigger | β Connect Link, same as Slack |
Gorgias | Composio-backed poller (agent-attributed ticket activity; customers never stored as actors) | β Connect, then the poll cron takes over |
Anything else | Bearer-key API push, or MCP | β Create an org API key and POST |
Adding a new Composio toolkit is a documented ~half-day code checklist (descriptor, triggerβtype map, actor extraction) β see SETUP.md Β§5. Outbound, customer agents deliver their own results (Slack messages, tasks, email) through the org's Composio connections via MCP β there is no separate delivery system.
π§ͺ The proof gate
This repo treats claims as liabilities. npm test runs the full gate, offline:
Verification scripts (
scripts/verify-*.sh) β each pins one contract: org isolation, bounded indexed reads, ingest idempotency, budget enforcement before session creation, docs-to-code drift (yes, this README is gated too).Behavioral tests (
convex-test+ vitest) β real function execution against an in-memory backend, including identity-based auth tests of the public write surfaces.
CI runs the same gate plus the offline typecheck on every push and pull request. Nothing in the gate calls a live vendor API; the live loop is proven by the smoke test on your own deployment. npm run typecheck (Convex codegen) needs a configured deployment; npm run typecheck:offline does not.
If you change a gated contract, update the proof in a separate commit that explains why the contract moved. See CONTRIBUTING.md.
ποΈ Repo layout
Path | What lives there |
| Every table, org-scoped, by domain |
| Auth wrappers, chat agent + tools, extraction, synthesis, linter, MCP handlers, managed-agent runtime |
| Source ingestion β chunking β extraction β pages/versions/attributions/references |
| Customer agents (lifecycle, scheduler, budgets) / the findings engine |
| Signal normalizer / HTTP routes ( |
| Usage accounting / chat knowledge tools / public Knowledge + Findings writes |
| Internal operator functions: proof probes, fixtures, repairs and backfills (run with |
| The proof gate |
| Deep dive, diagram sources, flow inventory, event pipeline design, Convex context for contributors and their AI assistants |
π€ Contributing
Contributions are welcome. Start with CONTRIBUTING.md: the proof-gate contract, tenancy and bounded-read rules, and the schema-change policy. Everything you need is in this repo β the code, the deep dive, the diagram sources in docs/diagrams/, and the proof scripts that state each contract.
Bugs and ideas: open an issue. For anything bigger than a fix, describe the change in an issue before you write the PR.
Security: report privately β see SECURITY.md. Please don't open public issues for vulnerabilities.
π License
MIT Β© 2026 Josh Day / WayFX
This server cannot be deployed
Maintenance
Related MCP Connectors
Make your knowledge agent-ready. One MCP endpoint, 5 connectors, 3 search modes.
Cloud or self-hosted knowledge for AI agents: hybrid search, reranking, GraphRAG, scoped MCP tools.
- LensHubOAuthai.lenshub
Team knowledge from 20 connectors, served to any MCP agent β classified, scored, access-controlled.
Knowledge base MCP for AI agents on iknow.dev. Search, read, and maintain via OAuth.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables searching a knowledge base and asking grounded questions with hybrid retrieval, reranking, and cited answers.-
- AlicenseAqualityCmaintenanceA knowledge base MCP server that aggregates team knowledge from multiple sources into Postgres. It provides hybrid search (full-text + vector + RRF) via MCP tools, and enables direct recording of decisions, learnings, and pitfalls.151MIT
- FlicenseNot gradedqualityBmaintenanceEnables to build and query a knowledge base with retrieval-augmented generation, supporting document ingestion, hybrid search, and live data integration from external APIs via MCP tools.2-
- FlicenseNot gradedqualityCmaintenanceEnables AI agents and users to search a unified organizational knowledge warehouse, create and iterate on documents, manage folders and reviews, and curate shareable knowledge packs through any compatible MCP client.-