ThoughtDAG
ThoughtDAG
AI conversations that branch on an infinite canvas.
Each exchange becomes a node. Wires are the context. Explore a side question, connect useful paths, and choose what the model sees next.
Download · Website · Docs · 中文
0.5 update · CLI · Harness · Desktop · How it works · How it differs · Research
New in 0.5 · ThoughtDAG × Jev
Bring relevant past conversations into the question you are asking now.
Find earlier work. The local index searches supported agent sessions and ThoughtDAG canvases. Topic dossiers collect decisions and open questions with links back to their sources.
Select what belongs. The optional Jev decision layer helps identify topics and rank relevant excerpts. Your chosen language model develops the answer.
Check what comes back. With recall enabled, the context panel lists the dossiers and excerpts added to a request. Inspect their sources or exclude individual items before continuing.
In a small relevance-selection pilot, Jev's median was 391 ms versus 24,813 ms for our GLM adapter. These are selection-stage timings, not end-to-end search or answer times.
Six runs per engine over the same 14 synthetic excerpts. Median selection latency: 391 ms for Jev-1.13 and 24,813 ms for the GLM-5.3-Flash adapter with default reasoning. These are different inference paths, not a controlled ranking of model speed. Retrieval and answer generation are excluded; this does not measure whole-product speed or accuracy gains.
Without a decision model, recall falls back to rules. The System 1 / System 2-style split describes software roles here: quick relevance decisions, then answer and dossier generation. It is not a claim about human cognition.
Set up history and recall · Configure Jev
Related MCP server: session-recall
Find past context from the command line
Remember a file, a phrase or a URL, but not the session? Search local conversations and jump to the matching turn, without opening the desktop app.
npx thoughtdag why src/lib/api.ts # conversations about this file
npx thoughtdag find "a phrase you remember" # matching conversation turns
npx thoughtdag topics # topics in your local indexFor regular use: npm install -g thoughtdag. Run thoughtdag setup mcp to expose read-only history tools to your agent. Retrieve the relevant turns rather than replaying a whole session. CLI guide →
Inside DeepSeek Harness
Switch between chat and ThoughtDAG's graph inside the harness. Choose the context on the canvas; the harness runs the next turn.
dsh plugin --profile web add dsh-thoughtdag
dsh webThe plugin bundles the canvas and memory layer. Requires Node 22.19+ (22.x) or 24+, and DeepSeek Harness 0.1.2-rc.1 or later. Plugin guide →
The desktop app
Read a document beside your conversation, branch from a passage, and connect the paths you want to explore together. Use your own model connection.
brew install --cask thoughtdagOr download for macOS, Windows or Linux, connect a model and open the example canvas.
The one rule
Wires are the context. Connect conversation paths to use them in the next question. Disconnect a path without deleting the work.
Branch from a detail, explore it separately, then connect the useful parts to a later question. The graph changes the model's input, not just the layout.
Preview what the model will receive before sending. Wires select the conversation paths; explicit references and enabled recall can add material alongside them. Context guide →
In action
✂️ Change the context, keep the exploration
Select text in an answer to start a side branch. Disconnect that branch from a later question, then regenerate to compare. Its nodes stay on the canvas: keep exploring from them or reconnect them later.
📖 Read, clip, and ask
Open a PDF, image or HTML alongside the graph. Ask about a passage or clip a figure into its own node. PDF clips keep their page reference, so you can check the source as the discussion develops.
💎 Condense the path; weave the highlights
Condense creates a shorter copy of a conversation path while preserving the original. Weave turns selected highlights into cited prose. Continue from the result, or export it as Markdown. Zooming out changes the view, not the context.
🧭 Session Atlas: continue an earlier conversation
Open a supported local agent session as a graph. Pick where to branch or continue; use the history index to find related discussions from other sessions. Atlas provides the view, and recall helps find what to bring in.
Supports local Claude Code, Codex, DeepSeek Harness and Pi sessions. Source sessions remain read-only.
How ThoughtDAG differs
Nodes and edges serve different purposes. Here is where ThoughtDAG fits:
Product category | ThoughtDAG's focus |
Linear chat | Keep several lines of inquiry visible and choose which ones continue into the next question. |
Mind maps and whiteboards | Use connections to change model input, not just organize ideas visually. |
Branching chat canvases | Connect several branches into one question, or disconnect a path while keeping its nodes. |
Agent workflow canvases | Edit conversational context as you explore, rather than design a pipeline of automated tasks. |
Retrieval and automatic memory | Inspect source-linked dossiers and recalled excerpts; edit or exclude what the next request uses. |
Code graphs and conversation search | Find the discussions behind a file or topic across supported agents, then continue from them. |
Harness context viewers | Move from inspecting a session to composing and sending its next turn. |
These categories overlap; individual tools may share capabilities. ThoughtDAG is not an autonomous research agent or a replacement for your coding harness. Retrieval can miss relevant history, and generated dossiers still need checking.
🗺️ Export the shape of your thinking
Export the canvas as a Thought Map: nodes, wires and structural counts, without the full conversation text. Use it to share how an investigation branched, narrowed and came together.
More ways to run
Run from source
npm install
npm run server # LLM proxy :3001
npm run dev # frontend :5173Configure a model in the app or through environment variables. Local setup →
Browser demo
The browser demo includes an example canvas that needs no API key. It is a subset: local session discovery, Session Atlas and the local history/memory layer require desktop or local hosting.
🧪 Research: Why editable context matters
Context Intervention Benchmark · Pilot v2
9 model endpoints · 1,485 scored responses · exact-match scoring
Deleting a wrong claim may leave its consequences in later replies. In our synthetic pilot, removing the source alone repaired 152 of 162 affected model-cases; removing the contaminated subgraph repaired 162, and recomputing descendants repaired 161. The report includes the protocol, results and limitations. This is a context-intervention experiment, not a general model leaderboard.
Read the case study · Methods and results · Suggest a model
More capabilities
Capability | What it adds to the same workflow |
Request preview | Check the conversation, references and recalled material assembled for the next call. |
Staleness and replay | Review dependent answers after an upstream edit; rerun in dependency order. |
Per-node model selection | Try a different model on a branch without changing the entire canvas. |
Read-only sharing | Share a graph for others to inspect; preview its contents before publishing. |
Folder backup | Save canvases as local files and keep a recoverable copy outside browser storage. |
Full capabilities and roadmap →
Models, cost & privacy
Canvases, documents, the index and dossiers are stored locally. Remote model calls send relevant content to your configured providers, including decision and dossier-generation calls. Provider charges may apply; disabling Jev does not disable ordinary model calls.
Connect local Ollama or an OpenAI-compatible endpoint. Inside DeepSeek Harness, inference uses the harness's providers and keys. Export backups and Markdown, and review text and metadata before sharing. Setup and privacy details →
Contributors
Contributions are welcome — start with CONTRIBUTING.md.
Supporters
With gratitude to @andreilaiter, ThoughtDAG's first supporter, and to everyone helping this independent open-source project grow.
Available Tools
4 toolsfindC
Where these exact words were asked (Q), answered (A) or attached (M) across local sessions, canvases and the memories Claude Code and Codex keep for themselves. Exact, case-insensitive match; every hit is a verbatim snippet with a pointer.
| Name | Required | Description | Default |
|---|---|---|---|
| in | No | ||
| limit | No | ||
| phrase | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that matching is exact and case-insensitive (not semantic/fuzzy) and that results are verbatim snippets with pointers, but says nothing about result limits, permissions, or what happens on no-match.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with little waste, but the lead phrase 'Where these exact words were asked' is grammatically awkward and front-loads a clause rather than the action or resource, forcing the reader to reconstruct the purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param, one-required tool with no annotations and no output schema, the description covers matching semantics and return shape reasonably, but omits the meaning of 'limit' and any routing guidance relative to its siblings, leaving gaps an agent would need to guess at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does decode the 'in' enum values (asked/answered/attached) which is genuinely additive, but 'limit' is completely unexplained and 'phrase' is only obliquely referenced as 'these exact words'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description scopes the corpus (local sessions, canvases, Claude Code/Codex memories) and the operation (exact, case-insensitive matching), but never states a clear leading verb and never distinguishes itself from siblings like recall_turn or why_file. An agent gets the search domain but must infer that this is a literal full-text 'find' rather than a semantic lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the sibling tools recall_turn, why_check, or why_file, and no stated prerequisites or exclusions. The Q/A/M breakdown hints at scope but is not framed as selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recall_turnA
One turn in full — the question, the answer, the tool calls with their diffs — by session id (or prefix) and turn number as shown by why_file.
| Name | Required | Description | Default |
|---|---|---|---|
| turn | Yes | ||
| session | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the return contents (question, answer, tool calls with diffs) and that session IDs can be prefixes, which is helpful. But it does not state that the operation is read-only, whether there are side effects, or what happens on invalid input, leaving key behavioral traits implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, densely packed sentence with zero wasted words. It front-loads the tool's output ('One turn in full') and then clarifies the parameters, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter retrieval tool with no output schema and no annotations, the description is fairly complete: it explains the return values and how to source the turn number. It omits edge-case behavior (e.g., missing session or turn) and does not state read-only status, but those are minor gaps given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds that 'session' can be a prefix and that 'turn' is a number obtained from why_file, which is useful. However, it does not explain the format of session (e.g., string pattern) or turn (e.g., integer range), so semantics remain partially specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource ('One turn in full') and enumerates its contents (question, answer, tool calls with diffs), making the tool's function clear. It references the sibling 'why_file' for turn numbers, but does not explicitly contrast itself with 'find' or 'why_check', so sibling differentiation is only partial.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: the turn number is 'as shown by why_file', suggesting a workflow where why_file is consulted first. However, there is no explicit when-to-use or when-not-to-use guidance, nor are alternatives like 'find' or 'why_check' named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
why_checkA
Cheap first question before editing a file: does this artifact have any history in local agent sessions? One line; history true/false.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | file path (absolute or relative to cwd), URL, or arxiv:<id> |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses cost ("cheap"), output shape ("one line; history true/false"), and scope (local agent sessions), but says nothing about permissions, latency beyond cost, or what qualifies as "history," leaving meaningful behavioral gaps for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight clauses, front-loading the usage trigger and then the cost/return profile. No filler sentences and nothing buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly supplies the return format (one line, boolean) and the cost profile, which is sufficient for a trivial query tool. The only gap is not routing against the similar-sounding why_file sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single path parameter is fully documented in the schema (file path, URL, or arxiv:<id>). The description adds no meaning beyond "this artifact," so the schema does the heavy lifting and baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: it checks whether an artifact has history in local agent sessions, and clarifies the return is a one-line true/false. It is clear, but it never distinguishes itself from the sibling why_file, which sounds closely related, so an agent could still hesitate between them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Cheap first question before editing a file" gives a concrete when-to-use trigger and implies a cost-based ordering of calls. It does not name alternatives (e.g., why_file) or state when-not to use it, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
why_fileC
The turns across local Claude Code, Codex, DeepSeek Harness, Pi and ThoughtDAG sessions that touched a file, URL or paper: when, what changed (Δ, observed), what was asked, what the answer said about it (≈, a candidate explanation, not a verified reason). Each hit carries a deep link.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| limit | No | max hits (default 10) | |
| include_read | No | also list turns that only read it (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does disclose useful semantics: the span of sources covered (local Claude Code, Codex, DeepSeek Harness, Pi, ThoughtDAG) and the epistemic caveat that '≈' is a candidate explanation, not a verified reason. It does not disclose read-only vs mutating behavior, auth requirements, result ordering, or volume/pagination behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense clause with no filler; each element (when, Δ, asked, answer, deep link) earns its place. It is slightly awkward to parse because it is a fragment with symbolic notation (Δ, ≈) and no leading verb, which costs a bit of front-loading.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description responsibly sketches the return contents (timing, change, prompt, candidate explanation, deep link). Missing for a 3-param tool with zero annotations are the mutation/safety profile and any hint about result volume or the meaning of the defaulted `limit`.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: limit and include_read are self-documented in the schema, while the required `path` has no description. The description partially compensates by revealing that the target can be a 'file, URL or paper', which is non-obvious from the parameter name `path`. That is real added meaning, but it does not cover the remaining ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a noun phrase describing the payload ('The turns ... that touched a file, URL or paper') rather than a verb-led statement of what the tool does. The intended action (find/list turns that touched an artifact, with explanation) is inferable but never stated outright, and there is no differentiation from siblings like recall_turn or why_check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no statement of prerequisites, and no mention of how this differs from find, recall_turn, or why_check. The agent must guess the routing among four siblings purely from their names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
find - First observed
recall_turn - First observed
why_check - First observed
why_file
TDQS
Scored across 4 tools
The four tools separate along clear axes: find does keyword search, why_file retrieves turns touching a specific artifact, why_check gives a quick boolean for artifacts, and recall_turn fetches a specific turn. why_check is a lightweight version of why_file and find overlaps with why_file for file-related queries, but the descriptions make intended usage clear.
All names use lowercase snake_case, but the patterns are mixed: find is a bare verb, recall_turn is verb_noun, while why_check and why_file use a why_ prefix. The inconsistency is readable but not a uniform verb_noun convention.
With only four tools, the set is tightly scoped for a session-history retrieval server, and each tool serves a distinct retrieval depth (keyword search, artifact existence check, artifact history, full turn). No obvious filler tools are present.
The surface covers keyword search, artifact-based history, existence checks, and full-turn retrieval, which supports the core recall workflows. It lacks session enumeration or time-range filtering, but agents can work around these gaps by searching known terms.
Related MCP Connectors
Search ATProto writing, annotations, identity, agents, and forum posts. 12 read-only tools.
Search and save your coding work in a CoralSwarm ocean: sessions, decisions, meetings.
Search your AI chat history (ChatGPT, Claude, Codex) from any MCP client. Remote, private, read-only
Cross-agent artifact workspace with provenance across Claude Code, Codex, Cursor, LangGraph.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceLocal memory search for Codex and Claude Code conversations. It keeps history on your machine, builds a local graph index, and returns compact evidence from past sessions.6MIT
- AlicenseBqualityCmaintenanceProvides local, agentic semantic recall over Claude Code session history, enabling the agent to search past discussions semantically, expand turns, and grep transcripts.512MIT
- AlicenseAqualityDmaintenanceSearches and browses local Claude Code and Cowork session histories stored on your machine, enabling questions about past work without uploading data.32MIT
- AlicenseAqualityBmaintenanceSearch past OpenCode conversation history before starting new work on a module or file, via a local read-only FTS5 index built from OpenCode's own SQLite database. No network calls, fully local. 7 tools for keyword search, file lookup, and session browsing.74MIT