Skip to main content
Glama
Ryu07-d

knowledge-curator-mcp

by Ryu07-d
README.md
# knowledge-curator-mcp

A **local, zero-cost** MCP server that fact-checks your Markdown / Obsidian notes.

It uses a **local LLM via [Ollama](https://ollama.com)** to extract claims and judge
them against **free, key-less sources** (Wikipedia + DuckDuckGo). No cloud API, no
API keys, nothing leaves your machine except the public search queries.

Unlike naive fact-checkers that count search hits, this server has the LLM **read each
source and decide whether it supports, contradicts, or is insufficient** for the claim.

## Why local?

- **No cost** — runs entirely on your own machine.
- **Private** — your notes are never sent to a cloud LLM.
- **Good enough** — a small instruct model (3–7B) is plenty for "does this evidence
  support this sentence?" entailment judgments.

## Requirements

- Node.js ≥ 18 (uses the global `fetch`)
- [Ollama](https://ollama.com) running locally with an instruct model pulled:

```bash
ollama pull qwen2.5:3b   # default — ~1.9GB, strong instruction-following
# alternatives: qwen3.5:4b (better), qwen2.5:7b (best, ~4.7GB)
```

> A general **instruct** model is recommended over a persona/style fine-tune:
> fact-checking needs neutral, accurate reading, not a personality.

## Install & build

```bash
npm install
npm run build
```

## Configure your MCP client

Add to your Claude Desktop config (`claude_desktop_config.json`):

```json
{
  "mcpServers": {
    "knowledge-curator": {
      "command": "node",
      "args": ["/absolute/path/to/knowledge-curator-mcp/build/index.js"],
      "env": {
        "OLLAMA_MODEL": "qwen2.5:3b"
      }
    }
  }
}
```

## Environment variables

| Variable       | Default                   | Description                                  |
| -------------- | ------------------------- | -------------------------------------------- |
| `OLLAMA_HOST`  | `http://localhost:11434`  | Ollama server URL                            |
| `OLLAMA_MODEL` | `qwen2.5:3b`              | Any installed Ollama model                   |
| `WIKI_LANG`    | `ja`                      | Wikipedia language edition (`en`, `ja`, ...) |

## Tools

| Tool                    | What it does                                                           |
| ----------------------- | --------------------------------------------------------------------- |
| `verify_claim`          | Verify a single statement; returns a verdict + citations.             |
| `fact_check_document`   | Extract & verify claims in a file; optional `auto_fix` inserts notes. |
| `add_citations`         | Fact-check a file and insert footnote citations for claims that need them. |
| `scan_vault_for_issues` | Walk an Obsidian vault and report files with contradicted/uncited claims. |
| `git_commit_corrections`| Commit corrected files to git.                                        |

### Verdicts

| Verdict          | Meaning                                                       |
| ---------------- | ------------------------------------------------------------- |
| ✅ `verified`     | Evidence clearly supports the claim.                          |
| ❌ `incorrect`    | Evidence clearly contradicts it (a correction is suggested).  |
| 📝 `needs_citation` | Plausible & on-topic, but evidence doesn't directly confirm. |
| ❓ `unverifiable` | Evidence unrelated or insufficient.                           |

## Example

> "Verify: 日本の首都は大阪である。"

```
**Verdict**: ❌ incorrect (90% confidence)
**Reasoning**: Evidence clearly contradicts the claim that '日本の首都は大阪である'.
**Sources**: Wikipedia: 大阪市, Wikipedia: 首都圏 (日本), ...
**Suggested correction**: 日本の首都は東京である。
```

## Case study

See [docs/case-study.md](docs/case-study.md) for an end-to-end walkthrough: ingesting
PDFs into Markdown locally, then fact-checking the claims — including the example above
where a naive hit-counter would be fooled but the LLM catches the error.

## How it works

```
note.md ──▶ [LLM extracts claims] ──▶ for each claim:
                                         ├─ Wikipedia search (intro extracts)
                                         ├─ DuckDuckGo instant answer
                                         └─ [LLM reads evidence → verdict + citation]
```

## Limitations

- Quality scales with the model. A 3B model occasionally produces a sloppy
  rationale; use `qwen3.5:4b`/`qwen2.5:7b` for tougher material.
- Free sources are shallow: Wikipedia covers general/encyclopedic facts well, but
  niche or very recent claims will often come back `unverifiable`.
- `auto_fix` edits files in place — keep your notes under version control.

## License

MIT

TDQS

B3.4/5.0

Scored across 5 tools

Disambiguation4/5

Tools are mostly distinct but add_citations and fact_check_document both operate on files and could be confused; one inserts citations while the other reports. scan_vault_for_issues is a batch version of fact_check_document, adding slight overlap.

Naming Consistency3/5

Most tools follow verb_noun (add_citations, fact_check_document, verify_claim) but git_commit_corrections and scan_vault_for_issues break the pattern with prepositional phrases or compound nouns, creating mild inconsistency.

Tool Count5/5

With 5 tools, the set is well-scoped for a knowledge curator. Each tool serves a clear purpose without redundancy, and the count feels neither sparse nor overwhelming.

Completeness4/5

The tools cover the essential workflow: verify claims, fact-check documents, add citations, scan vaults, and commit. Minor gaps exist (e.g., no tool for reverting changes or editing citations manually), but core operations are present.

Maintenance

ActivityInactive
ResponsivenessNo issues