SiteGraph
Resolves DOIs via doi.org to verify paper metadata, retraction status, and journal integrity as part of IntegrityLens.
Integrates with Google Fact Check Tools API to check if a claim or page has been fact-checked, returning ClaimReview data for ClaimTrail.
Integrates with Meta Ad Library to search for ads run on Meta platforms (Facebook, Instagram, etc.) as part of DSA-Lens transparency research.
Integrates with the npm registry to verify package existence, safety, and metadata for dependency verification in DepCheck.
Integrates with the PyPI registry to verify package existence, safety, and metadata for dependency verification in DepCheck.
Integrates with Trend Micro's slopsquatting dataset to detect potentially malicious or hallucinated package names in DepCheck.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@SiteGraphwho operates nytimes.com and what other sites share its operator?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
agent-osint
Seven organisation-focused OSINT servers for AI agents, speaking the Model Context Protocol. Attach them to Claude, Cursor, or any MCP client and your agent gains verifiable, evidence-backed answers to questions LLMs cannot answer from memory.
Server | Question it answers | Key sources | Needs a key? |
SiteGraph | Who operates this website, and which other sites share its operator? | ads.txt / app-ads.txt, sellers.json, analytics & tag IDs, legal notices (Impressum), security.txt, Certificate Transparency, Wayback, Common Crawl | No (optional |
CommentForensics | Which public comments on a US rulemaking are coordinated form-letter campaigns, and who is behind them? | Regulations.gov API v4, Mirrulations mirror, bulk CSV exports | Optional ( |
IntegrityLens | Is this paper/reference list/journal trustworthy? | Crossref (incl. Retraction Watch data), OpenAlex, doi.org, PubPeer, Problematic Paper Screener & hijacked-journal lists (you load them), tortured-phrase detector | Optional ( |
DSA-Lens | How do platforms moderate content in the EU? What ads ran? | EU DSA Transparency Database (Research API + daily dumps), Meta Ad Library |
|
Procurement Red Flags | Which public contracts show integrity risk indicators? | OCDS data (UK Find a Tender or any publisher), GLEIF, Companies House, OpenSanctions | Optional ( |
DepCheck | Is the package my coding agent wants to install real and safe? | PyPI, npm, popularity lists, Trend Micro slopsquatting dataset | No |
ClaimTrail | When and where did this claim/page/image first appear? Has it been fact-checked? | Wayback Machine, page metadata, ClaimReview, Google Fact Check Tools, C2PA | Optional ( |
48 tools in total — see docs/TOOLS.md. Verification status is documented honestly in VERIFICATION.md.
Design principles
Evidence on every claim. Each finding carries evidence IDs (URL, fetch time, HTTP status, SHA-256 of the raw bytes, and the bytes where lawful to keep). Read them via the
<server>://evidence/{id}MCP resource.Uniform envelope. Every tool returns
{"ok", "data", "evidence", "warnings", "limitations", "sources_checked"}or{"ok": false, "error": {"code", "message", "hint"}}— agents never see stack traces.Organisations, not people. Names that look like private individuals (sole traders, commenters, signatories, PubPeer users, ad payers) are withheld; personal emails and phone numbers are masked. There is no switch to disable this.
Polite by default. Conservative per-host token-bucket rate limits (matching providers' published limits where they publish them), retries with
Retry-After, robots.txt (RFC 9309) for HTML pages, response-size caps, and a local response cache.Safe to expose to an agent. Requests to loopback/private/link-local addresses are refused (SSRF guard), local file inputs can be confined with
AGENT_OSINT_FILE_ROOTS, and file types are checked.Honest outputs. Scores are tiers with stated rules, "no signal" is never presented as "clean", and every tool lists its limitations.
Related MCP server: BizIntel MCP
Install
Requires Python 3.10+.
pip install ./agent-osint # from this folder
pip install "./agent-osint[c2pa]" # + C2PA reading for ClaimTrailOr run without installing, with uv: uvx --from ./agent-osint agent-osint-depcheck.
Docker: docker build -t agent-osint . then docker run -i --rm agent-osint agent-osint-depcheck.
Connect to your AI
Claude Desktop — edit claude_desktop_config.json (Windows: %APPDATA%\Claude\, macOS:
~/Library/Application Support/Claude/):
{
"mcpServers": {
"depcheck": { "command": "agent-osint-depcheck" },
"sitegraph": { "command": "agent-osint-sitegraph", "env": { "AGENT_OSINT_CONTACT": "you@example.org" } },
"integritylens": { "command": "agent-osint-integritylens", "env": { "AGENT_OSINT_CONTACT": "you@example.org" } },
"commentforensics":{ "command": "agent-osint-commentforensics", "env": { "REGULATIONS_GOV_API_KEY": "..." } },
"dsalens": { "command": "agent-osint-dsalens", "env": { "DSA_TDB_TOKEN": "...", "META_AD_LIBRARY_TOKEN": "..." } },
"procurement": { "command": "agent-osint-procurement", "env": { "COMPANIES_HOUSE_API_KEY": "..." } },
"claimtrail": { "command": "agent-osint-claimtrail", "env": { "GOOGLE_FACTCHECK_API_KEY": "..." } }
}
}Claude Code — claude mcp add depcheck -- agent-osint-depcheck (repeat per server; add -e KEY=value for env).
Cursor / other clients — any MCP client that launches stdio servers works with the same commands. For remote use,
run agent-osint-<server> serve --transport streamable-http.
Configuration
Variable | Purpose |
| Data directory (default |
| Your email, sent in the User-Agent and to Crossref/OpenAlex "polite pools" (recommended) |
|
|
|
|
| Override rate limits, e.g. |
| Cache lifetime (s), timeout (s), response cap (bytes) |
| Record live responses as fixtures / replay them offline |
| Extra tortured-phrase list ( |
| Free key from api.data.gov |
| DSA Transparency Database Research API token (EU Login + request to the DSA helpdesk) |
| Meta Ad Library access token / Graph API version (default |
| PubPeer API key (request from PubPeer) |
| Google Cloud key with the Fact Check Tools API enabled |
| UK Companies House REST API key |
| OpenSanctions hosted API key, or the URL of your self-hosted yente |
| SSLMate Cert Spotter key for higher CT query limits |
A template is in .env.example.
Command-line utilities
agent-osint-depcheck check requirements.txt # CI gate: exit 1 if any dependency is 'block'
agent-osint-sitegraph index domains.txt # bulk-fingerprint a domain list into the index
agent-osint-sitegraph ingest-warc CC-MAIN-*.warc.gz # seed the sibling index from Common Crawl WARC files
agent-osint-commentforensics ingest EPA-HQ-OAR-2021-0317 --source mirrulations
agent-osint-commentforensics ingest-csv export.csv # Regulations.gov bulk download
agent-osint-dsalens ingest sor-global-2026-09-01-light.zip --platform TikTok
agent-osint-integritylens load retraction_watch retraction_watch.csv
agent-osint-procurement load https://example.org/ocds/release-package.json
python scripts/live_smoke.py # live end-to-end check of every server
python scripts/slopsquat_benchmark.py # re-check LLM-hallucinated package names on PyPIDevelopment
pip install -e ".[dev,c2pa]"
pytest # offline suite (replay fixtures, real MCP stdio sessions)
pytest -m live # live registry tests (needs internet)
ruff check src testsResponsible use
These tools surface public information about organisations. Shared identifiers, form-letter campaigns and procurement red flags are indicators for further review, not proof of ownership, fraud or corruption. Verify before you publish, give organisations a chance to respond, and respect each data source's terms of use (see NOTICE and VERIFICATION.md).
License
Apache License 2.0. Third-party data and test fixtures are listed in NOTICE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Evidence-backed x402 web verification for AI agents, with auditable decisions for every condition.
Live web checks for AI agents: sitemaps, robots.txt, URL status, broken links, feeds, citations.
Scan any public site for AI-agent visibility; get scored findings, a machine-readable fix pack, and
Website facts for agents: HTTPS/certificate checks, public site info, small local-business search.
31
Related MCP Servers
- AlicenseAqualityBmaintenanceProvides tools for AI agents to audit websites, including stack detection, DNS snapshots, and security checks.817 npmMIT
- AlicenseNot gradedqualityDmaintenanceProvides real-time website audits, lead scoring, tech-stack detection, and local-business search for AI agents doing sales outreach and competitor research.MIT
- AlicenseAqualityAmaintenanceEnables AI agents to investigate corporate ownership, trace ultimate beneficial owners, screen sanctions, detect offshore exposure, and access fully cited dossiers from 130M+ entities across 31 global registries.2196 npm1MIT
- FlicenseNot gradedqualityCmaintenanceEnables AI agents to investigate any website's technology stack, identifying CMS, ecommerce platforms, JavaScript frameworks, analytics, CRM, marketing automation, payments, chat, CDN, and hosting.-