Skip to main content
Glama

HKEx Filing Scraper

HKEx Filing Scraper — one scraper, many databases

CI GitHub Release PyPI License: MIT Python 3.10+ Docs MCP Glama MCP Ruff

An open-source scraper for 25+ years of HKEx regulatory filings — into any of nine databases, with full-text extraction, graph linking, and a read-only MCP server for AI agents.

The US has EDGAR full-text search. Japan has EDINET. Hong Kong has a search form that returns one page at a time. There is no bulk, machine-readable, full-text corpus of HKEx filings. This builds one.

An open-source Python tool that scrapes 25+ years of Hong Kong Stock Exchange (HKEx) regulatory filings and ingests them into any combination of nine databases — with full-text and table extraction, chunk-level coverage, optional graph linking, and a read-only MCP server so AI agents can query the corpus or the live site.

It speaks the undocumented HKEx JSON API directly, which is faster and more resilient than driving a browser.

Vendors & Integrations

Databases — nine first-class destinations, in documented popularity order (see the support matrix):

  • PostgreSQL — production-grade open-source relational

  • MySQL / MariaDB — GPL relational servers, one driver

  • SQLite — zero-server file database, no install needed

  • MongoDB — document database

  • Neo4j — property-graph database

  • ClickHouse — columnar analytics engine

  • DuckDB — in-process analytical engine

  • SurrealDB — multi-model graph + document database

AI clients — any MCP-capable agent; ready-made configuration for Claude, ChatGPT, Cursor, VS Code/Copilot, Gemini CLI, opencode, Manus, and Perplexity.

Available on — PyPI · Glama · MCP Registry · hosted gateway.

Related MCP server: internet-context-mcp

Two Ways to Use It

Hosted MCP gateway

Local pipeline

What

A public endpoint you point an AI agent at

The hkex-scraper CLI

Setup

None — paste a URL

pip install + one environment variable

Data

Live from HKEx, nothing stored

Stored in your database(s)

Docs

Live MCP gateway · AI agent support

Getting started

Example: install, scrape filings into SQLite, then query the hosted MCP gateway from an AI agent

Use the Hosted MCP Gateway

POST, Streamable HTTP, no API key:

https://hkex-listco-updates.ascent-partners.com/api/mcp

Four read-only tools: get_server_info, search_filings (a window of at most 31 days, with optional stock-code, title, document-type, category, and stock-name filters), list_filing_facets (browse what a window contains), and get_filing (downloads one document and extracts its text and tables).

Two ways to reach HKEx filings from an AI agent: the hosted MCP gateway or the local stdio server

Point a client at it — for example opencode:

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "hkex-live": {
      "type": "remote",
      "url": "https://hkex-listco-updates.ascent-partners.com/api/mcp"
    }
  }
}

Then ask:

Use hkex-live to list the filings published between 2026-09-01 and 2026-09-18,
then summarise the interim report.

Ready-made configuration for Claude, ChatGPT, Cursor, VS Code/Copilot, Gemini CLI, opencode, Manus, and Perplexity is in AI agent support — and for a stored corpus, the stdio MCP server exposes a wider tool catalog and is published on Glama. The gateway is listed in the official MCP Registry as io.github.simonmak-ascent/hkex-filings.

Featured on Glama — the read-only stdio MCP server is also published on Glama, where Glama scans the built server and scores tool-definition quality (currently 4.7/5).

Quick Start (Local)

pip install hkex-filing-scraper        # core; SQLite needs no server
pip install "hkex-filing-scraper[all]" # Excel + dotenv + every driver + the MCP server
cp .env.example .env                   # then set DATABASE_TARGET (below)
hkex-scraper --metadata-only --limit 100

Optional extras: excel, postgres, mysql, duckdb, mongodb, clickhouse, neo4j, mcp, pdf, all, dev.

DATABASE_TARGET is an ordered, comma-separated list of sink ids; the order decides which sink serves reads. To start with no server:

DATABASE_TARGET=sqlite
SQLITE_PATH=hkex.db

hkex-scraper runs the full pipeline (metadata + documents + graph); hkex-scraper --full-history covers everything since April 1999. The schema is created automatically. Full install options and per-sink settings are in Getting started.

Database Support

Every sink is a first-class destination; rows are in documented popularity order. The full matrix — licenses, capability differences, per-engine notes — is in Database sinks.

Sink

Model

License

Extra

Idempotent upsert

postgres

relational

PostgreSQL License

postgres

ON CONFLICT DO UPDATE

mysql / mariadb

relational

GPLv2

mysql

ON DUPLICATE KEY UPDATE

sqlite

relational

Public domain

—

ON CONFLICT DO UPDATE

mongodb

document

SSPL¹

mongodb

update_one(upsert=True)

neo4j

graph

GPLv3 (Community)

neo4j

MERGE

clickhouse

columnar

Apache-2.0

clickhouse

ReplacingMergeTree + read-merge

duckdb

relational

MIT

duckdb

ON CONFLICT DO UPDATE

surrealdb

graph + document

BSL 1.1¹

—

UPSERT / RELATE

¹ Source-available, not OSI-approved — labeled exceptions per ADR 0003.

Valid sink ids, in documented order: postgres, mysql, sqlite, mongodb, mariadb, neo4j, clickhouse, duckdb, surrealdb. Set one variable and the same run feeds every sink:

# Order sets read precedence.
DATABASE_TARGET=postgres,sqlite
POSTGRES_DSN=postgresql://user:password@localhost:5432/hkex
SQLITE_PATH=hkex.db

How This Compares

Four ways to get HKEx filings, and what each one costs you.

This project

HKEXnews web search

Browser automation you write

Licensed HKEx feed

Bulk export

Yes

No — page-at-a-time

Yes

Yes

History to April 1999

Yes

Yes, manually

Depends on your code

Yes

Full text of documents

Extracted from PDF/HTML/Excel

No — you open each file

You build the extractor

Varies by contract

Structured tables

Extracted to Markdown

No

You build it

Varies

Coverage verification

Per-chunk, auditable

Not applicable

You build it

Vendor SLA

Lands in your engine

9 engines, any combination

No

Whatever you wire up

Usually one format

Speed

JSON API, no browser

Manual

Slower — renders pages

Fast

Cost

Free, MIT

Free

Your time

Subscription

Commercial redistribution

See docs/legal.md

Restricted

Restricted

Licensed

If you need licensed, redistributable, SLA-backed data, buy the feed. If you need a complete local corpus for research, compliance, or RAG, this replaces the pipeline you would otherwise write yourself.

How It Works

flowchart LR
    A[HKEx JSON API] --> B[Phase 1: metadata]
    B --> C[Canonical record]
    C --> D{DATABASE_TARGET}
    D --> E[(PostgreSQL)]
    D --> F[(MySQL / MariaDB)]
    D --> G[(SQLite)]
    D --> H[(MongoDB)]
    D --> I[(Neo4j)]
    D --> J[(ClickHouse)]
    D --> K[(DuckDB)]
    D --> L[(SurrealDB)]
    B --> M[Graph linking]
    M --> D
    B --> N[Phase 2: download and extract]
    N --> C
  • Phase 1 scrapes filing metadata through a JSF session, splitting the range into monthly chunks and deduplicating on a 16-character MD5 filingId.

  • Phase 2 downloads each filing's PDF/HTML/Excel document, extracts text and tables to Markdown, and writes the payload.

  • Graph linking (optional) writes has_filing and references_filing edges when COMPANY_TABLE is set.

  • Failure isolation — a failure on one sink is logged and counted but never blocks another; the run exits non-zero if any configured sink failed.

Deeper detail: Architecture · ADR 0002.

Features

  • Fast API scraping — direct HKEx JSON API; no browser or Selenium.

  • Full history — every filing from April 1999 to today, with chunk-level coverage checks.

  • Document processing — PDF/HTML/Excel text and structured tables, extracted to Markdown.

  • Multi-sink — any ordered combination of nine databases, each with native idempotent upserts.

  • AI-ready — a hosted live MCP gateway plus a local stdio MCP server.

  • Resumable and observable — batching, parallel downloads, stalled-job detection, per-sink counters, and --coverage-report / --parity-report / --verify.

  • Optional dependencies — the core is requests + beautifulsoup4; drivers and document extraction are extras with graceful fallbacks.

Documentation

Development

pip install -e ".[dev,all]"
ruff check           # lint (py310, line-length 100)
ruff format --check  # formatting
pytest               # unit tests (no DB or network required)

Tests are pure unit tests; SQLite and DuckDB contract tests run in-process, and integration tests that need a server are skipped unless that sink is configured. See Testing.

Contributing

See CONTRIBUTING.md; report security issues per SECURITY.md. Ideas and questions are welcome in Discussions.

Built by Ascent Partners.

If this saves you time, a ⭐ on GitHub helps others find it.

Use with Context7

Up-to-date HKEx Filing Scraper documentation is indexed on Context7, so coding agents can pull it into context on demand. With the Context7 MCP server or ctx7 CLI installed, name the library in your prompt:

use library /simonmak-ascent/hkex-filing-scraper for API and docs

License

MIT — see LICENSE. That covers this project's code only; optional dependencies carry their own licenses, notably the pdf extra (PyMuPDF / pymupdf4llm), which is AGPL-3.0 and deliberately excluded from .[all]. See docs/legal.md.

Data & Terms of Use: this is a research tool for the undocumented HKEx JSON API, and it is not affiliated with or endorsed by HKEx. Commercial redistribution of HKEx data may require a licensed HKEx feed; see docs/legal.md.

Available Tools

12 tools
describe_schemaA
Read-onlyIdempotent
Inspect

Use this before filtering or interpreting results to learn the canonical fields.

Use search_filings or search_documents to query filings once you know the field names. Pass section to return just one section (saves tokens); default returns everything: filing and document fields, the known filing types/categories/statuses/types, graph edge kinds, and the supported query filters. Reads no filings.

ParametersJSON Schema
NameRequiredDescriptionDefault
sectionNoNarrow to one section: filing, document, types, query, or all (default).all

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context: it reads no filings, defaults to returning all schema sections, and can save tokens when narrowed with section.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the key guidance. Each sentence earns its place: when to use, what to use instead, how to narrow output, and what the default returns. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only schema-introspection tool with one optional parameter and an output schema, this description is complete. It covers use timing, sibling routing, parameter behavior, and side-effect profile without needing to document return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the parameter is already fully documented. The description adds value by explaining the token-saving rationale for passing section and enumerating what the default return includes, going beyond the schema's enum definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific purpose: learn the canonical fields before filtering or interpreting results. It names the resource (the schema) and clearly distinguishes the tool from query siblings like search_filings and search_documents.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use the tool (before filtering/interpreting results) and points to alternatives for querying once field names are known. This gives an agent a concrete decision rule rather than leaving usage to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_coverageA
Read-onlyIdempotent
Inspect

Use this to report scrape coverage per monthly chunk, with totals.

Use get_statistics for filing counts grouped by a dimension, or verify_sinks for cross-sink comparison. date_from/date_to (YYYY-MM-DD) filter by chunk month. limit caps how many chunks are returned, but totals always aggregate the full filtered range. Rows are newest-first. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum coverage rows to return (1..200).
date_toNoFilter by chunk month, YYYY-MM-DD.
date_fromNoFilter by chunk month, YYYY-MM-DD.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, closed-world, so repeating 'Read-only' adds nothing. However the description discloses two non-obvious behaviors beyond the annotations: rows are newest-first, and totals aggregate the full filtered range regardless of limit. That is real behavioral value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then alternatives, then parameter behavior, then ordering and read-only note. Every sentence earns its place, though the trailing 'Read-only' is redundant with the annotations and the multi-clause structure is slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return format need not be explained. With purpose, alternatives, filtering semantics, ordering, and aggregation behavior all covered, an agent has everything needed to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds meaning the schema does not: date_from/date_to filter by chunk month (not filing or ingestion date), and limit caps rows while totals stay unaffected by it. That clarification materially changes how limit is used.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — reporting scrape coverage per monthly chunk, with totals — and explicitly contrasts itself with get_statistics and verify_sinks in the sibling list. An agent can distinguish it from every listed sibling without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names two alternatives with the exact condition that selects each: get_statistics for filing counts grouped by a dimension, verify_sinks for cross-sink comparison. Explicit routing guidance, not merely implied context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_filingA
Read-onlyIdempotent
Inspect

Use this to read one filing's metadata and extracted document content.

Returns the canonical filing and document fields. document_text is a window of max_text_chars from text_offset; when text_truncated is true, call again with next_text_offset for more. Tables (document_tables) are included only when include_tables is true. Obtain ids from search_filings. This tool is read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
filing_idYes16-character filing id; obtain from search_filings.
text_offsetNoCharacter offset into document_text for paging.
include_textNoInclude the extracted document text window.
include_tablesNoInclude extracted document tables.
max_text_charsNoMaximum characters of text to return (0..200000).

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description adds a precise paging contract: document_text is a window of max_text_chars from text_offset, text_truncated signals another page, and next_text_offset is the continuation point. It also discloses that tables are only included when include_tables is true, which is valuable behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: core purpose first, then return fields, paging behavior, table inclusion, and the source of IDs. Every sentence contributes operational guidance with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a full input schema, rich annotations, and an output schema present, the description covers the remaining operational essentials: paging, table inclusion, and how to get the filing ID. An agent has everything needed to invoke the tool correctly and process paginated results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers all five parameters with descriptions, so the baseline is solid. The description adds extra relational meaning by explaining how text_offset, max_text_chars, and next_text_offset work together for pagingapper. It does not add much on filing_id, but the schema already covers that sufficiently.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'read one filing's metadata and extracted document content.' It clearly identifies this as a single-filing accessor and distinguishes it from search-oriented siblings by directing agents to obtain IDs from search_filings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the intended use case ('Use this to read one filing's...') and the prerequisite workflow of obtaining IDs from search_filings. However, it does not explicitly name alternatives like get_filings or describe when not to use them, so the guidance is clear but not fully exclusionary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_filingsA
Read-onlyIdempotent
Inspect

Use this to read several filings in one call (up to 50 ids).

Use get_filing instead to read a single filing. Returns each filing's metadata plus, when include_text is true, a bounded text window. Ids not found are listed in not_found. Text is off by default. Parameter relationships: filing_ids must come from search_filings (1..50 ids); max_text_chars applies per filing and only takes effect when include_text is true; include_tables is independent of the text window. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
filing_idsYesFiling ids to fetch (1..50); obtain from search_filings.
include_textNoInclude the extracted document text window (off by default).
include_tablesNoInclude extracted document tables.
max_text_charsNoMaximum characters of text to return (0..200000).

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly=true, idempotent=true and destructive=false, so safety is covered. The description nonetheless adds real behavior: missing ids surfaced in ``not_found``, text extraction off by default, and the text window being bounded. It stops short of describing pagination or ordering, so not a full 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the batch purpose and the sibling contrast, then groups parameter relationships compactly. Minor redundancy with schema defaults (text off by default is stated twice) keeps it just shy of 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, yet the description still flags the ``not_found`` field and the default-off text behavior. Combined with the id-count limit and source constraint, an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the 3 baseline would apply on defaults alone. The description goes further by spelling out inter-parameter relationships: max_text_chars is per filing and only effective when include_text is true, and include_tables is independent of the text window.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource ('read several filings in one call, up to 50 ids') and explicitly contrasts with the sibling get_filing for single-filing reads. An agent can distinguish it from get_filing and search_filings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative ('Use get_filing instead to read a single filing') with the exact condition that selects it, and states the source constraint for ids ('must come from search_filings'). Both the routing decision and the prerequisite are explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_server_infoA
Read-onlyIdempotent
Inspect

Use this first to learn the server version, configured sinks, and read sink.

Use include="config" for the raw configuration values or list_sinks for per-sink detail. include is one of: summary (server metadata only, default), sinks (also returns the list_sinks payload), or config (also returns the configuration payload) — so you can pull the summary and the detail in a single call. Returns server metadata only; it reads no filings. include="sinks" is an alias for the list_sinks payload — call list_sinks directly when sinks are the only thing you need. This tool is read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
includeNosummary (server metadata only), sinks, or config; sinks/config fold in list_sinks output and the configuration payload.summary

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered; the description nonetheless adds real behavioral context ('Returns server metadata only; it reads no filings') and clarifies that the call can fold in other payloads in a single request. The closing 'This tool is read-only' is redundant with readOnlyHint, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the primary instruction and organized around the include values. Mild redundancy between 'Returns server metadata only; it reads no filings' and the trailing 'This tool is read-only', plus a slightly awkward em-dash sentence, costs a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained; the description covers purpose, routing, and the semantics of the single parameter. For a zero-required-parameter read tool, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the enum is self-describing, but the description adds meaning beyond the schema: it explains the consequence of each value (folds in the list_sinks payload or the configuration payload) and names the defaults, which the schema alone does not fully convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (learn server version, configured sinks, read sink) with a concrete priming instruction ('Use this first'). It also distinguishes itself from list_sinks, which an agent can act on without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('use this first'), explicit alternative routing ('call list_sinks directly when sinks are the only thing you need'), and explicit per-value intent for include=config vs include=sinks. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_statisticsA
Read-onlyIdempotent
Inspect

Use this to count filings grouped by one dimension, or per sink.

Choose between: this tool gives a grouped breakdown of one population (or per-sink totals when group_by="sink"); get_coverage reports scrape coverage by month; verify_sinks compares across sinks. group_by is one of: company_ticker (default), filing_type, filing_category, document_status, exchange, or sink. Optional filters narrow the population and combine with AND semantics (a filing must match every filter you set); group_by="sink" honours only ticker, filing_type, filing_category, document_status, date_from, and date_to, and returns per-sink totals (relational sinks only). top_n caps the returned buckets and min_count drops small buckets, but total still counts every matching filing. Buckets are sorted by count descending. This tool is read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
top_nNoMaximum buckets to return (1..200).
sourceNoSource of the filing, e.g. HKEx.
tickerNoCompany ticker filter, e.g. 0700.HK; comma-separate to match several.
date_toNoLatest filing date, YYYY-MM-DD inclusive.
exchangeNoExchange code, e.g. HK.
group_byNoDimension to count filings by: company_ticker (default), filing_type, filing_category, document_status, exchange, or sink (per-sink totals across configured sinks).company_ticker
date_fromNoEarliest filing date, YYYY-MM-DD inclusive.
min_countNoOnly return buckets with at least this many filings (0 = all).
filing_typeNoFiling type(s), e.g. 'Annual Report'; comma-separate to match several.
title_queryNoCase-insensitive substring matched against the filing title.
document_typeNoDocument type: pdf, html, xlsx, docx, or unknown.
document_statusNoDocument status(es): processed, skipped, failed, or unprocessed; comma-separate to match several.
filing_categoryNoFiling category(ies), e.g. LISTED_COMPANY; comma-separate to match several.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds substantial behavior beyond them: filters combine with AND semantics, group_by="sink" only honours a subset of filters and returns relational-sink totals, top_n/min_count cap display buckets while total still counts every match, and buckets are sorted count-descending. This is exactly the extra context the annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then siblings, then semantics — a logical order. It is dense and long, but nearly every clause carries a distinct behavioral fact (filter interaction, sink caveat, bucket vs total), so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter read-only aggregation tool with 100% schema coverage and an output schema present, the description covers the interaction rules and edge cases an agent needs (AND filters, sink-mode caveats, total vs buckets). Return-format explanation is correctly left to the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline would be 3, but the description adds real meaning on top: the AND semantics across the 13 filters, the filter-subset restriction under group_by="sink", and the distinction that min_count/top_n never affect the reported total. It does not document every individual filter, which the schema already does.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: counting filings grouped by one dimension or per sink. It also explicitly contrasts itself with its two closest siblings (get_coverage, verify_sinks), so an agent can route without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use criteria ('grouped breakdown of one population') and names the alternatives with their distinguishing conditions: get_coverage = scrape coverage by month, verify_sinks = comparison across sinks. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_companiesA
Read-onlyIdempotent
Inspect

Use this to list companies, or just their ticker codes, that have filings.

view="companies" (default) returns each company's ticker, name, and filing count, ordered by filing count descending (ties broken by ticker ascending). view="tickers" returns only the distinct ticker codes, alphabetically sorted and paged. ticker is a case-insensitive substring match (e.g. '0700' matches '0700.HK'). Not for a company's filings (use search_filings) or for aggregate counts (use get_statistics with group_by="company_ticker"). This tool is read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
viewNoWhat to list: "companies" (default) returns ticker + name + filing count; "tickers" returns just the distinct ticker codes.companies
limitNoMaximum rows (1..100 for companies; 1..1000 for tickers).
offsetNoZero-based offset for paging.
tickerNoCase-insensitive substring match, e.g. '0700'; empty lists all.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, and the closing sentence confirms read-only. Beyond that, the description discloses non-obvious behavior the annotations cannot: default ordering by filing count descending with ticker tie-break, alphabetical sorting in tickers view, and which view is paged. It does not discuss rate limits or the cost of the tickers view, but the behavioral picture is notably richer than the annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then view semantics, then the ticker filter, then exclusions and safety — a sensible order with no filler sentences. Slightly dense and the '0700' example duplicates the schema's own example, which is the only real redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the remaining agent-facing questions: which view to pick, how results are ordered, how paging behaves, and which sibling to use instead. Nothing needed to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: the view enum's return shape and sort order (count-descending for companies, alphabetical for tickers) and paging scope, plus the case-insensitive substring semantics of ticker with a concrete example. It goes beyond restating the schema even though the ticker example is duplicated from the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('list companies... that have filings') and immediately distinguishes the two modes of operation (view="companies" vs view="tickers"). It also names the sibling tools it is not, so an agent can separate it from search_filings and get_statistics without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly carves out the two wrong-use cases: 'Not for a company's filings (use search_filings) or for aggregate counts (use get_statistics with group_by="company_ticker")'. This is textbook when-to-use/when-not-to-use with named alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_referencesA
Read-onlyIdempotent
Inspect

Use this to explore the graph edges between companies and filings.

Use search_filings(ticker=...) for a company's own filings or search_filings( referenced_ticker=...) to find mentions via the filing column. This reads the canonical edge tables keyed on ticker (format like 0700.HK; discover valid values via list_companies(view="tickers")): kind="referenced_by" returns filings whose title mentions the company (cross-references), kind="owned" returns the company's own filings. Results are paged via limit/offset and stay empty until graph linking has been run. This tool is read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoreferenced_by = filings that mention this company; owned = this company's own filings.referenced_by
limitNoMaximum edges to return (1..100).
offsetNoZero-based offset for paging.
tickerYesCompany ticker, e.g. 0700.HK; use list_companies(view='tickers') to discover.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/destructiveHint, so the safety profile is covered. The description adds real value beyond them: results are paged via limit/offset and 'stay empty until graph linking has been run' — a non-obvious precondition an agent would otherwise misread as a bug.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then alternatives, then semantics and caveats. Every sentence carries information, though the parenthetical list_companies reference and repeated ticker format make it slightly longer than strictly necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return shapes need no explanation, and the description still covers the paging model, the empty-until-linked precondition, the kind semantics, ticker validation, and read-only status. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description goes beyond the schema by explaining what kind='referenced_by' vs 'owned' mean in graph terms (title-mention cross-references vs the company's own filings) and reiterates the ticker format with a discovery path, giving the enum semantics real operational meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('explore the graph edges between companies and filings') and immediately names the sibling tools (search_filings) it must not be confused with. An agent can distinguish it from get_filings/search_filings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: search_filings(ticker=...) for a company's own filings vs search_filings(referenced_ticker=...) for mentions, and explains that this tool reads canonical edge tables. It also names list_companies(view='tickers') as the discovery path for valid ticker values.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_sinksA
Read-onlyIdempotent
Inspect

Use this when the user asks which databases are configured or their capabilities.

Use get_server_info (with include="config") for the raw DATABASE_TARGET string or a one-line summary instead. Pass sink_id to inspect one sink without the full list. Returns every known sink id with its license, optional extra, configured/available status, per-sink capabilities, and which sink serves reads. Reads no filings. get_server_info with include="sinks" returns this same payload; call this tool directly when per-sink detail is all you need.

ParametersJSON Schema
NameRequiredDescriptionDefault
sink_idNoNarrow to one configured sink id, e.g. 'postgres'; empty returns all. Call with no argument to discover the valid ids.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower, and the description still adds real context: it discloses the return contents (ids, license, optional extra, configured/available status, per-sink capabilities, read-serving sink) and explicitly states 'Reads no filings,' which is a meaningful scope constraint. It stops short of noting any cost or rate-limit behavior, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the trigger condition, and every sentence is on-topic. There is mild redundancy: get_server_info is invoked twice and the equivalence 'get_server_info with include="sinks" returns this same payload' restates the routing already implied, which adds length without new decision value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite a single optional parameter and an existing output schema, the description stays complete: trigger, alternatives, narrowing path, and return-contents summary are all present. Nothing an agent needs to call it correctly is missing, and no return-value documentation is required given the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already explains sink_id (narrow to one configured sink id, empty returns all). The description reinforces the semantics contextually ('pass sink_id to inspect one sink without the full list'), tying the parameter to an intent rather than just a type, which is a modest but genuine addition over the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (list configured sinks/databases and their capabilities) and explicitly differentiates itself from the closest sibling, get_server_info, including the exact argument variant that would substitute for it. An agent can select this over siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('user asks which databases are configured or their capabilities') and when-not-to-use (use get_server_info with include='config' for the raw DATABASE_TARGET string, or include='sinks' for the same payload). It also names the sink_id path for narrowing to a single sink, covering all three invocation modes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_documentsA
Read-onlyIdempotent
Inspect

Use this for full-text search over extracted document text.

Use search_filings instead to filter by metadata without a text query. Matches text_query case-insensitively inside document_text and returns filing rows with a snippet when the sink supports it (see snippets_supported). document_type accepts pdf, html, xlsx, docx, or unknown; source is a free-text origin like HKEx. All filters combine with AND semantics (a filing must match every filter you set). Returns nothing until documents are processed. This tool is read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
offsetNoZero-based offset for paging.
sourceNoSource of the filing, e.g. HKEx.
tickerNoCompany ticker filter, e.g. 0700.HK; comma-separate to match several.
date_toNoLatest filing date, YYYY-MM-DD inclusive.
order_byNoSort order of the result.filing_date_desc
date_fromNoEarliest filing date, YYYY-MM-DD inclusive.
page_sizeNoMaximum filings to return (1..100).
text_queryYesCase-insensitive term matched against extracted document text.
filing_typeNoFiling type(s), e.g. 'Annual Report'; comma-separate to match several.
document_typeNoDocument type: pdf, html, xlsx, docx, or unknown.
document_statusNoDocument status(es): processed, skipped, failed, or unprocessed; comma-separate to match several.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral context beyond the annotations: case-insensitive matching, snippet availability depending on sink support, AND-filter semantics, and the dependency on documents being processed before results appear. It also restates the read-only nature consistently with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized and front-loaded with the core use case, followed by routing guidance and behavioral notes. Every sentence adds value without unnecessary repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 100% schema coverage, an output schema, and strong annotations, the description covers the remaining contextual gaps: when to use it, how filters behave, and what prerequisites affect results. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema for key parameters: text_query's matching behavior, document_type's allowed values, source's example, and the AND combination of filters. It doesn't discuss every parameter, but the schema already covers those individually.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: full-text search over extracted document text, matching text_query inside document_text. It clearly distinguishes itself from the sibling search_filings by contrasting full-text search with metadata-only filtering.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative: 'Use search_filings instead to filter by metadata without a text query.' It also clarifies when search_documents applies and notes that all filters combine with AND semantics, giving an agent a clear decision rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_filingsA
Read-onlyIdempotent
Inspect

Use this to find filings by ticker, type, status, date range, or title text.

Prefer search_documents for full-text search over extracted text, and get_filing or get_filings to read filings whose ids you already have. Filters are optional and combinable; comma-separate a value to match several (e.g. filing_type="Annual Report,Dividend"). document_status accepts the real statuses plus unprocessed (no document yet). document_type accepts pdf, html, xlsx, docx, or unknown, and source is a free-text origin like HKEx. date_from/date_to are YYYY-MM-DD inclusive. order_by is one of filing_date_desc (default), filing_date_asc, title_asc, filing_id_asc. Returns paged filing rows (no document text); call get_filing for the document. This tool is read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
offsetNoZero-based offset for paging.
sourceNoSource of the filing, e.g. HKEx.
tickerNoCompany ticker filter, e.g. 0700.HK; comma-separate to match several.
date_toNoLatest filing date, YYYY-MM-DD inclusive.
exchangeNoExchange code, e.g. HK.
order_byNoSort order of the result.filing_date_desc
date_fromNoEarliest filing date, YYYY-MM-DD inclusive.
page_sizeNoMaximum filings to return (1..100).
stock_codeNoNumeric stock code filter, e.g. 00700; comma-separate to match several.
filing_typeNoFiling type(s), e.g. 'Annual Report'; comma-separate to match several.
title_queryNoCase-insensitive substring matched against the filing title.
document_typeNoDocument type: pdf, html, xlsx, docx, or unknown.
document_statusNoDocument status(es): processed, skipped, failed, or unprocessed; comma-separate to match several.
filing_categoryNoFiling category(ies), e.g. LISTED_COMPANY; comma-separate to match several.
referenced_tickerNoTicker referenced by the filing (graph edge).

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered structurally. The description adds genuine value by disclosing return behavior ('Returns paged filing rows (no document text)') and explaining the 'unprocessed' status edge case — details an agent could not infer from annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and routing are front-loaded in the first two sentences, and the filter-semantics block is dense but every sentence earns its place for a 15-parameter, all-optional tool. It is a long single paragraph and could benefit from light grouping, but there is no fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex search tool with 15 optional parameters, the description covers everything needed for correct invocation: purpose, sibling routing, filter combinability, value domains, date semantics, sort options, return behavior, and read-only confirmation. Output schema covers return fields and annotations cover safety, so no critical gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, baseline is 3, but the description adds meaning above that: it generalizes comma-separation across filters with a concrete example (filing_type="Annual Report,Dividend"), clarifies date formats as YYYY-MM-DD inclusive, enumerates order_by values with the default, and defines document_status's special 'unprocessed' value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb plus resource and scope: 'find filings by ticker, type, status, date range, or title text.' It differentiates from siblings by naming what it is not — search_documents (full-text over extracted text) and get_filing/get_filings (by known ids) — so an agent can route correctly without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing: 'Prefer search_documents for full-text search over extracted text, and get_filing or get_filings to read filings whose ids you already have.' It also states that filters are optional and combinable, covering both when to use and when not to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_sinksA
Read-onlyIdempotent
Inspect

Use this to check that configured sinks agree on their filings.

mode="hashes" (default) compares (filing_id, document_sha256) sets across comparable sinks and returns a bounded sample of any missing/extra/mismatched ids. mode="counts" runs a faster count-only parity check, returning per-sink counts and the spread (parity is OK when the spread is zero). For a per-sink total or a grouped breakdown use get_statistics, and for monthly scrape coverage use get_coverage. Parameter relationships: sinks must list two or more comparable sink ids from list_sinks (empty compares every comparable pair); sample_size bounds the examples returned per problem bucket, not the comparison itself, and is ignored by mode="counts". This tool is read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoComparison depth: "hashes" (default) compares per-filing ids and document hashes; "counts" is a quick per-sink count parity check.hashes
sinksNoRestrict to a subset of sink ids, comma-separated; empty = all.
sample_sizeNoMax examples to return per problem bucket (1..50).

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive behavior, but the description adds substantial behavioral detail: it explains the return shape (bounded samples of missing/extra/mismatched ids, per-sink counts and spread, parity OK when spread is zero) and clarifies that sample_size bounds examples rather than the comparison and is ignored in counts mode. It also notes the tool is read-only, reinforcing annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose, then mode details, alternatives, and parameter relationships, with no filler. It is a bit dense and includes one redundant sentence ("This tool is read-only"), but the length is justified by the tool's comparison modes and cross-parameter constraints.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the 100% schema coverage, annotations, and an existing output schema, the description covers everything an agent needs: when to use each mode, what the tool returns, parameter constraints, and explicit alternatives. No critical gap remains for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful cross-parameter semantics beyond the schema: sinks must list two or more comparable sink ids from list_sinks, empty compares every comparable pair, and sample_size bounds examples per problem bucket rather than the comparison itself and is ignored by mode="counts". These relationships are not captured in the input schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: check that configured sinks agree on their filings, and clarifies the two modes (hashes vs counts). It also names sibling tools get_statistics and get_coverage for adjacent tasks, so an agent can distinguish this tool's purpose without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance for each mode and names alternatives: use get_statistics for per-sink totals or grouped breakdowns, and get_coverage for monthly scrape coverage. It also states that sinks must list two or more comparable sink ids from list_sinks, and that empty compares every comparable pair.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv2.4.1
    • Changedlist_companies1 field changed
      • addedInput schema / properties / view / enum
        Added value: +[
        +  "companies",
        +  "tickers"
        +]
    • Changedlist_sinks1 field changed
      • changedInput schema / properties / sink_id / description
        Previous value: -"Narrow to a single sink id, e.g. 'postgres'; empty returns all."New value: +"Narrow to one configured sink id, e.g. 'postgres'; empty returns all. Call with no argument to discover the valid ids."
  2. 5 tool updatesv0.1.9
    • Removedcount_filings
    • Changedget_statistics2 fields changed
      • changedInput schema / properties / group_by / description
        Previous value: -"Dimension to count filings by."New value: +"Dimension to count filings by: company_ticker (default), filing_type, filing_category, document_status, exchange, or sink (per-sink totals across configured sinks)."
      • changedInput schema / properties / group_by / enum
        Previous value: -[
        -  "company_ticker",
        -  "filing_type",
        -  "filing_category",
        -  "document_status",
        -  "exchange"
        -]New value: +[
        +  "company_ticker",
        +  "filing_type",
        +  "filing_category",
        +  "document_status",
        +  "exchange",
        +  "sink"
        +]
    • Changedlist_companies2 fields changed
      • changedInput schema / properties / limit / description
        Previous value: -"Maximum companies to return (1..100)."New value: +"Maximum rows (1..100 for companies; 1..1000 for tickers)."
      • addedInput schema / properties / view
        Added value: +{
        +  "default": "companies",
        +  "description": "What to list: \"companies\" (default) returns ticker + name + filing count; \"tickers\" returns just the distinct ticker codes.",
        +  "title": "View",
        +  "type": "string"
        +}
    • Changedlist_references1 field changed
      • changedInput schema / properties / ticker / description
        Previous value: -"Company ticker, e.g. 0700.HK; use list_tickers to discover."New value: +"Company ticker, e.g. 0700.HK; use list_companies(view='tickers') to discover."
    • Removedlist_tickers
  3. 2 tool updatesv0.1.8
    • Removedget_parity
    • Changedverify_sinks1 field changed
      • addedInput schema / properties / mode
        Added value: +{
        +  "default": "hashes",
        +  "description": "Comparison depth: \"hashes\" (default) compares per-filing ids and document hashes; \"counts\" is a quick per-sink count parity check.",
        +  "title": "Mode",
        +  "type": "string"
        +}
  4. 3 tool updatesv0.1.6
    • Removedget_config
    • Changedget_server_info1 field changed
      • changedInput schema / properties / include / description
        Previous value: -"summary (server metadata only), sinks, or config; sinks/config fold in list_sinks/get_config output."New value: +"summary (server metadata only), sinks, or config; sinks/config fold in list_sinks output and the configuration payload."
    • Removedlist_pending_filings
  5. 2 tool updatesv0.1.5
    • Changedlist_companies1 field changed
      • addedInput schema / properties / ticker
        Added value: +{
        +  "default": "",
        +  "description": "Case-insensitive substring match, e.g. '0700'; empty lists all.",
        +  "title": "Ticker",
        +  "type": "string"
        +}
    • Changedlist_tickers1 field changed
      • addedInput schema / properties / ticker
        Added value: +{
        +  "default": "",
        +  "description": "Case-insensitive substring match, e.g. '0700'; empty lists all.",
        +  "title": "Ticker",
        +  "type": "string"
        +}
  6. 12 tool updatesv0.1.3
    • Changedcount_filings6 fields changed
      • addedInput schema / properties / date_from
        Added value: +{
        +  "default": "",
        +  "description": "Count only filings on/after this date, YYYY-MM-DD inclusive.",
        +  "title": "Date From",
        +  "type": "string"
        +}
      • addedInput schema / properties / date_to
        Added value: +{
        +  "default": "",
        +  "description": "Count only filings on/before this date, YYYY-MM-DD inclusive.",
        +  "title": "Date To",
        +  "type": "string"
        +}
      • addedInput schema / properties / document_status
        Added value: +{
        +  "default": "",
        +  "description": "Count only this document status: processed, skipped, failed, or unprocessed.",
        +  "title": "Document Status",
        +  "type": "string"
        +}
      • addedInput schema / properties / filing_category
        Added value: +{
        +  "default": "",
        +  "description": "Count only this filing category, e.g. LISTED_COMPANY.",
        +  "title": "Filing Category",
        +  "type": "string"
        +}
      • addedInput schema / properties / filing_type
        Added value: +{
        +  "default": "",
        +  "description": "Count only this filing type, e.g. 'Annual Report'; comma-separate for several.",
        +  "title": "Filing Type",
        +  "type": "string"
        +}
      • addedInput schema / properties / ticker
        Added value: +{
        +  "default": "",
        +  "description": "Count only this ticker, e.g. 0700.HK; comma-separate to match several.",
        +  "title": "Ticker",
        +  "type": "string"
        +}
    • Changeddescribe_schema1 field changed
      • addedInput schema / properties / section
        Added value: +{
        +  "default": "all",
        +  "description": "Narrow to one section: filing, document, types, query, or all (default).",
        +  "enum": [
        +    "all",
        +    "filing",
        +    "document",
        +    "types",
        +    "query"
        +  ],
        +  "title": "Section",
        +  "type": "string"
        +}
    • Changedget_config1 field changed
      • addedInput schema / properties / key
        Added value: +{
        +  "default": "",
        +  "description": "Return a single setting, e.g. 'database_target'; empty returns all.",
        +  "title": "Key",
        +  "type": "string"
        +}
    • Changedget_parity1 field changed
      • addedInput schema / properties / sinks
        Added value: +{
        +  "default": "",
        +  "description": "Restrict to a subset of sink ids, comma-separated; empty = all.",
        +  "title": "Sinks",
        +  "type": "string"
        +}
    • Changedget_server_info1 field changed
      • addedInput schema / properties / include
        Added value: +{
        +  "default": "summary",
        +  "description": "summary (server metadata only), sinks, or config; sinks/config fold in list_sinks/get_config output.",
        +  "enum": [
        +    "summary",
        +    "sinks",
        +    "config"
        +  ],
        +  "title": "Include",
        +  "type": "string"
        +}
    • Changedget_statistics4 fields changed
      • addedInput schema / properties / document_type
        Added value: +{
        +  "default": "",
        +  "description": "Document type: pdf, html, xlsx, docx, or unknown.",
        +  "title": "Document Type",
        +  "type": "string"
        +}
      • addedInput schema / properties / min_count
        Added value: +{
        +  "default": 1,
        +  "description": "Only return buckets with at least this many filings (0 = all).",
        +  "title": "Min Count",
        +  "type": "integer"
        +}
      • addedInput schema / properties / source
        Added value: +{
        +  "default": "",
        +  "description": "Source of the filing, e.g. HKEx.",
        +  "title": "Source",
        +  "type": "string"
        +}
      • addedInput schema / properties / top_n
        Added value: +{
        +  "default": 20,
        +  "description": "Maximum buckets to return (1..200).",
        +  "title": "Top N",
        +  "type": "integer"
        +}
    • Changedlist_pending_filings1 field changed
      • addedInput schema / properties / offset
        Added value: +{
        +  "default": 0,
        +  "description": "Zero-based offset for paging.",
        +  "title": "Offset",
        +  "type": "integer"
        +}
    • Addedlist_references
    • Changedlist_sinks1 field changed
      • addedInput schema / properties / sink_id
        Added value: +{
        +  "default": "",
        +  "description": "Narrow to a single sink id, e.g. 'postgres'; empty returns all.",
        +  "title": "Sink Id",
        +  "type": "string"
        +}
    • Changedsearch_documents2 fields changed
      • addedInput schema / properties / document_type
        Added value: +{
        +  "default": "",
        +  "description": "Document type: pdf, html, xlsx, docx, or unknown.",
        +  "title": "Document Type",
        +  "type": "string"
        +}
      • addedInput schema / properties / source
        Added value: +{
        +  "default": "",
        +  "description": "Source of the filing, e.g. HKEx.",
        +  "title": "Source",
        +  "type": "string"
        +}
    • Changedsearch_filings2 fields changed
      • addedInput schema / properties / document_type
        Added value: +{
        +  "default": "",
        +  "description": "Document type: pdf, html, xlsx, docx, or unknown.",
        +  "title": "Document Type",
        +  "type": "string"
        +}
      • addedInput schema / properties / source
        Added value: +{
        +  "default": "",
        +  "description": "Source of the filing, e.g. HKEx.",
        +  "title": "Source",
        +  "type": "string"
        +}
    • Changedverify_sinks2 fields changed
      • addedInput schema / properties / sample_size
        Added value: +{
        +  "default": 5,
        +  "description": "Max examples to return per problem bucket (1..50).",
        +  "title": "Sample Size",
        +  "type": "integer"
        +}
      • addedInput schema / properties / sinks
        Added value: +{
        +  "default": "",
        +  "description": "Restrict to a subset of sink ids, comma-separated; empty = all.",
        +  "title": "Sinks",
        +  "type": "string"
        +}
  7. 16 tool updates
    • First observedcount_filings
    • First observeddescribe_schema
    • First observedget_config
    • First observedget_coverage
    • First observedget_filing
    • First observedget_filings
    • First observedget_parity
    • First observedget_server_info
    • First observedget_statistics
    • First observedlist_companies
    • First observedlist_pending_filings
    • First observedlist_sinks
    • First observedlist_tickers
    • First observedsearch_documents
    • First observedsearch_filings
    • First observedverify_sinks

TDQS

A4.7/5.0

Scored across 12 tools

Disambiguation4/5

Most tools are clearly separated by resource and action, and descriptions explicitly cross-reference alternatives. The main overlap is list_sinks versus get_server_info(include="sinks"), which return the same payload; get_server_info also overlaps with describe_schema and the sink-listing behavior.

Naming Consistency5/5

All tool names use consistent snake_case verb_noun patterns (list_sinks, get_statistics, search_filings, get_filing, verify_sinks, etc.). Singular/plural distinctions like get_filing vs get_filings are handled cleanly and predictably.

Tool Count5/5

Twelve tools is well-scoped for a filing-data server covering discovery, search, retrieval, statistics, coverage, sink verification, and references. Each tool has an apparent role, and the set avoids both thinness and bloat.

Completeness5/5

For a read-only HKEX filing server, the surface covers the full exploration lifecycle: schema discovery, sink inspection, company lookup, metadata search, full-text search, single/bulk filing retrieval, statistics, coverage checks, sink parity, and graph references. No obvious core read operation is missing.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    SEC EDGAR filing MCP for equity research agents: search 10-K/10-Q/8-K with CompanyFacts metrics, preview a free sample, and purchase full structured JSON via x402 USDC on Polygon. Public endpoint on xpay.tools.
    3
    2
    MIT
  • F
    license
    A
    quality
    C
    maintenance
    A read-only MCP server that enables AI assistants to search files, list directories, retrieve system info, and get file metadata on the local file system.
    4
    -