postgres-mcp-hardened
It is a hardened read-only PostgreSQL MCP server for safely inspecting and analyzing databases through an AI agent.
Query read-only SQL — run validated
SELECT/WITH/VALUES/EXPLAIN/SHOWstatements with row limits, offsets, timeouts, cost guards, and automatic refusal of writes.Explain queries — get PostgreSQL execution plans, optionally with real run times and buffer usage.
Explore schema — list schemas/tables/views, describe columns, types, nullability, defaults, primary/foreign keys, and comments.
Assess database health — cache hit ratio, connection usage, long-running statements, abandoned transactions, vacuum backlog, invalid indexes, sequence ceilings, replication lag, and unanalyzed tables.
Find slow statements — use
top_queries(viapg_stat_statements) and drill into the worst ones withexplain_query.Check security posture — see whether the connected role can write, bypass RLS, or read files; whether transport/auth/audit are sound; and what command fixes issues.
Simulate indexes — test hypothetical indexes with hypopg without creating anything, comparing plan/cost and whether the planner uses it.
Analyze existing indexes — spot unused/duplicate indexes and tables where a seq scan suggests an index would pay off.
Multi-database and resource browsing — serve several databases from one server, with schema/resources exposed for MCP clients.
Operational safeguards — optional OAuth/bearer auth, tamper-evident audit log, column redaction, schema/table allowlists, rate limits, and a strict startup that rejects misconfiguration.
Provides a hardened read-only interface to PostgreSQL databases, allowing AI agents to run validated SQL queries with timeouts, cost limits, and an audit trail.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@postgres-mcp-hardenedQuery the database for the top 5 products by sales"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
postgres-mcp-hardened
🚧 Version 0.1.10 — the rest of an outside review, and a fix that was too blunt
Published: binaries for five platforms with checksums, Sigstore signatures and build provenance;
.mcpbbundles for one-click install; an image onghcr.iofor amd64 and arm64; a package on npm; and an entry in the official MCP registry.0.1.8 and 0.1.9 closed six bypasses that four independent reviewers found in 0.1.7, none of them found by us. Two days of our own adversarial work had come back mostly clean the day before. Passing the tests you thought to write is not the same as looking. The one that mattered most needs no privileges at all: with a column redacted, a join on it through
USINGanswered whether a given value was present, which is a complete equality oracle against the least-privilege reader this project tells you to configure.0.1.10 closes the three findings that were left open. With
MCP_ALLOW_SCHEMAS=public,SELECT * FROM secret.salarieswas refused whiledescribe_tablehanded over every column, type and default of that same table. An ordinary connection string carrying nosslmodeused the driver'sprefer, which sends everything in the clear whenever the server declines TLS, and anyone on the wire can make it decline. AndMCP_SSLROOTCERTadded a private certificate authority to 242 public ones instead of replacing them, so an operator who believed they had pinned trust to their own issuer had not.The first version of that TLS fix was wrong in a way worth reading about. It asked "is this loopback" and demanded TLS from everything else, which took down six PostgreSQL version jobs, the conformance run, the container check and the adversarial corpus in a single push. All of them reach the database the way ordinary deployments do:
postgres://user:pass@postgres:5432/db, a service name on a private network. Docker Compose and Kubernetes are not the public internet, and a server that demands TLS from a container on a bridge network is a server people switch off entirely. The question is now whether somebody untrusted can sit on the wire, not whether the address is loopback, and a real socket to a PostgreSQL reachable only by service name is in the test suite, because a rule about networks should be tested over a network.A resource limit is documented and not solved, in
THREAT_MODEL.md: 49 bytes of SQL make PostgreSQL fold a constant into 5.9 GB of backend memory during planning, and a five secondstatement_timeoutdoes not stop it. The obvious shapes are refused; the general problem is upstream of anything this server can do.Every change here was reproduced against a running server before being fixed, and each is in
CHANGELOG.mdwith the query.Everything here is 0.1.x because nobody outside this project has run it against their own data.
The official Postgres MCP server was deprecated in 2024 and still gets 391k downloads a month. Its entire defence is one database-level read-only transaction — and that alone does not stop every write. This is a maintained Rust replacement with defence in depth.
A drop-in Model Context Protocol server that lets an AI agent query PostgreSQL — read-only, enforced at the database level, with real SQL validation, timeouts, cost limits, OAuth 2.1, and an audit trail. Speaks Streamable HTTP and stdio, and negotiates the MCP revision: 2026-07-28 (current, and the default since upstream released it on 2026-08-03), 2025-11-25, and 2025-06-18 — what most shipping clients still speak today. A client asks for what it knows; it is not negotiated down.
Try to break it — one command, no database
The read-only guard has an offline mode. Hand it a statement and it says what it decided: no database, no configuration, nothing installed permanently.
npx postgres-mcp-hardened --validate "/* comment */ DROP TABLE users"
# REJECT: non-read-only statement: Drop
npx postgres-mcp-hardened --validate "SELECT 1; DROP TABLE users"
# REJECT: multiple statements are forbidden
npx postgres-mcp-hardened --validate "WITH d AS (DELETE FROM t RETURNING *) SELECT * FROM d"
# REJECT: non-read-only statement: non-read-only query (CTE / SELECT INTO / FOR UPDATE)
npx postgres-mcp-hardened --validate "SELECT * FROM orders WHERE id = 1"
# ALLOWIf something that writes comes back ALLOW, that is the most valuable thing anyone can send us.
It needs no working exploit and no write-up — one line of SQL and "this should not be allowed" is a
complete report. Anything that gets past the guard goes through SECURITY.md;
everything else is an ordinary issue, and the bar for opening one is this looks wrong to me, not
I am certain.
The fuzzer is deterministic and prints its seed, so whatever it finds reproduces on a machine that has never seen yours — a million mutations take about a minute:
npx postgres-mcp-hardened --fuzz 1000000
# fuzz: 1000000 iterations, seed 1592594996, slowest validation 8 ms
# RESULT: 0 invariant violationsFor the whole thing against a real database, docker compose -f examples/docker-compose.yml up -d
brings up PostgreSQL with sample data and the server in front of it, connecting as a role that holds
SELECT and nothing else.
Every bypass found so far lives in the MUST_REJECT corpus in src/validate.rs and runs on every
commit, recorded with what it cost rather than tidied away. Yours would join them.
Related MCP server: PostgreSQL MCP Server
Why
@modelcontextprotocol/server-postgres is deprecated on npm (last publish December 2024) and
still sees 475,790 downloads in the 30 days to 9 August 2026. Credit where it is due: its approach is not naive — it
wraps each query in BEGIN TRANSACTION READ ONLY and always ROLLBACKs, which is a real defence
and one this server now adopts as well.
The problem is that it is the only defence, and it is not complete:
A read-only transaction does not block every write, and a rollback does not undo everything it lets through. Two separate facts, and the second is the one that matters.
gin_clean_pending_list()runs insideSET TRANSACTION READ ONLYand its work survives the rollback: an index with 25 pending pages has 0 after the transaction is rolled back.pg_backup_start()puts the session into backup state, survivesDISCARD ALL, and with the defaultfast => falsewaits for a spread checkpoint while forcingfull_page_writeson, which is a real cost on a busy server.pg_import_system_collations()also executes without raisingSQLSTATE 25006, but be careful how much weight you put on it: that one IS undone by a rollback, so against a server that always rolls back it is a curiosity rather than a bypass.Reproduce it, but read the two preconditions first, because without them you will see a zero or an error and conclude we made this up. All three need superuser or ownership of the object. And the import only restores collations that are missing, so something has to be removed first:
-- as superuser, and note these are three separate transactions: a statement that errors -- inside a block aborts the whole block, so they cannot be run as one. DELETE FROM pg_collation WHERE oid IN (SELECT oid FROM pg_collation ORDER BY oid DESC LIMIT 200); BEGIN READ ONLY; DELETE FROM pg_collation WHERE collname LIKE 'zu%'; -- ERROR: cannot execute DELETE ... ROLLBACK; BEGIN READ ONLY; SELECT pg_import_system_collations('pg_catalog'); -- 200, no error COMMIT; -- and now the rows are thereBoth are writes, both are inside a read-only transaction, and one is refused while the other is not. That asymmetry is why this server does not treat the transaction as its only defence. It is also why the role matters more than any of this: every example above needs privileges a least-privilege reader does not have, and this server refuses to start as a network listener when the role it was given can write. What it cannot control is which connection string somebody pastes into a client config, and the usual answer is whichever one they already had.
No statement timeout, no cost guard, no row limit — one query can run until the server gives up.
No authentication, no audit trail, no handling of prompt injection through returned row data.
One source file of 143 lines, unmaintained since December 2024, no test suite.
This server keeps the rollback, adds AST validation in front of it, and adds the operational layers the original never had.
postgres-mcp-hardened vs the archived original
archived | postgres-mcp-hardened | |
Read-only enforcement |
| AST validation (sqlparser) plus the same read-only transaction and rollback, plus a denylist for functions that write despite it |
Multi-statement / | reaches the database and is stopped only by the transaction | rejected by the parser, before it reaches the database |
Statement timeout | none |
|
Runaway / expensive queries | run unbounded |
|
Prompt injection via row data | raw output | wrapped |
Error messages | leak schema ( | structured, non-leaking, actionable |
Auth | none | OAuth 2.1 (RS256 JWT, scope + audience + issuer) |
Audit | none | tamper-evident hash-chained log |
Schema as MCP resources | ✅ | ✅ — plus comments, primary and foreign keys |
Tests / CI | none | unit + end-to-end suites against live PostgreSQL, a deterministic fuzz harness, conformance driven by the official MCP SDK, clippy + |
Transport | stdio / deprecated SSE | Streamable HTTP + stdio |
Maintained | ❌ deprecated since 2024 | ✅ |
Install
Five ways in, in the order most people want them.
One click, for a client that accepts .mcpb bundles: download
postgres-mcp-hardened-<your-platform>.mcpb from the
latest release and open it. The
bundle asks for the connection string and stores it in the OS keychain rather than in a plain-text
config file. Nothing to install, nothing to edit.
Through npm — shortest, and the one your MCP client config can point at directly. There is no Node runtime involved at run time: the package is a launcher that fetches the native binary for your platform and verifies its checksum before running it.
{
"mcpServers": {
"postgres": {
"command": "npx",
"args": ["-y", "postgres-mcp-hardened", "--stdio"],
"env": { "DATABASE_URL": "postgres://readonly_user:PASSWORD@localhost:5432/mydb" }
}
}
}The connection string goes in env, not in args, on purpose: arguments show up in ps output and
in shell history on a shared machine, and a database password does not belong there.
A binary from the releases page — one file, nothing to keep up to date, and the option to take if
your machine has no Node at all. (Not a static binary, as this page claimed until 0.1.7: the
-gnu and macOS targets link the system C library like any other native program. There is simply
nothing to install alongside it.) Every release carries builds for Linux, macOS and Windows
on x86-64 and arm64, each with a checksum and a signature; verifying them is the next section.
As a container, if that is how you run things. The image is distroless and runs as a non-root user, and the same signatures cover it as cover the binaries.
docker run --rm -p 127.0.0.1:8080:8080 --memory=512m \
-e DATABASE_URL="postgres://readonly_user:PASSWORD@db-host:5432/mydb" \
-e MCP_ADDR=0.0.0.0:8080 \
-e MCP_BEARER_TOKEN="$(openssl rand -hex 32)" \
ghcr.io/eszetael/postgres-mcp-hardened:latest--memory is not decoration. The server idles at 7.7 MB and a normal request costs single-digit
megabytes, but a caller can write SELECT repeat('x', 100000000) and drive peak memory to 400 MB —
not through the result, which stays bounded at 300 bytes, but through the cost guard's own
EXPLAIN, which PostgreSQL fills with the constant it folded while planning. That is a named
residual risk in THREAT_MODEL.md, with the three repairs that were tried and
what each one broke. Until it is closed, the memory limit is the thing that holds, so set one:
--memory here, MemoryMax= under systemd.
MCP_ADDR must bind 0.0.0.0 and not 127.0.0.1, or the server listens on an interface that only
exists inside the container and the published port answers nothing. The other easy one: localhost
in DATABASE_URL means the container, not your machine, so a PostgreSQL running on the host needs
host.docker.internal (Docker Desktop) or the host's address on the bridge (172.17.0.1 by default
on Linux). Both of these were walked end to end against the published image before being written
here, including that a read returns rows and DROP TABLE comes back as
-32602 non-read-only statement: Drop.
From source — cargo build --release in a clone. Not cargo install: this crate is not on
crates.io, and an instruction that fails is worse than one that is missing.
Checking what you downloaded
Every released binary is signed with Sigstore keyless signing — there
is no private key for us to lose, and the certificate names the workflow, repository and tag that
produced the file. Each artefact ships with a .sig and a .pem beside it:
F=postgres-mcp-hardened-x86_64-unknown-linux-gnu.tar.gz
cosign verify-blob "$F" --bundle "$F.bundle" \
--certificate-identity-regexp '^https://github.com/Eszetael/postgres-mcp-hardened/' \
--certificate-oidc-issuer https://token.actions.githubusercontent.comPin the identity, not just the signature. Without --certificate-identity-regexp and
--certificate-oidc-issuer the check answers "somebody signed this", which is not the question.
A verified certificate names the workflow, the repository and the tag that built the file — you
can read it with base64 -d "$F.pem" | openssl x509 -noout -text (cosign writes the certificate
base64-encoded, which surprises people who try openssl on it directly).
Older cosign builds predate --bundle; separate .sig and .pem files are published alongside
for them, used as --signature "$F.sig" --certificate "$F.pem". Current cosign marks those flags
deprecated, so prefer the bundle.
Public releases additionally carry SLSA build provenance, verifiable with
gh attestation verify <file> --repo Eszetael/postgres-mcp-hardened.
Use it in Claude Desktop / Cursor (stdio)
{
"mcpServers": {
"postgres": {
"command": "postgres-mcp-hardened",
"args": ["--stdio"],
"env": { "DATABASE_URL": "postgres://readonly_user:YOUR_PASSWORD@localhost:5432/mydb" }
}
}
}Or run it as a remote server (Streamable HTTP)
DATABASE_URL="postgres://readonly_user:YOUR_PASSWORD@host:5432/mydb" \
MCP_ADDR="0.0.0.0:8080" \
postgres-mcp-hardened
# POST /mcp · GET /health · GET /ready · GET /metricsTLS: connections to PostgreSQL are encrypted whenever the server supports it, and
sslmode=require,verify-caandverify-fullare all accepted (the certificate chain and the hostname are always verified, sorequirebehaves likeverify-full) — so managed Postgres (RDS, Supabase, Neon, Render) works out of the box. Certificates and host names are always verified — verified (acceptance: "a certificate naming another host is refused, by name"); for a private CA, pointMCP_SSLROOTCERTat the PEM bundle. There is no "trust anything" switch.Tip: point
DATABASE_URLat a least-privilege read-only role. The server enforces read-only itself, but a scoped DB role is defense-in-depth.
Or run it on a container platform (Apify Standby)
The server needs no code changes to run as an Apify Actor in Standby mode. It reads the port the
platform assigns from ACTOR_WEB_SERVER_PORT and binds 0.0.0.0 there — that port wins over
MCP_ADDR, loudly, on stderr, because binding anywhere else means the run is never marked ready and
the failure looks like a mysterious timeout. GET / answers the platform's readiness probe
(x-apify-container-server-readiness-probe) without touching the database: container readiness is
not database readiness, and a probe that waits on a busy pool turns a slow database into a container
that never starts.
Endpoint | Method | Purpose |
| POST | the MCP endpoint (Streamable HTTP). |
| GET | readiness probe; otherwise a signpost naming the real endpoint |
| GET | the process is alive |
| GET | the process and a database connection are available |
| GET | counters (needs |
| GET | what a registry reads: revisions, transports, tools |
Input is a JSON-RPC request in the POST body — initialize, tools/list, tools/call,
resources/list, resources/read, server/discover. Output is a JSON-RPC response; from
2025-11-25 a refused statement comes back as a tool execution error (isError: true) with the
reason in the content, so the model can rewrite the query. tools/list is the authoritative
description of every argument.
Authentication there is the platform's, not ours. Apify checks the caller's token before routing
to the container, so the server does not additionally demand MCP_BEARER_TOKEN — requiring a second
secret would mean an agent that finds this server cannot call it. That exemption is narrow: it needs
both APIFY_IS_AT_HOME and ACTOR_WEB_SERVER_PORT, one alone changes nothing, and the server
card then reports "type": "apify-platform" rather than claiming a lock we do not hold. Everywhere
else the server still refuses to start on a network address with no authentication. Set
MCP_BEARER_TOKEN as well if you want a second lock on the same door.
The other start gate is unchanged and matters more here: a role that can write is refused a network
listener. Point DATABASE_URL at a read-only role — --print-setup-sql writes the statements.
Migrating from the deprecated server
The most-discussed problems reported against @modelcontextprotocol/server-postgres were
reproduced against this server; here is how each behaves:
What people reported | Here |
Two instances (prod + dev) are indistinguishable, the client picks one | Resource URIs carry the database name ( |
One database per instance, because the connection string is a command-line argument |
|
| The error says the server requires TLS and names the fix ( |
Read-only bypassed by injecting | Rejected — the multi-statement gate works on tokens, before the parser, and |
| A single native binary; no Node, no npx, no |
Hangs indefinitely against RDS with no output or error | Bounded: an unreachable host answers in ~8 s with the reason, never silently |
| Point |
Connection string only as a command-line argument |
|
| The error says which characters to percent-encode, and how |
Partition children flood the table and resource lists | Hidden by default; |
|
|
No row limit — one query floods the context | Auto- |
Testing
Beyond unit tests, the repository carries two harnesses that run in CI on every change:
--fuzz— a deterministic fuzzer that mutates a corpus of known writes with transformations that do not change SQL meaning (comments, case, dollar-quoting, invisible Unicode, parentheses) and asserts that none of them ever becomes an allowed statement.tests/acceptance.sh— an end-to-end suite that starts its own PostgreSQL and checks 317 behaviours: every write-bypass reported against the deprecated server (including theCOMMIT/ENDinjection), truthful results, schema introspection, protocol conformance, configuration mistakes failing loudly, audit tamper detection, fair use under load, and multi-database deployments.
Every reported problem, answered
docs/COMMUNITY_ISSUES.md is the complete ledger: every problem
reported against the deprecated server and every open issue against the maintained alternatives,
each with what happens here — including the handful we could not fix in code, said plainly.
What we learned from the alternatives
Every server in this space has an issue tracker, and those trackers are a map of what goes wrong. The ones we deliberately built against:
A published image that lags the code. The most-supported open complaint against the leading alternative. Our container is built and pushed from the same tag that produces the binaries, so it cannot drift.
A hardcoded query timeout. Also among their most requested settings.
MCP_STATEMENT_TIMEOUTis configurable and validated at startup.Unrestricted access by default. Some servers default to read/write and rely on the operator to restrict it. This one has no write path at all.
Credentials in the client configuration.
MCP_PASSWORD_FILEkeeps the password out of it.Tables in a non-default schema silently not found.
MCP_SEARCH_PATHfixes the lookup, and the tools take an explicitschemaanyway.Deprecated transport. HTTP+SSE was replaced by Streamable HTTP in 2025-03-26, three revisions ago (this page said 2025-06-18 until 0.1.7, which was wrong by one revision; the specification's own changelog for 2025-03-26 records the replacement). We speak the current transport.
Troubleshooting
Answers to the questions people actually asked about the deprecated server, so nobody has to open an issue to find them.
spawn npx ENOENT / "which Node version do I need?" — none. This is a single native binary, with no runtime to install beside it.
Download it from the releases page and point your client
at the file. There is no node_modules, no npx, nothing to keep up to date.
"The server starts but nothing is listening on a port." — that is stdio mode, which is correct
for Claude Desktop and Cursor: the client talks to the process over its standard input and output,
not over a socket. If you want a network endpoint, start it without --stdio; it then prints
MCP HTTP listening on http://… and speaks Streamable HTTP.
"Can my client on another machine reach the database?" — yes: run the server next to the
database in HTTP mode, expose it, and enable OAuth (JWT_PUBKEY_PEM, JWT_AUD, JWT_ISS). The
database credentials then never leave the host the server runs on.
"Could not attach to MCP server." — the process exited before the handshake. Run the same command in a terminal: a configuration mistake prints its reason and exits with status 2 rather than dying quietly, and a connection problem is reported on the first query with the cause.
self-signed certificate in certificate chain / unable to verify the first certificate —
your provider uses a private CA (Supabase, GCP and RDS all do). Download their CA bundle and set
MCP_SSLROOTCERT to it. The error message names the step for your provider. We do not offer a
"trust anything" switch.
Managed providers
This server always verifies the database certificate, including with sslmode=require. That is
a deliberate deviation from libpq, where require encrypts without verifying and a machine in the
middle can therefore read and rewrite every query and result without anyone noticing. The cost of
being strict is that a provider with a private CA needs one extra step; the cost of being lax is
that you never find out. If you disagree with the trade-off, verify-full with the bundle below is
the same amount of work and leaves no doubt either way.
Provider | What to expect |
Supabase | Private CA. Dashboard → Project Settings → Database → SSL configuration → download the certificate, then set |
Amazon RDS / Aurora | Private CA: |
Google Cloud SQL | Private CA: Connections → Security → |
Azure Database for PostgreSQL | Public CA — nothing to download. Azure rotated its root to DigiCert Global Root G2 during Q1 2026; we ship the Mozilla root store, so the rotation needs nothing from you. |
Neon | Public CA (ISRG Root X1, Let's Encrypt) — nothing to download. Pooled and direct endpoints both work. |
DigitalOcean | Private CA: download the certificate from the cluster's Overview page. |
Not sure which case you are in? Ask the server itself, before configuring anything:
echo | openssl s_client -starttls postgres -connect YOUR_HOST:5432 2>/dev/null \
| openssl x509 -noout -issuerA well-known issuer (DigiCert, ISRG, Google Trust Services) means it will just work; anything naming your provider means you need their bundle.
With a connection pooler (Supavisor, PgBouncer) in transaction mode, note that this server sets
statement_timeout and idle_in_transaction_session_timeout per session and runs every query in an
explicit read-only transaction. Both are compatible with transaction pooling; session-level
SET outside a transaction is not, which is why we do neither.
no pg_hba.conf entry … no encryption — the server accepts only TLS connections for that
host and user. Add ?sslmode=require to the connection string.
INVALID_URL / invalid connection string — a password containing @, :, /, # or ?
must be percent-encoded (@ → %40, : → %3A, / → %2F, # → %23).
"My table has hundreds of partitions and the list is unusable." — partition children are hidden
by default; the parent is listed. Set MCP_SHOW_PARTITIONS=1 if you need them.
"I need production and staging at the same time." — either run one server per database (they
are distinguishable: set MCP_SERVER_LABEL), or configure both in one server with
MCP_DATABASE_URLS and pass database in the tool arguments.
Resources
Every table and view is exposed as an MCP resource (postgres:///<schema>/<table>/schema), so a
client can browse the schema without issuing a query — the same capability the deprecated server
offered, plus column comments, primary keys and foreign keys in the payload.
When the database has not answered, resources/list returns an empty list with the reason in
_meta rather than a protocol error. A catalogue inspecting this server starts it with no
database at all and calls resources/list straight after initialize; answering that with an error
reads as a server that does not work. The empty list is not a claim that there are no tables —
initialize says the database has not answered, security_posture gives the detail, and the reason
travels with the list itself. A database that does answer and refuses is still an error, because
reporting "no resources" for a missing privilege is the silent failure this server exists to avoid.
verified (acceptance: "a host can inspect the server with no database and mcp-proxy in front")
Tools
explain_query— the execution plan; withanalyzeit runs the statement and reports real timings and buffer usage, which is safe here because the statement is validated read-only and runs inside a transaction that is always rolled back. The plan comes with asummary: which node spent the time (self time, not inclusive), and where the planner's row estimate was furthest from reality — because a bad estimate is usually why the plan is bad.database_health— cache hit ratio, connections (this database and the cluster), the longest running statement and the longest abandoned transaction as separate figures, vacuum backlog, invalid indexes, sequences near their ceiling, replication lag, tables that have never been analysed (no planner statistics — the usual reason a database looks healthy and runs slowly), and the window the counters cover. Anything the role cannot see is declared rather than returned as a confident zero.analyze_indexes— unused indexes, duplicates, and tables scanned sequentially often enough that an index would pay off.top_queries— the heaviest statements, frompg_stat_statements.security_posture— what this deployment is actually able to do to your database, asked of PostgreSQL rather than assumed: whether the role can write, bypass row-level security or reach server files; whether the transport is authenticated; whether the audit chain is keyed; whether the connection is encrypted. Returns a grade — the worst finding, never an average — and, for anything wrong, the command that fixes it. The same summary reaches the model throughinitialize, because under stdio nobody sees stderr and the agent is the only messenger the operator has.query— run a read-only SQL query (validated, auto-LIMIT, cost-guarded). The response states what it did:returnedRows,appliedLimit,truncated, plusrequestedLimitwhen a larger request was capped at the 10000-row maximum,offsetwhen paging, andredactedColumnswhen masking is configured — so an agent never has to guess whether it received the whole answer.list_schemas,list_tables,describe_table— progressive schema discovery (parameterized, injection-safe).describe_tablereturns the schema comments (COMMENT ON TABLE/COLUMN), primary keys, foreign keys and defaults, so the agent reads what a column means and what it points at instead of guessing from its name — and a missing table is an error, not an empty column list.
Protocol revisions
The server answers initialize with the revision the client asked for when it implements it, and
with its newest otherwise. Over HTTP the revision comes from the MCP-Protocol-Version header, per
request — one client's negotiation cannot change the contract another client is served under.
If a request carries no header, the server does not fall straight back to the oldest contract. It
reads the revision this session agreed on at initialize, which is what the transport
specification asks for: the default applies only "if the server does not receive an
MCP-Protocol-Version header, and has no other way to identify the version — for example, by
relying on the protocol version negotiated during initialization". A session is that other way, so a
client that negotiated 2025-11-25 and then omitted the header keeps the contract it agreed to
rather than being silently demoted.
A header we cannot parse is a different matter from a header that is absent, and the specification
is explicit about it: "If the server receives a request with an invalid or unsupported
MCP-Protocol-Version, it MUST respond with 400 Bad Request." A version we do not implement —
2025-03-26 or not-a-date — is refused with 400
and the list of revisions we do speak, rather than served under a contract the client never agreed
to. Falling back is for silence, not for disagreement.
Only a request with neither a header nor a session falls back, and it falls back to 2025-06-18 —
the oldest revision this server implements — rather than the 2025-03-26 the specification names.
That revision is not implemented here, and answering under a contract the server cannot honour would
be worse than answering under the oldest one it can.
The difference that matters is where a refusal goes. Under 2025-06-18 "this statement is not
read-only" was a JSON-RPC error: the client saw a broken call and the model often never saw the
reason. From 2025-11-25 (SEP-1303) it arrives as a tool execution error — isError: true with the
reason in the content — so the model rewrites the query instead of handing the user a failure. What
does not change is the audit: the refusal is recorded by the code that refuses, and the acceptance
suite asserts both halves together, so friendlier errors can never quietly mean a quieter log.
Mcp-Method and Mcp-Name are held to agreement, not presence. The draft requires them; earlier
revisions do not, and demanding them would break every client shipping today. But a gateway that
routes or authorises on Mcp-Method while the server executes the body has decided about a
different request than the one that runs — and that is true whatever revision is in force. So a
header that is present must match the body under every revision, while a client that sends none is
untouched.
Protocol failures stay protocol failures. A malformed envelope, an unknown method or a missing token is not something a model can fix by rewriting SQL, and a client's error handling expects those where they have always been.
The next revision, before it lands
2026-07-28 is the largest break MCP has had: no initialize, no session header, no ping. That
identifier comes from LATEST_PROTOCOL_VERSION in the draft schema, and it is not a release date —
MCP names a revision for the last date a backwards-incompatible change was made, so it describes the
draft's history rather than a schedule. Every request carries its own protocol version in _meta,
and a new server/discover replaces the handshake. We implemented it early behind a switch, because
a draft moves and a server advertising support for a moving target will be wrong in public. Upstream
cut schema/2026-07-28 on 2026-08-03 — the released schema differs from the draft we had verified
against in four documentation URLs and nothing else — so the switch is gone and this is what the
server speaks by default. Clients on 2025-11-25 and 2025-06-18 are answered as before.
server/discover answers under every revision, because the specification
expects clients to use it as a backwards-compatibility probe — which only works if older servers
answer it. Ours answers with the revisions we speak and, in _meta, the full security posture. That
is deliberate: a client can learn it is talking to a server connected as a superuser before it
sends a query, as structured data rather than prose a model has to notice.
Two of the draft's rules are security controls here, not formalities. Mcp-Method and Mcp-Name
must agree with the request body, and we refuse the mismatch (-32020) — the headers exist so a
gateway can route and authorise without parsing the body, and if header and body may disagree, then
the thing that authorised and the thing that executes saw two different requests. And a client
stating a version we do not implement is told so (-32022) rather than quietly served under a
contract it never agreed to.
What the safety costs
Measured, not asserted: tests/bench/, against the pg driver running the same query on the same
machine (PostgreSQL 18.6 in Docker, 50k-row table, 300 sequential requests, rate limit off).
Re-measured 2026-08-17 on a shared VPS at load average 2.6 — median of two runs:
query | driver | this server | difference |
point lookup | 0.43 ms | 5.9 ms | +5.5 ms |
small scan | 1.0 ms | 8.3 ms | +7.3 ms |
aggregate | 4.7 ms | 11.3 ms | +6.6 ms |
An earlier table here said +3.6/+5.2/+3.7 and "about 4 ms". Those came from a quieter machine, and
the driver floor moved with them — 0.28 ms against today's 0.43 for the same lookup — so it was the
hardware talking, not the code. Two things were checked before changing the number, because the
obvious suspect was our own build: the 0.1.6 binary, built before lto was switched on, measures
+5.4/+7.8/+5.9 on this machine within the same hour. Identical. The release profile halved the
binary and cost nothing here.
Expect 5 to 8 ms per query, and treat any single figure on this page as a reading from one machine on one day. The shape matters more than the size: the overhead is nearly constant. If the AST validation were the cost, it would grow with the query. It does not. The time goes on round trips — the session is reset, the timeouts and read-only flag are set, a read-only transaction is opened, the cost guard plans the statement, then the query runs and the transaction is rolled back. Five or six exchanges where the driver has one.
That is a deliberate trade and you can see exactly what it buys. For an agent making tens of calls it is invisible; if you are putting this in front of a latency-critical serving path, you are using the wrong tool, and it is not one.
Under concurrency the interesting number is not throughput but what happens past the limits: at 8 concurrent clients it served 373 requests/second and turned away 192 more with "too many requests in flight", which is the in-flight cap doing its job rather than a queue growing until something falls over.
Would this index help? — answered without creating one
The one capability the leading alternative is genuinely known for is index tuning: it can tell you an index would pay off before you build it. It gets there by defaulting to a connection that can create real indexes — safe only if you remembered to restrict it.
simulate_index answers the same question from a connection that cannot write anything.
hypopg registers a hypothetical index in backend memory: the
planner sees it, storage never does, and it is gone when the call returns. You get the plan and cost
with and without, and — separately — whether the planner actually reached for it, because a cost
that barely moves and an index the planner ignored are different answers.
The tool takes a table and a list of columns. Not a CREATE INDEX statement. The definition is
assembled server-side from identifiers the catalogue confirmed exist, quoted by PostgreSQL itself,
so there is no path from a tool argument to arbitrary DDL — a column name carrying SQL dies on the
catalogue lookup, and there is a test that fires exactly that. The numbers are planner estimates:
treat a large improvement as a reason to test the index, not as proof.
Conformance is checked by somebody else's client
Every other test here is our harness talking to our server. If we misread the specification, we
misread it the same way in both halves and everything passes. So CI also drives the server with the
official MCP SDK — the client library the ecosystem uses — over stdio and Streamable HTTP:
handshake, tool listing and schemas, a read, a refused write arriving as a tool execution error
rather than a protocol one, resource listing and reading. A protocol mistake shows up as a client
that cannot talk to us. tests/conformance/.
Working on this
git config core.hooksPath .githooks # once, per clone.githooks/pre-push runs format, clippy, the unit tests and the documentation-claim checks before
anything leaves your machine. It exists because of a specific mistake: a commit went out with a
failing clippy lint, and the first anyone knew of it was a failure email. Note that cargo test
passes on that code — clippy lints are not compiler errors — so "it builds locally" is not the
same answer as "CI will be green".
It deliberately skips the acceptance suite and the PostgreSQL matrix: those need Docker and about
twenty minutes, and a hook people cannot afford to run is a hook people bypass. CI runs everything.
git push --no-verify when you want to see something fail in CI on purpose.
Setting up the role
DATABASE_URL=postgres://admin@host/mydb postgres-mcp-hardened \
--print-setup-sql --role mcp_reader --schemas public --redact ssn,email > setup.sql
# read it, then:
psql -v pw="$(openssl rand -base64 24)" -f setup.sql mydbRun with a connection string and the table and column lists come from the catalogue; without one you get the same document with placeholders. The difference matters most for redaction: the columns to grant back have to be read from the database, because writing them from memory is how a column meant to stay hidden gets handed back.
The output ends with checks that return no rows when it worked, and a reminder that the server itself will tell you what the role can do the moment you point it at the database.
Limiting what the server can reach
MCP_ALLOW_SCHEMAS and MCP_ALLOW_TABLES restrict which relations a query may touch. Either one
turns the allowlist on; schema.* and schema.table both work.
MCP_ALLOW_TABLES='public.customers,public.orders,analytics.*'The check reads the query plan, not the SQL. That is the whole design: the planner has already
applied search_path, resolved every alias, expanded views to base tables, and knows that a CTE
named customers is not the table customers — so WITH customers AS (SELECT 1) SELECT * FROM customers runs and touches nothing, while WITH x AS (SELECT * FROM salaries) SELECT * FROM x
is refused. Reading the statement instead is what lost three rounds of adversarial review.
Two consequences worth knowing before you turn it on:
A partition rides on its parent. You allow
events; PostgreSQL decides which children to read.A view needs its base tables allowed too, because the plan names those. Allow both, and let the database privileges keep the base table unreachable directly — that is the boundary in any case.
pg_catalog and information_schema are outside the surface unless MCP_ALLOW_CATALOG=1: an agent
that can still read the catalog can enumerate exactly what the allowlist was meant to hide. The
schema tools keep working, because they run fixed queries rather than caller SQL.
That covers three routes to the same facts, because for a while it covered only one. A catalogue view
plans to scans over real relations and the plan names them — but pg_settings plans to a single
Function Scan on pg_show_all_settings, naming no relation at all, and current_setting() is a
scalar call that never appears as a scan. Both used to return the server's configuration under an
active allowlist. Functions whose name begins with pg_, plus current_setting, inet_server_addr
and inet_server_port, are now refused alongside the catalogue relations, and MCP_ALLOW_CATALOG=1
opens all of them together. Ordinary set-returning functions — generate_series, jsonb_each,
unnest, regexp_split_to_table — carry no such prefix and are unaffected. current_user,
session_user, current_database and version stay readable on purpose: an agent already knows what
it connected to and as whom.
The corpus of things that got through
Every shape that defeated a control during review lives in tests/adversarial/, with the round in
which it stopped working. It runs on every build, and — because the cases are written against
placeholders rather than our fixture — you can point it at your own database:
ADV_URL='postgres://…' \
ADV_TABLE=people ADV_TABLE2=orders ADV_REDACT_COL=ssn \
./tests/adversarial/run.shADV_TABLE needs the sensitive column, ADV_TABLE2 is any other readable table, ADV_REDACT_COL
is the column to redact. All three matter: this example omitted ADV_TABLE2 until 0.1.7, so it kept
its default of film — a table from the Pagila sample database — and following the instruction
exactly produced three "mismatches" that were only a missing relation. The harness now checks the
three up front and says which one is wrong, because a security corpus that reports a typo as a
failed control teaches you to ignore the failures that matter.
A security claim you can only check by reading our source is a claim you have to take on trust, and this project's own history is the argument against that.
Security model
The full statement of what this server guarantees, what it does not, and which control enforces which promise is in THREAT_MODEL.md — including the controls that have been defeated in review and are therefore described as depth rather than as boundaries.
Encrypted transport: TLS to PostgreSQL via rustls (no OpenSSL in the image), certificate verification always on, private CAs via
MCP_SSLROOTCERT.Read-only, two ways: every statement is parsed with
sqlparserand rejected unless it's aSELECT/WITH/EXPLAIN/SHOW; the DB session is additionally setdefault_transaction_read_only = on.Anti-DoS: enforced
statement_timeout, auto-injectedLIMIT, and anEXPLAIN-based cost guard that rejects expensive plans before they run.Sensitive columns — defence in depth, and honest about it:
MCP_REDACT_COLUMNSmasks values at every depth and refuses to run a query that references those columns, including the ways round it that an adversarial panel actually found — renaming (SELECT password AS pw), wrapping (md5(password)), serialising the whole row (row_to_json(t),t::text,json_agg(t)), and naming the column as a string rather than an identifier (to_jsonb(t) ->> 'password',#>> '{password}',$.password), whole-row wildcards (ROW(t.*)::text) and positional renaming ((SELECT * FROM staff) AS x(c1, …, c9)). It is still name-based filtering, and name-based filtering cannot be a boundary against the whole SQL language — four adversarial rounds each got past it through a shape nobody had listed.The fourth, in 0.1.7, did not find a new shape of name. It went around names altogether:
SELECT get_raw_page('people', 0)hands back 8192 bytes of the table as the disk holds it, and every value on that page is in there, including the redacted one. Demonstrated, not theorised — withMCP_REDACT_COLUMNS=ssn,SELECT ssn FROM peoplewas refused while the raw page came back with the social security numbers in plain ASCII.pageinspectand its relatives are now refused as a category of their own, because "this returns storage rather than columns" is a different problem from "this writes", and telling an operator which one they hit is worth a separate message.The same round found the quieter version of it. PostgreSQL's planner keeps a sample of each column's real values, and
pg_statspublishes them: with 3,000 rows,SELECT * FROM pg_stats WHERE tablename='people'returned{123-45-6789,555-00-1111,987-65-4321}whileSELECT ssn FROM peoplewas refused. That query never names the redacted column, so a name-based rule has nothing to act on. The value-bearing statistics columns —most_common_vals,histogram_bounds,most_common_elems,stavalues1…stavalues5— now join whatever you configure, whenever you configure anything. The rest of the view is untouched:n_distinct,null_fracandcorrelationare what the index advice below is built from and they carry no values, so removing the whole relation would have broken ten columns to fix four.Two more doors turned out to open on the same room. A value too long for its row is stored in a TOAST table, and that table is readable by name:
SELECT chunk_data FROM pg_toast.pg_toast_16384returned the redacted text in the clear.pg_largeobjectis the same thing for large objects — the bytes underneathlo_get. Both are refused now, as relations rather than functions, with a message that says why: they hold physical storage rather than columns. The catalogue that merely describes the database is untouched —pg_tables,pg_stat_activity,pg_largeobject_metadata, and the statistics columns index advice needs.All four say the same thing, and it describes this feature's limit better than any list of patches: a column filter protects columns, so anything that reads underneath columns is outside what it can promise. Raw pages, planner samples, TOAST chunks and large-object bytes are four doors into that space, all four found in a single afternoon by looking for the shape rather than the names. The honest assumption is that there are more, which is why the database role and the read-only transaction are the real boundary and this stays what it says it is: defence in depth.
So the server stops asserting and asks the database: at startup it reports every table where the connected role can still read a redacted column, with the exact statements that fix it, and
MCP_REDACT_REQUIRE_REVOKE=1turns that report into a refusal to run. Note the fix is a table-levelREVOKEfollowed by aGRANTof the columns that stay — a bareREVOKE SELECT (password) ON staffis silently a no-op while the role holds SELECT on the whole table. With column-level grants PostgreSQL then refusesSELECT *on that table, so callers name columns instead;describe_tablelists them and marks the redacted one.Prompt-injection aware: row data is returned inside a
trusted="false"provenance block with delimiters escaped, so a malicious cell can't hijack the agent.It generates the role you should be running as:
--print-setup-sqlwrites the DDL for a role that inherits nothing, bypasses nothing, creates nothing, reads only the relations you name, and — where you have named sensitive columns — has them revoked in the order that actually works. It prints; it never executes. Applying this needs administrative rights, and a tool whose whole identity is "read-only" has no business holding an administrator's password.It will not expose a role that can write: when the listen address is reachable from the network, the server asks PostgreSQL what the connected role is actually allowed to do — superuser,
BYPASSRLS, membership ofpg_write_all_dataand friends, and write privileges on a bounded sample of tables — and refuses to start if the answer is more than "reader", naming each reason and pointing at--print-setup-sql. It refuses an unauthenticated network listener for the same reason. Loopback and stdio are left alone: there the caller is the operator. The overrides (MCP_ALLOW_EXCESSIVE_ROLE,MCP_ALLOW_ANONYMOUS_NETWORK) take the literal valuei-accept-the-riskso they cannot be switched on by a typo, and they are recorded in the audit log. This server enforces read-only itself, but that enforcement is code, and code has been wrong before; a role that cannot write is the part no bug of ours can undo.A browser cannot reach it: a request carrying an
Originis refused with 403 unless the operator listed that origin, and on a loopback listener aHostthat is not localhost is refused too — the shape a DNS-rebinding attack takes when it aims at a database server on your laptop.The audit knows the configuration: the chain opens with a
startuprecord naming the version, the transport and every setting in force, with connection passwords stripped and secrets reduced to fingerprints, plus aconfig_fpan operator can pin across restarts. A log that says what happened but not under which settings cannot answer the first question an incident asks.The audit notices being shortened: a hash chain proves entries were not altered, but a log with its tail cut off is internally consistent — recomputing it finds nothing wrong. Alongside
MCP_AUDIT_LOGthe server therefore keeps<log>.hwm, a one-line record of the last sequence number and hash it wrote, updated only after the entry is durably appended. On start the two are compared, and a disagreement is reported: entries missing from the end, a rewritten last entry, or a log that has gone away entirely. This is not proof of tampering — an unclean shutdown looks the same — but a tamper-evident trail owes you the question, not the verdict. Keep the sidecar with the log when you archive or move it; deleting it only loses the truncation check, never an entry. The offline verifier is unchanged and still needs an external anchor:--verify-audit <file> --expect-last <hash>. verified (acceptance: "a shortened log is noticed at startup, without any external anchor")A wrong setting is fatal, not merely wrong: an unparsable listen address, an audit file that cannot be written,
sslmode=disableto a database on another machine, a metrics token that is also the database credential, a boolean speltyes— each used to be accepted and quietly do something other than what was meant. Startup now stops and names the setting.A misspelt setting is fatal — somebody else's setting is not:
MCP_REDACT_COLUMN(singular) used to start the server with redaction quietly switched off, so a near miss of a real setting still stops startup and names the intended spelling. A name that resembles nothing we define was set by another program sharing the environment: it is reported and ignored.mcp-proxy, which every catalogue puts in front of a server to inspect it, exportsMCP_PROXY_DEBUG— until 0.1.6 that one variable made this server exit before reading a request.MCP_X_*remains reserved for the operator's own use. verified (acceptance: "a misspelling is still fatal")No schema leaks: database errors are mapped to structured, actionable messages that never echo table/column names.
OAuth 2.1: optional RS256 bearer-token validation (signature,
exp,aud,iss) with scope enforcement; disabled when unconfigured for local/self-host use.Audit: every tool decision is logged as a tamper-evident, hash-chained JSON line (no raw SQL).
Supply chain: dependency licences, sources and advisories enforced in CI (
cargo deny,cargo audit); a CycloneDX SBOM is attached to every release.Runtime: ships as a distroless, non-root container — 14.8 MB to download, 41 MB on disk for
linux/amd64at 0.1.6, built and smoke-tested in CI. Both numbers, because a single one is always the flattering one:docker imagesshows the second, your bandwidth pays the first.
Footprint
Measured on an ordinary VPS against a 16k-row sample database, so you can check the "written in Rust" claim rather than take it:
Resident memory, idle | 7.7 MB — median of five separate starts, all within 0.1 MB of each other |
Resident memory, after 200 requests | 9.4 MB, and flat afterwards |
Median request latency | ~8 ms — including the |
Start to first validated statement | 5 ms — median of five |
Binary | 9.2 MB (linux x86_64, 0.1.7 onward). Nothing to install alongside it — no Node, no Python, no shared library we ship. It is not statically linked: like any |
Container image | 12.6 MB compressed, 31.8 MB unpacked (linux/amd64, 0.1.7), distroless, non-root |
Two lines here were wrong until 0.1.7. Idle memory said 5.2 MB and measures 7.7 — five starts under
identical conditions landed within 0.1 MB of one another, so the old figure is not noise, it is a
different measurement whose method was not written down. The binary line said "11 MB, static", and the file people actually
downloaded was 18.9 MB and dynamically linked. The size was never measured on a release build,
because this crate had no [profile.release] at all, so more than five megabytes of debug symbols
shipped to every user. Setting strip, lto and codegen-units = 1 took it to 9.2 MB. The word
"static" was simply not true of any of the five targets we publish, none of which is a musl build.
The image shrank with the binary, from 14.8 MB compressed in 0.1.6 to 12.6 MB — both figures read
from the registry manifest of the published image rather than from a local build, because a local
build is not what anybody pulls.
A twelve-minute soak of mixed traffic (reads, refusals, errors, aborted requests, session churn, unauthenticated requests) served 51,499 requests and ended with the same 15 open file descriptors it started with. Resident memory went from 8.1 MB to 11.2 MB, and the shape of that is the interesting part: 8.1 to 10.7 happened inside the first 400 requests, and the remaining 51,000 added 0.5 MB in a curve that flattened as it went. That is an allocator settling, not a leak. This page used to say memory stayed "flat", which was true of everything after the first few seconds and not true of the number, so here is the number.
Reproduce it with tests/soak.sh rather than believing the paragraph.
Configuration
Env | Purpose |
| PostgreSQL connection string (use a read-only role) |
| HTTP listen address (default |
| reject queries whose |
| enable OAuth 2.1 token validation (omit to disable auth); the key may be the PEM text or a path to a PEM file |
| path to the append-only audit log (hash-chained); verify with |
| key that turns the audit chain into HMAC-SHA256 — keep it off the host so the log cannot be rewritten (a trailing newline in the file is ignored) |
| comma-separated previous keys, so a log that survived a key rotation still verifies. verified (acceptance: "a chain spanning a key rotation verifies with both keys") |
| columns to keep out of results, e.g. |
| shared token required on every request, for deployments without an identity provider. Ignored when OAuth is configured — accepting it as an alternative would give its holder full scope and leave the audit with no identity |
| query time limit (PostgreSQL interval, default |
| schemas to search when a table name is unqualified, e.g. |
| read the database password from a file instead of putting it in the connection string |
| several databases from one server: |
| name shown in the client UI, e.g. |
|
|
| comma-separated catalog functions to permit that we do not know to be read-only |
| path to a PEM CA bundle for TLS to PostgreSQL (e.g. the AWS RDS bundle); system and Mozilla roots are trusted by default |
| max concurrent requests from one client (default 4; |
| per-client request rate limit (default 120/min; |
| the same limit for stdio (default 600/min): an agent exploring a schema legitimately makes dozens of calls a minute, but a runaway loop against a production database is still what a DBA fears most |
| a name for this client in the audit log over stdio, e.g. |
| burst allowance for that limit (default |
| token required on |
|
|
|
|
| database slots kept for authenticated traffic so an anonymous flood cannot take the pool (default: a quarter) |
| this server's public base URL, used in the OAuth discovery metadata |
| authorization server URLs advertised in that metadata |
| Retired. It gated |
| schemas a query may reach, e.g. |
| relations a query may reach, e.g. |
|
|
| set to |
| set to |
| set to |
| browser origins permitted to call this server, e.g. |
| extra |
| development only: makes |
| set to |
License
MIT — see LICENSE. The core is MIT and stays that way.
Commercial and team use
Pointing this at a production database inside an organisation raises questions the MIT core does not answer: policy bound to an identity rather than to a scope, an audit log shipped somewhere it cannot be quietly edited, evidence a compliance reviewer will accept, a deployment somebody is on the hook for.
A team edition covering those is being scoped right now, and what goes into it is not decided. If that is what your organisation needs, write to eskulapstudio@gmail.com and say what would have to be in it. There is nothing to buy yet — the answers are what decides whether it gets built at all, and in what order.
Available Tools
10 toolsanalyze_indexesIndex findingsARead-onlyIdempotent
Indexes nobody uses, genuine duplicates, and tables scanned sequentially often enough that an index would likely pay off. Counters come from pg_stat_*, which reset with the server — read them after real traffic, not after a restart. Primary-key and unique indexes are excluded from the unused list on purpose: they earn their keep by enforcing a constraint.
| Name | Required | Description | Default |
|---|---|---|---|
| schema | No | public | |
| database | No | which configured database to use; omit when only one is configured |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only and non-destructive, so the description adds extra value by disclosing that the data source resets with the server and that exclusion rules apply. This behavior is not visible in annotations and is important for interpretation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each earning its place: findings, the reset caveat, and the exclusion rationale. It is front-loaded with the most important output information and has no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the key findings, data freshness caveat, and exclusions, which is fairly complete for a read-only analysis tool. It lacks a description of the return format and does not explain the 'schema' parameter, but overall it gives sufficient context for correct use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not mention either parameter. Schema coverage is only 50%, with 'database' described and 'schema' having only a default and no description, leaving the semantics of the 'schema' parameter unclear. The tool description should compensate for this gap but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly enumerates the tool's outputs (unused indexes, duplicates, profitable sequential scans), which identifies it as an index analysis tool. It does not explicitly use a verb like 'analyze' or 'list', and it doesn't directly contrast with sibling tools, but the scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides important timing guidance: counters come from pg_stat_* and reset with server restart, so they should be read after real traffic. It also explains why primary-key and unique indexes are excluded, which helps the agent know when to trust results. It doesn't mention alternatives explicitly, but the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
database_healthHealth snapshotARead-onlyIdempotent
One snapshot of the things an operator would otherwise assemble by hand: cache hit ratio, connections, long-running statements and abandoned transactions, vacuum backlog, invalid indexes, sequences near their ceiling, replication lag. Scoped to the current database; anything the connected role cannot read is reported as unavailable rather than left out.
| Name | Required | Description | Default |
|---|---|---|---|
| database | No | which configured database to use; omit when only one is configured |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, but the description adds valuable context: it is scoped to the current database and that unreadable items are reported as 'unavailable' rather than omitted. This explains permission handling and output completeness, going beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one well-structured sentence that front-loads the core purpose ('One snapshot...') and efficiently lists the included metrics. Every clause adds information (scope, permission behavior). No wasted words or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, the presence of annotations, and the absence of an output schema, the description is complete. It explains what metrics are included, the scope, and how permission limitations are handled. This is sufficient for an agent to invoke the tool and interpret the results without further documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage for the single optional parameter 'database'. The description does not add meaning beyond the schema, but because the schema is sufficient, the baseline of 3 is appropriate. The description's reference to 'current database' indirectly relates to the parameter but does not enhance it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool produces a snapshot of database health metrics (cache hit ratio, connections, long-running statements, vacuum backlog, etc.). It uses a specific verb ('snapshot') and resource ('database health'), and the scope ('current database') is explicit. It distinguishes from siblings like top_queries or list_tables, which focus on narrower concerns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage as a one-stop health overview ('things an operator would otherwise assemble by hand'), which gives clear context for when to use it. However, it does not explicitly name alternatives or state when not to use it (e.g., when a specific metric is needed). This is a minor gap from full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_tableDescribe a tableARead-onlyIdempotent
Column names, types, nullability, defaults and primary key for one table, plus each column's comment. description is null unless somebody ran COMMENT ON — that means undocumented, not unused, so do not infer a column is dead from it.
| Name | Required | Description | Default |
|---|---|---|---|
| table | Yes | ||
| schema | Yes | ||
| database | No | which configured database to use; omit when only one is configured |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, and non-destructive. The description adds crucial behavioral context: null 'description' means undocumented, not unused, preventing a common analytical error. This goes beyond the annotations to clarify interpretation of results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long. The first sentence efficiently lists the returned metadata categories, and the second adds a high-value clarification about null descriptions. Every word earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple describe tool, the description covers the return contents (columns, types, nullability, defaults, PK, comments) and an important edge case (null comment semantics). Annotations handle safety, and no output schema exists, so the description fully carries the explanatory burden.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, with only 'database' having a description. The tool description does not explain the semantics of 'schema' or 'table' parameters, nor does it clarify that two are required and one optional. Parameter meaning relies entirely on names, which is inadequate given the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly specifies the tool returns column metadata (names, types, nullability, defaults, PK, comments) for one table. This distinguishes it from sibling tools like list_tables and list_schemas, which focus on listing rather than describing a single table's structure.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for inspecting a table's schema but provides no explicit 'when to use' or named alternatives. It gives context about what information is returned, which hints at suitable scenarios, but lacks exclusions or comparisons to sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_queryExplain one queryARead-onlyIdempotent
Why THIS statement is slow: the PostgreSQL execution plan for a query you provide. With analyze=true it actually runs the query and reports measured timings and buffer usage (still read-only, still rolled back). Use it on a specific statement; use top_queries to find out which statement to look at.
| Name | Required | Description | Default |
|---|---|---|---|
| sql | Yes | ||
| analyze | No | run the query and report real timings instead of estimates | |
| database | No | which configured database to use; omit when only one is configured |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds crucial context beyond annotations: with analyze=true, it actually runs the query and reports measured timings and buffer usage, while remaining read-only and rolled back. This clarifies the real-world impact of the tool and aligns with the readOnlyHint and destructiveHint annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each serving a distinct purpose: purpose, analyze behavior, and usage guidance. No redundant or extraneous content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the primary use case, the behavior of the analyze flag, and safety guarantees. While no output schema exists, it hints at what the report includes (execution plan, timings, buffer usage). Could be slightly richer in detailing return structure or limitations, but it is largely complete for agent selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% (sql lacks a description). The description compensates by explaining that sql is the query to analyze and adding detail to analyze (runs the query, reports timings and buffer usage, still rolled back). Database is already described in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool explains PostgreSQL execution plans for slow queries, with a specific verb and resource. It also distinguishes itself from sibling tools by pointing users to top_queries for which statement to analyze.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear when-to-use guidance: apply to a specific statement, and directs users to top_queries as the alternative for discovering which statement to examine. Does not explicitly mention other sibling tools or exclusions, but the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_schemasList schemasARead-onlyIdempotent
List the schemas in the database, excluding PostgreSQL's own catalogs. Start here when you do not know the layout yet.
| Name | Required | Description | Default |
|---|---|---|---|
| database | No | which configured database to use; omit when only one is configured |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context by noting that PostgreSQL's own catalogs are excluded, which is valuable beyond the annotations. However, it does not describe the return format or other details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences, both of which add value: the first states what the tool does and its scope, the second provides usage context. No unnecessary words or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with a single optional parameter, no output schema, and comprehensive annotations, the description is complete. It tells the agent what the tool does, what it excludes, and when to use it, which is sufficient for effective selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single optional 'database' parameter, so the schema fully documents the parameter. The description does not add any additional meaning about the parameter, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists schemas in the database and explicitly excludes PostgreSQL's own catalogs, which distinguishes it from sibling tools like list_tables and describe_table. The verb 'list' is specific and the resource is well-defined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Start here when you do not know the layout yet' provides clear contextual guidance for when to use the tool. It does not explicitly mention alternatives or exclusions, but for a simple discovery tool this is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tablesList tables in a schemaARead-onlyIdempotent
List tables, views and materialized views in one schema, with their comments. Only objects the connected role may read are shown.
| Name | Required | Description | Default |
|---|---|---|---|
| schema | Yes | ||
| database | No | which configured database to use; omit when only one is configured |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds valuable context: it includes views and materialized views, returns comments, and filters by the connected role's read permissions. This goes beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every phrase adds meaning: object types, scope, comments, and access control.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list operation with annotations providing safety context, the description covers the key aspects: object types, scope, comments, and permission filtering. There is no output schema, but the return contents are sufficiently implied. It could mention ordering or limit behavior, but that is not critical for this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50% because the 'schema' parameter lacks a description. The tool-level description mentions 'in one schema' but does not clarify the format or accepted values for the schema parameter. The 'database' parameter is well described in the schema, but the description does not compensate for the missing schema parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly specifies the verb 'list' and the resource (tables, views, materialized views) within a single schema, and includes comments. This distinguishes it from sibling tools like list_schemas and describe_table, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when you need to enumerate objects within a schema and indicates the scope ('in one schema'). It does not explicitly exclude alternatives, but the context is clear enough for an agent to distinguish from list_schemas or describe_table.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
queryRun SQL (read-only)ARead-onlyIdempotent
Run a read-only SQL query and return rows. Writes, DDL and administrative functions are refused before the statement reaches the database. At most 1000 rows come back unless you pass limit (server maximum 10000); truncated: true in the response means there is more data — page through it with offset, and give the query an ORDER BY when you do, or the rows you get on page two depend on the planner's mood.
| Name | Required | Description | Default |
|---|---|---|---|
| sql | Yes | a single read-only statement: SELECT, WITH, VALUES, EXPLAIN or SHOW (write `SELECT * FROM t`, not `TABLE t`) | |
| limit | No | maximum rows to return; larger values are capped at 10000 and the response says so | |
| offset | No | rows to skip, for paging; pair it with ORDER BY for stable pages | |
| database | No | which configured database to use; omit when only one is configured |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by disclosing that writes/DDL/admin functions are refused, the default row limit of 1000, the maximum of 10000, the truncated flag, offset paging, and the need for ORDER BY to ensure stable pagination. This provides rich behavioral context beyond the readOnlyHint and destructiveHint annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the core purpose, then adding safety and pagination details. Every sentence carries useful information without redundancy or rambling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Combined with comprehensive parameter descriptions and annotations, the description covers safety, row limits, truncation, paging, and ordering guidance. There is no output schema, but the description mentions the truncated flag, and the level of detail is sufficient for an AI agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides 100% description coverage for all four parameters, so the baseline is 3. The description adds operational context about truncation, paging, and ORDER BY best practices, but it does not significantly alter parameter meanings beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Run a read-only SQL query and return rows', using a specific verb and resource. It distinguishes itself from specialized siblings by being the generic query runner, and the title reinforces the read-only scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by focusing on read-only SQL and allowed statement types, but it does not explicitly tell when to use this tool versus alternatives like explain_query or list_schemas. There are no exclusions or alternative recommendations, leaving the choice somewhat inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
security_postureWhat this server can and cannot doARead-onlyIdempotent
What this deployment is actually able to do to your database, asked of PostgreSQL rather than assumed: whether the connected role can write, bypass row-level security or reach server files; whether the transport is authenticated; whether the audit chain is keyed; whether the connection is encrypted. Returns a grade and, for anything wrong, the command that fixes it. Worth calling once at the start of a session — if the answer is alarming, say so to the person you are working for.
| Name | Required | Description | Default |
|---|---|---|---|
| database | No | which configured database to use; omit when only one is configured |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only, idempotent, non-destructive behavior. The description adds context by stating the tool asks PostgreSQL rather than assuming, enumerates the specific checks performed, and reveals that it returns a grade plus fix commands. It also advises acting on alarming results. This goes beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no fluff. It front-loads the purpose, lists specific checks in a compact list, then states the output format and usage recommendation. Every clause adds value, making it appropriately sized for a security audit tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lacking an output schema, the description explains the return value (a grade and fix commands) and covers all relevant behavioral aspects. It also provides a usage context (session start) and what to do with results. For a tool with one optional parameter, this is complete and self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% coverage with a single optional 'database' parameter described as 'which configured database to use; omit when only one is configured.' The description does not add parameter-specific details, but the schema already fully covers it, meeting the baseline for schema-heavy tools.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool assesses the deployment's actual capabilities against the database (write, RLS bypass, file access, auth, encryption) and returns a grade with fix commands. It distinguishes itself from sibling tools like database_health and query by being a session-start security audit rather than a data operation. The verb is specific and the resource (security posture) is well defined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly recommends calling this tool once at the start of a session and instructs the agent to communicate alarming results to the user. It does not explicitly mention alternatives, but given sibling tools, the context is clear. The guidance is useful and actionable, though not exhaustive on when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
simulate_indexWould this index help?ARead-onlyIdempotent
Answers whether an index would change the plan for a given query — WITHOUT creating it. Uses the hypopg extension, which registers the index in backend memory only: the planner sees it, storage never does, and it is gone when the call returns. Give the query, the table and the columns; the index definition is assembled here, so there is no way to send DDL through this tool. Returns the plan and cost with and without, and whether the planner actually reached for it — a cost that barely moves and an index the planner ignored are different answers. These are planner ESTIMATES, not measured times: treat a big improvement as a reason to test the index, not as proof. Needs hypopg installed; says so plainly, with the package name, when it is missing.
| Name | Required | Description | Default |
|---|---|---|---|
| sql | Yes | the read-only query the index is supposed to help | |
| table | Yes | the table to index, `schema.table` or just `table` (defaults to the `schema` field, then to public) | |
| using | No | access method | btree |
| schema | No | used only when `table` is unqualified | public |
| columns | Yes | column names, in index order | |
| database | No | which configured database to use; omit when only one is configured |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the readOnly/idempotent/destructive annotations by explaining the hypopg mechanism: index registered in backend memory only, gone when the call returns. It also discloses the failure mode if hypopg is missing and clarifies that results are planner estimates, not measured times. This aligns with annotations with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense yet well-structured, front-loading purpose and safety, then covering inputs, outputs, interpretation, and prerequisites. Every sentence earns its place; there is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description fully explains what is returned: plan and cost with and without, plus whether the planner used the index. It also covers prerequisite, absence of side effects, and interpretation nuance, making it complete for a complex tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers 100% of parameters with descriptions, so baseline is 3. The description adds value by stating 'the index definition is assembled here,' clarifying that users supply table and columns rather than a full DDL statement, which goes beyond the schema's individual parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Answers whether an index would change the plan for a given query — WITHOUT creating it,' a specific verb and resource. It explicitly distinguishes itself from sibling tools like analyze_indexes and explain_query by emphasizing 'the planner sees it, storage never does,' making the hypothetical nature unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage guidance: provide query, table, and columns; the index is assembled automatically; no DDL can be sent. It also warns that these are estimates, not proof. However, it does not explicitly name alternative tools or state when not to use this tool, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
top_queriesSlowest statements server-wideARead-onlyIdempotent
WHICH statements cost the most, ranked by total execution time across the whole server. Requires the pg_stat_statements extension; if it is missing the answer says how to enable it. Take the statement you find here to explain_query for the plan.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| database | No | which configured database to use; omit when only one is configured |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive behavior. The description adds a key behavioral trait: it requires the pg_stat_statements extension and will return instructions to enable it if absent. It also implies the output contains statement text suitable for explain_query. This adds meaningful transparency beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, tightly packed with purpose, prerequisites, and a cross-reference to explain_query. It front-loads the primary action and avoids unnecessary elaboration. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity of the tool (2 params, no output schema), the description covers the essential aspects: what it does, the extension dependency, and a pointer to the next step. It doesn't mention ordering direction (descending) or output limit behavior, but these are reasonably inferred. It is nearly complete for a read-only listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides descriptions for only one of two parameters (database) and the limit parameter has constraints but no description. The tool description does not mention parameters at all. With 50% schema coverage, the description should compensate, but it doesn't. The limit parameter is self-explanatory from its name and constraints, so a mid-score is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool identifies the statements costing the most, ranked by total execution time server-wide. It distinguishes from siblings by specifying the resource (statements) and the metric (total execution time), and it contrasts with explain_query which is for plans.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: use to find top expensive statements, and then take the result to explain_query for the plan. It also notes the extension requirement and that the tool will explain how to enable it if missing. However, it doesn't explicitly state when not to use this tool versus other query tools, so it gets a 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.1.0- First observed
analyze_indexes - First observed
database_health - First observed
describe_table - First observed
explain_query - First observed
list_schemas - First observed
list_tables - First observed
query - First observed
security_posture - First observed
simulate_index - First observed
top_queries
TDQS
Scored across 10 tools
Each tool targets a distinct concern: health snapshot, query ranking, generic read-only query, schema enumeration, table description, query explanation, security audit, hypothetical index testing, and index usage analysis. The two index-related tools are clearly separated by hypothetical vs. actual, and the descriptions reinforce the boundaries.
All names are lowercase with underscores, and most follow a verb_noun pattern (list_schemas, describe_table, explain_query, simulate_index, analyze_indexes). A few are noun or adjective phrases (database_health, top_queries, security_posture), which is a minor deviation from the verb pattern but still consistent in style and predictably scoped.
Ten tools is within the sweet spot for a database diagnostics server. Each tool covers a distinct aspect of operation and analysis, and there is no redundancy or bloat. The scope is broad enough to be useful without overwhelming an agent.
The tool surface covers the full lifecycle of database investigation: discover schema, inspect tables, run read-only queries, diagnose slow queries via top_queries and explain, simulate and analyze indexes, and assess security posture. The generic query tool provides an escape hatch for anything not explicitly covered, ensuring no dead ends.
Maintenance
Related MCP Connectors
Safe, read-only Postgres and MySQL access for AI agents. Audit log + column-level controls.
Deterministic safety, correctness & cost gate that vets Postgres SQL before your AI agent runs it.
Query PostgreSQL databases in plain English — LLM-generated, safety-validated SQL.
Query your Postgres from ChatGPT or Claude without exposing the database or handing over credentials. Run npx boltschema connect next to your database and it dials out over HTTPS — no inbound firewall rule, no open port, works with localhost and VPC-private databases. Read-only is enforced by a SQL guard, a Postgres READ ONLY transaction, and a scoped role generated for you.
Related MCP Servers
- -licenseNot gradedqualityAmaintenanceA Model Context Protocol server that provides read-only access to PostgreSQL databases. This server enables LLMs to inspect database schemas and execute read-only queries.93,517 npm90,399MIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that provides AI assistants with secure, read-only access to PostgreSQL databases while offering comprehensive tools for schema exploration, query validation, and performance optimization.MIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server providing read-only access to PostgreSQL databases, enabling LLMs to inspect database schemas and execute read-only SQL queries.93,517 npmMIT
- FlicenseNot gradedqualityDmaintenanceAn MCP server that gives an AI agent scoped, safe access to your Postgres databases with per-connection access control, row caps, timeouts, and defense-in-depth read-only enforcement.-