Skip to main content
Glama
deBilla

BigQuery MCP

by deBilla

BigQuery MCP

CI

A read-only Model Context Protocol server over Google BigQuery. It lets an AI client (Claude Code, Claude Desktop, …) answer plain-language data questions by discovering schema and running SELECT queries.

The AI does the natural-language → SQL translation; this server just safely executes against BigQuery under your own Google credentials.


Tools exposed

Tool

Purpose

Cost

list_datasets

List datasets in the project

free

list_tables

List tables/views in a dataset

free

get_table_schema

Columns (nested paths expanded), partitioning, size, row count

free

check_table_freshness

When each table was last written — catches stale sources

free

list_environments

Which BigQuery environments are configured, and the default

free

list_scheduled_queries

Which scheduled query writes a table, and whether it is disabled or failing

free

get_scheduled_query

One query's SQL, destination and recent runs

free

list_code_assets

Notebooks, saved queries and data canvases in BigQuery Studio

free

get_code_asset

One notebook or saved query's body, notebook outputs stripped

free

find_code_assets_using_table

Which notebooks/saved queries read a table

quota, not $

list_notebook_schedules

Scheduled Colab notebooks, with how many recent runs failed

free

list_notebook_runs

Individual notebook runs across every schedule — failures by default

free

get_notebook_schedule

One schedule's cron, notebook and recent errors

free

run_query

Run a validated, read-only SELECT and return rows

scans data

Only run_query costs anything, so the discovery tools are the ones to spend first. Two of them exist to prevent specific, repeated mistakes:

  • get_table_schema reports partitioning from table metadata, never from column names. A table with a partition_date column may not be partitioned — in which case no WHERE clause reduces the scan and every query reads the whole table. The response flags this explicitly when the table is large.

  • check_table_freshness finds tables that stopped being written to without being dropped. Those return stale data rather than an error, which is the failure mode nobody notices.

  • find_code_assets_using_table answers the other half. The scheduled-query tools say what writes a table; this says who reads it, which is the question before a schema change. It is also the one discovery tool that is not free: it opens every asset it considers, spending Dataform read quota, so it is capped and reports how much of the project it actually covered. A result is evidence about the assets scanned, never proof that nothing else uses the table — assets it could not read are listed separately rather than counted as misses.

  • list_scheduled_queries says why. A stale table is usually a scheduled query that was disabled or is failing, and that lives in a different API (BigQuery Data Transfer) needing roles/bigquerydatatransfer.viewer. Without that role the two tools return an error naming it and everything else works normally. Most scheduled queries declare no destination because they write with DDL, so the target is read out of the SQL and reported as writes_to_from_sql — a heuristic, labelled as one.


Related MCP server: BigQuery MCP Server

BigQuery Studio notebooks and saved queries

These are not BigQuery resources. BigQuery Studio stores each code asset as a Dataform repository holding a single file, which means a third API and a third permission — roles/dataform.viewer — beyond BigQuery and the Data Transfer Service. Without it the three tools return an error naming the role and everything else works normally. They are also invisible in the Dataform UI, so nothing in the console hints that this is where they live.

Two things about that storage are worth knowing before you configure it:

  • Code assets are regional, and it is not the dataset region. Dataform rejects multi-regions, so a platform whose datasets are US keeps its notebooks in something like us-central1. No configuration is needed: when location is a multi-region the server probes the regions inside it, uses the one holding the assets, and says so — a multi-region cannot simply be inherited, because using it is guaranteed to fail rather than merely likely to. Pin code_asset_location (or BQ_CODE_ASSET_LOCATION) to skip the probing; an explicit value is never second-guessed, so a wrong one returns an empty list rather than an error. Every result echoes back the location it read.

  • Notebook bodies are mostly output. Across 52 real notebooks, cell outputs were 77% of the bytes — one was 1.44 MB of file for 80 KB of code. Outputs are stripped before anything is returned, and the saving is reported so you can see that what is missing was rendered charts rather than logic.

Reads are quota-limited by volume rather than by concurrency, and the quota refills over tens of seconds. Exhaustion is retried with backoff and, if it persists, reported as something to retry shortly rather than as a failure.


Scheduled Colab notebooks

A notebook in BigQuery Studio can have a schedule attached to it, and those runs are where scheduled work goes unwatched: nothing reports on them, so a notebook can fail nightly for weeks and the only symptom is a table that quietly stopped moving. On the platform this was built against, 354 of 7,061 runs had failed and none of it was visible.

Answering one question means joining three resources, which is most of why it was hard to see. The schedule is a Vertex AI Schedule (cron, timezone, paused or not); each run is a NotebookExecutionJob; the notebook is the same Dataform code asset list_code_assets lists — so a failing schedule leads straight to get_code_asset for the code that failed. That means a fourth API and a fourth permission, roles/aiplatform.viewer. Without it the three tools return an error naming the role and everything else works normally. No new dependency: this talks to Vertex AI over REST rather than pulling in google-cloud-aiplatform for two list endpoints.

Three things are worth knowing before trusting what you see:

  • A schedule reports itself healthy while its notebook fails. Every one of 49 live schedules reports its last scheduled run as OK — including one whose previous 79 runs had failed. OK means the scheduler successfully launched a job, not that the notebook ran. That field is what the console shows first, and it is exactly how this went unnoticed, so health here is always computed from execution jobs and the schedule's own status is never reported as one. A PAUSED schedule, meanwhile, is the single most common reason a notebook-written table went stale.

  • Outcome cannot be filtered server-side. Vertex AI rejects a jobState filter outright and caps pages at 100, so "show me the failures" means reading pages and filtering locally. Runs are ordered newest-first, which is what makes a bounded window affordable: lookback_days (30 by default, so monthly schedules show at least one run) stops the walk instead of scanning a year of history. 30 days is about 7 API calls; 90 is about 23.

  • Runs outlive the schedule that created them. Deleting and recreating a schedule is the normal way to edit one, so 7,061 live runs referenced 75 distinct schedules of which only 49 still existed. Those runs are real and their failures count, so they are labelled (schedule no longer exists) rather than dropped. For the same reason display names are not unique — two live schedules shared one — so an ambiguous name is an error listing the candidates and their ids, never a silently chosen first match.

data-platform-mcp doctor reports how many schedules are readable, how many are paused and how many failed in the last week.


Environments

One server answers questions about several targets — a warehouse and its staging copy, or two regions of the same project. Every tool takes an optional environment; omitting it uses the default.

# ~/.config/data-platform-mcp/config.toml
default_environment = "warehouse"

[environments.warehouse]
project = "my-data-platform"
impersonate = "data-platform-mcp-ro@my-data-platform.iam.gserviceaccount.com"
dataset_allowlist = ["sales", "events"]

[environments.central]          # same project, different region
project = "my-data-platform"
location = "us-central1"

See config.toml.example for every setting, or set BQ_MCP_ENVIRONMENTS to the same structure as JSON. A single BQ_PROJECT still works unchanged — it becomes one environment named default.

An environment can be named by its own name, an alias, the built-in shorthands (prod, stg, dev, live) or its project id. An unknown name is an error naming the valid options, never a silent fall back to the default: a typo that answered a production question from staging would be invisible in the reply. Every result echoes back the environment it came from.

Regions are why this matters most here. BigQuery cannot query across locations, and its error for trying names neither location, so it reads as a missing table. One environment per location; doctor reports which datasets are where.


Read-only as a property of the identity

The SELECT-only guard and the readOnlyHint annotations are promises about this code. Pointing the server at a service account that holds only roles/bigquery.jobUser and a dataset-scoped roles/bigquery.dataViewer makes it a fact about the credentials — enforced by IAM whatever the code does, and whatever your own roles allow:

data-platform-mcp setup --project my-data-platform --datasets sales,events

Creates the account, grants those two roles, and gives you roles/iam.serviceAccountTokenCreator on it so the server can impersonate it. Add --dry-run to see the commands first; it is safe to re-run.

With --datasets, the dataset allowlist stops being an if statement in this process and becomes a grant Google enforces.


macOS setup

Terminal.app is not Xcode. It ships with every Mac. What does not ship is the Xcode Command Line Tools, and an analyst's laptop usually has neither those nor Homebrew. Nothing here needs them — but it is easy to trip over by accident, because git, make, clang and the stock /usr/bin/python3 are stubs for that bundle: running any of them pops a system dialog offering to install about a gigabyte of developer tooling.

None of the commands below invoke one. They use only utilities macOS already has — curl, tar, sh, uname — because both installs are self-contained:

Install

Why it needs nothing else

uv

A standalone binary. Its installer never mentions Python, and it downloads its own to run the server.

Google Cloud CLI

The macOS tarball bundles its own Python (.install/bundled-python3-unix-darwin-*).

The whole terminal requirement is the three blocks below, once.

1. Install uv

curl -LsSf https://astral.sh/uv/install.sh | sh
which uvx      # note this absolute path — Claude Desktop will need it

Typically /Users/<you>/.local/bin/uvx.

2. Install the Google Cloud CLI

Pick the build for your chip — uname -m prints arm64 for Apple Silicon, x86_64 for Intel:

# Apple Silicon
curl -O https://dl.google.com/dl/cloudsdk/channels/rapid/downloads/google-cloud-cli-darwin-arm.tar.gz
tar -xzf google-cloud-cli-darwin-arm.tar.gz

# Intel — same, with the other file
# curl -O https://dl.google.com/dl/cloudsdk/channels/rapid/downloads/google-cloud-cli-darwin-x86_64.tar.gz
# tar -xzf google-cloud-cli-darwin-x86_64.tar.gz

./google-cloud-sdk/install.sh --quiet

Avoid brew install --cask google-cloud-sdk: Homebrew itself requires the Command Line Tools, which is the thing this section exists to avoid.

3. Authenticate

./google-cloud-sdk/bin/gcloud auth application-default login
./google-cloud-sdk/bin/gcloud auth application-default set-quota-project your-gcp-project

This writes a credentials file that the Google libraries read directly. gcloud does not need to be on your PATH afterwards — it is needed once, here. That is why a GUI-launched Claude Desktop can query BigQuery even though it cannot see your shell.

Your account needs BigQuery Job User on the project the query runs in, and BigQuery Data Viewer on each dataset it reads — often a different project.

4. Check it worked

BQ_PROJECT=your-gcp-project uvx data-platform-mcp doctor

Then register with your client: Claude Desktop or Claude Code.

Alternative: no terminal at all for the analyst

If even that is too much, an admin can do the credential half centrally and the analyst installs nothing but uv — skipping step 2 and step 3 entirely. (Nothing about this is macOS-specific; it works the same on any OS.)

# the admin, once, on their own machine
data-platform-mcp setup --project your-gcp-project --datasets sales,events
gcloud iam service-accounts keys create analyst-key.json \
  --iam-account data-platform-mcp-ro@your-gcp-project.iam.gserviceaccount.com

The analyst saves that file and points the config at it:

{
  "mcpServers": {
    "bigquery": {
      "command": "/Users/YOU/.local/bin/uvx",
      "args": ["data-platform-mcp@latest"],
      "env": {
        "BQ_PROJECT": "your-gcp-project",
        "GOOGLE_APPLICATION_CREDENTIALS": "/Users/YOU/keys/analyst-key.json"
      }
    }
  }
}

The trade-off is real and worth stating. A key file is a long-lived credential sitting on a laptop, where gcloud auth application-default login issues short-lived tokens tied to a person. It is defensible here because the account created by setup --datasets can only read the datasets you name, and because a key can be revoked centrally the moment a laptop is lost — but it is strictly weaker, and it is a per-analyst secret, so do not put it in a shared config file or a repository.


Linux

The same two installs, with no equivalent of the Command Line Tools problem:

curl -LsSf https://astral.sh/uv/install.sh | sh
curl -O https://dl.google.com/dl/cloudsdk/channels/rapid/downloads/google-cloud-cli-linux-x86_64.tar.gz
tar -xzf google-cloud-cli-linux-x86_64.tar.gz
./google-cloud-sdk/install.sh --quiet
./google-cloud-sdk/bin/gcloud auth application-default login

Then check it worked and register with your client.


Windows

Not verified end to end. The download URLs and install locations below were checked; the flow itself has not been run on a Windows machine. CI tests Linux only. Treat this as a careful derivation, not a tested recipe — and please open an issue if a step is wrong.

1. Install uv

In PowerShell:

powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

This installs uv.exe and uvx.exe into %USERPROFILE%\.local\bin. Confirm the exact path, because the Desktop config needs it in full:

(Get-Command uvx).Source

2. Install the Google Cloud CLI

Download and run GoogleCloudSDKInstaller.exe. Leave "Bundled Python" ticked — it is what lets the SDK run without a separate Python install, the same property the macOS tarball has.

3. Authenticate

In a new PowerShell window, so it picks up the updated PATH:

gcloud auth application-default login
gcloud auth application-default set-quota-project your-gcp-project

This writes credentials to %APPDATA%\gcloud\application_default_credentials.json, which the Google libraries read directly — so gcloud need not be on PATH afterwards.

4. Check it worked

$env:BQ_PROJECT="your-gcp-project"; uvx data-platform-mcp doctor

5. Configure Claude Desktop

%APPDATA%\Claude\claude_desktop_config.json — create it if absent. Backslashes must be doubled in JSON, and the path must be absolute:

{
  "mcpServers": {
    "bigquery": {
      "command": "C:\\Users\\YOU\\.local\\bin\\uvx.exe",
      "args": ["data-platform-mcp@latest"],
      "env": {
        "BQ_PROJECT": "your-gcp-project"
      }
    }
  }
}

Replace C:\Users\YOU\... with what (Get-Command uvx).Source printed, with each \ written as \\. Then fully quit and reopen Claude Desktop.

If it fails, the logs are in %APPDATA%\Claude\logs\. ENOENT there means the command path is wrong or its backslashes were not doubled — the same failure macOS has, with one extra way to get it wrong.


Quick start (per user)

Each person runs their own local copy. Queries execute under their own BigQuery/IAM permissions, so existing access controls decide who can see what.

Install the prerequisites for your platform first — macOS, Linux, Windows — then come back here.

1. Install

The package is published as data-platform-mcp (bigquery-mcp was already taken on PyPI by an unrelated project). No checkout is needed — the client can fetch and run it directly:

uvx data-platform-mcp --version

From source, for development:

git clone git@github.com:deBilla/bigquery-mcp.git
cd bigquery-mcp

python3 -m venv .venv
./.venv/bin/pip install -e .

Either way you get a data-platform-mcp command, which is what the client runs.

2. Authenticate to Google (one time)

Covered in the platform sections above: macOS step 3, or the equivalent gcloud auth application-default login elsewhere. Queries then run under your own credentials via Application Default Credentials.

Using a service-account key instead? Set GOOGLE_APPLICATION_CREDENTIALS to its path — but set it where the MCP server is launched, not in a shell:

// in your client's MCP config, alongside BQ_PROJECT
"env": {
  "BQ_PROJECT": "your-gcp-project",
  "GOOGLE_APPLICATION_CREDENTIALS": "/absolute/path/to/key.json"
}

The client spawns the server as a subprocess with only the environment its config declares. Exporting the variable in a terminal has no effect on it — that is a distinct failure from having no credentials at all, and it looks identical from the outside.

3. Check your setup

BQ_PROJECT=your-gcp-project data-platform-mcp doctor

Checks credentials, job permission, dataset visibility and — the one that catches people — dataset regions. BigQuery cannot query a dataset from a different location, and its own error names neither the location it wanted nor the one the dataset is in, so it reads as a missing table. doctor names both:

[  ok  ] run a query in my-project (location US)
[  ok  ] 39 datasets visible (no allowlist; all are readable)
[ warn ] 6 of 39 datasets are outside location US
         US-CENTRAL1: analytics_raw, business_data, ds_public, pg_public, public, recommendations
         BigQuery cannot query these from US, and cannot join them with
         datasets that are in it.
         Fix:  set BQ_LOCATION to the region you need, and run a separate
               server for datasets in another one.

A dataset in another region is a warning; one on your BQ_DATASET_ALLOWLIST is a failure, because no tool call could ever read it.

4. Register with your AI client

Replace your-gcp-project with your GCP project ID.

Claude Code — once published:

claude mcp add bigquery \
  --env BQ_PROJECT=your-gcp-project \
  -- uvx data-platform-mcp

From a source install, point at the checkout instead (replace /abs/path/bigquery-mcp):

claude mcp add bigquery \
  --env BQ_PROJECT=your-gcp-project \
  -- /abs/path/bigquery-mcp/.venv/bin/data-platform-mcp

Claude Desktop — see the dedicated section below; it needs absolute paths.

5. Restart the client and ask a question

"Which datasets are available? In the sales dataset, how many rows does the orders table have?"


Claude Desktop

Most of a data team will use Desktop rather than the CLI, and it has one failure mode the CLI does not.

Claude Desktop does not inherit your shell PATH. It launches from the Finder, so uvx, python and anything installed by Homebrew or uv are invisible to it. A config that says "command": "uvx" fails with ENOENT — the server never starts, and the error names the command rather than the reason. Every path in this file must be absolute.

What does not break: credentials. Application Default Credentials are a file that the Google libraries read directly, so gcloud does not need to be on PATH for queries to work — it is only needed once, in a terminal, to create that file. Verified by running this server with an entirely empty environment: the query succeeded.

1. Install and authenticate

Do the platform setup first — macOS (two pastes, no Xcode tools needed) or Linux, Windows. You need two things from it: the absolute path that which uvx printed, and a completed gcloud auth application-default login.

2. Edit the config

Claude Desktop's Settings → Connectors lists hosted connectors; a local server like this one is not added there. It goes in a JSON file instead:

Settings → Developer → Edit Config opens it. Or edit it directly:

OS

File

macOS

~/Library/Application Support/Claude/claude_desktop_config.json

Windows

%APPDATA%\Claude\claude_desktop_config.json

The file usually already exists and holds your Desktop preferences. Add mcpServers as one more top-level key — do not replace the file, or you will lose those settings. If it genuinely does not exist, create it with just the block below.

{
  "mcpServers": {
    "bigquery": {
      "command": "/Users/YOU/.local/bin/uvx",
      "args": ["data-platform-mcp@latest"],
      "env": {
        "BQ_PROJECT": "your-gcp-project"
      }
    }
  }
}

Replace /Users/YOU/.local/bin/uvx with what which uvx printed. On Windows the path looks like C:\\Users\\YOU\\.local\\bin\\uvx.exe, and backslashes must be doubled in JSON.

Merged into a file that already has settings, it looks like this — mcpServers sits alongside whatever is there, not instead of it:

{
  "preferences": { "...": "your existing settings, left alone" },
  "mcpServers": {
    "bigquery": {
      "command": "/Users/YOU/.local/bin/uvx",
      "args": ["data-platform-mcp@latest"],
      "env": { "BQ_PROJECT": "your-gcp-project" }
    }
  }
}

Check it still parses before restarting — a stray comma disables every server, silently:

python3 -m json.tool ~/Library/Application\ Support/Claude/claude_desktop_config.json

3. Restart Claude Desktop

Fully quit and reopen — reloading the window is not enough. The server appears under the tools icon in the message box.

Managing several warehouses

Rather than growing the JSON, put the environments in ~/.config/data-platform-mcp/config.toml (see Environments). The Desktop config then needs no env block at all, and is identical on every machine:

{
  "mcpServers": {
    "bigquery": {
      "command": "/Users/YOU/.local/bin/uvx",
      "args": ["data-platform-mcp@latest"]
    }
  }
}

This is the better shape for a team: one config file to share, and the JSON stops carrying project ids.

When it does not work

Desktop hides the reason, so check in this order:

  1. Run the doctor in a terminal. It reports credentials, roles, dataset visibility and regions in one pass, and is the fastest way to tell a setup problem from a Desktop problem:

    BQ_PROJECT=your-gcp-project /Users/YOU/.local/bin/uvx data-platform-mcp doctor
  2. Read the logs. macOS: ~/Library/Logs/Claude/mcp*.log. ENOENT or "command not found" there means the command path is wrong — go back to which uvx.

  3. Check the JSON parses. A trailing comma silently disables every server:

    python3 -m json.tool ~/Library/Application\ Support/Claude/claude_desktop_config.json

Safety

  • Every query is dry-run first to validate it and estimate bytes scanned.

  • Only SELECT / WITH statements run — no writes, DDL, or DML.

  • Cost confirmation: a query estimated to scan more than BQ_WARN_BYTES (default 1 GB) does not run. It returns status: "confirmation_required" with the estimated scan size and dollar cost so the client can ask before proceeding. Re-call with confirm_expensive=true to run it.

  • Hard cap: queries above BQ_MAX_BYTES_BILLED (default 5 GB) never run, even with confirmation — a runaway-cost backstop.

  • Optional dataset allowlist restricts what can be read.

  • Refusals are protocol errors. Anything the server declines to do — a non-SELECT statement, a disallowed dataset, a query over the hard cap — arrives with MCP's isError set, so it cannot be mistaken for a result. confirmation_required is the deliberate exception: it is a normal result, because the agent is meant to relay it and come back.

  • Responses are size-bounded. run_query stops adding rows once the serialised response reaches ~40k characters and sets stopped_for_size, so a wide result cannot quietly consume the whole context window. A partial answer always says that it is partial.

  • SQL is never written to the audit log — only a hash and a length. Query text routinely contains the user IDs or emails it filters on.

Cost-confirmation flow

run_query(sql)
   │  dry run estimates the scan
   ├── ≤ 1 GB ........... runs, returns rows + estimated_cost_usd
   ├── 1–5 GB .......... status: confirmation_required (size + $ estimate) → ask user
   │                      → run_query(sql, confirm_expensive=true) runs it
   └── > 5 GB ........... rejected, never runs

Configuration (environment variables)

Var

Default

Meaning

BQ_MCP_ENVIRONMENTS

(none)

JSON map of environment name to settings. Takes precedence over the config file.

BQ_MCP_DEFAULT_ENVIRONMENT

(safest, else first)

Environment used when a call omits environment. Prefers a staging/dev environment when unset.

BQ_MCP_CONFIG

~/.config/data-platform-mcp/config.toml

Path to the TOML config file

BQ_IMPERSONATE_SERVICE_ACCOUNT

(none)

Read-only service account to impersonate

BQ_PROJECT

(ADC project)

GCP project ID whose BigQuery datasets you query. Falls back to the project associated with your credentials; tools error with instructions if neither is set.

BQ_LOCATION

US

BigQuery location

BQ_WARN_BYTES

1073741824 (1 GB)

Above this, ask the user to confirm before running

BQ_MAX_BYTES_BILLED

5368709120 (5 GB)

Hard per-query scan cap — never exceeded

BQ_COST_PER_TIB_USD

6.25

On-demand price used to render the cost estimate

BQ_ROW_LIMIT

200

Default rows returned

BQ_DATASET_ALLOWLIST

(empty = all)

Comma-separated dataset IDs

BQ_MCP_TRANSPORT

stdio

stdio (subprocess) or http/sse (serve over network)

BQ_MCP_HOST

127.0.0.1

Bind host when transport is http/sse. run-http.sh overrides this to 0.0.0.0 so containers can reach it — see the security note below.

BQ_MCP_PORT

8765

Bind port when transport is http/sse

BQ_MCP_AUDIT_LOG

~/.local/state/data-platform-mcp/audit.jsonl

JSONL record of every tool call. off disables it. SQL text is never written — only a hash and length.

BQ_MCP_LOG_LEVEL

INFO

Verbosity of the stderr log

By default the server speaks stdio — the right choice when a client spawns it (Claude Code, Claude Desktop), and what the Quick start above uses.


Advanced: serve over HTTP

To reach the server from a remote or containerized client instead of having each client spawn its own, run it over HTTP:

BQ_PROJECT=your-gcp-project ./run-http.sh
# Serving … on http://0.0.0.0:8765/mcp

Clients then connect by URL (Claude Code):

claude mcp add --transport http bigquery http://<host>:8765/mcp

⚠️ Security: the HTTP endpoint has no authentication, and every query runs under the host's ADC credentials — not the connecting user's. Anyone who can reach the port gets full read access to BQ_PROJECT under your identity. Only expose it on a trusted network (bind BQ_MCP_HOST=127.0.0.1 and use an SSH tunnel/VPN, or an authenticating proxy). See docs/nanoclaw.md for the containerized-client setup this mode was designed for.

For server deployments, point GOOGLE_APPLICATION_CREDENTIALS at a service-account key with BigQuery Data Viewer + Job User roles instead of using personal ADC.


Development

./.venv/bin/pip install -e ".[dev]"
./.venv/bin/python -m pytest

The suite needs no credentials and no network — every test runs against fakes in tests/conftest.py, so it is deterministic and free. Layers:

File

Covers

test_protocol.py

The MCP contract through a real in-memory client session: tool set, read-only annotations, generated schemas, isError on refusal

test_query_guard.py

The cost gate — what runs, what is refused, what is handed back to the user, and what the caller is told about limits

test_payload_shape.py

Response shapes against fake tables, including the partitioning trap and nested-field flattening

test_observability.py

The audit trail, and the promise that SQL text never reaches it

test_diagnostics.py

doctor's report, including the region and allowlist failures it exists to catch early

test_environments.py

Routing between environments, per-environment limits, and impersonation targeting

test_config.py

The environment registry, aliases, the TOML file, and the missing-project error that used to be an import-time crash

test_errors.py

Auth failures carry the command that fixes them

test_formatting.py

The size and cost figures a user is asked to approve

test_eval_scoring.py

The eval scorer, fed the trajectories each case exists to reject

Evals

Two further layers need live credentials, so they are not part of pytest: evals/measure.py records what a client actually receives from each tool, and evals/tool_use_evals.py asks real questions through the claude CLI and scores the trajectory from the server's own audit log — which tool ran, against which environment, with which arguments.

./.venv/bin/python evals/measure.py                     # payload sizes
./.venv/bin/python evals/tool_use_evals.py              # 6 cases, spends tokens
./.venv/bin/python evals/tool_use_evals.py --rescore    # re-score saved replies, free

See evals/README.md for what each case catches and evals/BASELINE.md for what the last run measured. Tool and server descriptions are the highest-leverage thing to change in this server, and nothing except an eval tells you they need changing.

Mutation testing

A suite that passes on its first run proves nothing, so the guarantees above were checked by breaking them: reverting refusals to error-shaped returns, logging raw SQL, guessing partitioning from column names, removing the response budget, dropping functools.wraps from the audit wrapper, letting confirmation bypass the hard cap, silencing stale-table detection, and removing the allowlist check. Each one fails the suite.

Upgrading

uvx resolves the latest version on its first run only, then reuses the cached environment indefinitely. A server left as ["data-platform-mcp"] keeps running the build it first downloaded: tools added by a later release are simply absent, which reads as "this server cannot do that" rather than as an upgrade that has not landed. Pin @latest so every start re-resolves:

"args": ["data-platform-mcp@latest"]

Then fully quit and reopen the client. MCP servers are spawned at client startup, so a reload leaves the old process running.

To upgrade a one-off without editing config, uv cache clean data-platform-mcp (or uv tool upgrade data-platform-mcp if it was installed with uv tool install). list_environments reports server_version, so you can confirm what is actually running from inside the conversation.

Releasing

Version numbers live in two files and CI refuses a tag where they disagree — a mismatch would ship a tag pointing at different code than the package claims. (__version__ is read from the installed distribution, so it cannot drift.)

# 1. bump both to the same value
#      pyproject.toml   project.version
#      server.json      version  AND  packages[0].version

# 2. tag and push
git tag v0.3.3 && git push origin v0.3.3

The tag triggers .github/workflows/release.yml, which verifies the versions agree, builds, publishes to PyPI via Trusted Publishing, then registers the release with the MCP registry. Neither step stores a token: PyPI uses OIDC from this repository and the pypi environment, and the registry uses GitHub OIDC. Both need one-time setup before the first release:

What CI checks

.github/workflows/ci.yml runs on every push and pull request:

Job

Checks

test

The suite on Python 3.11, 3.12 and 3.13 — with no GCP credentials on the runner, which is the point

safety

No credential-shaped strings in tracked files; .env/.mcp.json untracked; no mutating BigQuery client calls anywhere in src/

package

Builds, twine checks, asserts no local config leaked into the sdist, then installs the wheel into a clean venv and drives the real protocol — 5 tools, every one annotated read-only and documented, instructions intact

The last one is the important one: it catches a package that installs cleanly and dies on its first request, which is a failure no unit test sees.

License

MIT — see LICENSE.

Available Tools

14 tools
check_table_freshnessA
Read-onlyIdempotent

Report when tables were last written, to catch stale or dead sources.

Several plausible-looking tables on this platform stopped being updated without being dropped, so a query against one silently returns old data. Check before trusting a table you have not used before.

Free — reads table metadata only, scanning no data.

Args: dataset_id: The dataset to check, e.g. "events_raw". table_id: A single table to check. Omit to report every table in the dataset, which is the faster way to spot a dead one. environment: Which configured BigQuery environment to use. Omit to use the default. Call list_environments to see what exists.

ParametersJSON Schema
NameRequiredDescriptionDefault
table_idNo
dataset_idYes
environmentNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), but the description adds genuinely new context: it reads metadata only and scans no data, so the agent knows there is no cost or data-scan risk. It explains the failure mode it detects (silent stale data) but says nothing about the shape or size of the report.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and the motivating problem, then the arguments. The middle rationale paragraph is slightly longer than strictly needed but earns its place by justifying when to call the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter, read-only metadata tool with no output schema, the definition covers purpose, trigger, cost profile, and every argument. The one gap is that it never hints at what the freshness report contains, which an agent would have to discover at runtime.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden and does so well: dataset_id gets a concrete example, table_id is explained including the omit-to-report-all behavior and why that is preferred, and environment documents the default plus the list_environments route. All three parameters gain meaning beyond the bare schema types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('report when tables were last written') and immediately frames the goal ('to catch stale or dead sources'). This distinguishes it from siblings like list_tables and get_table_schema, which enumerate structure rather than freshness.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('Check before trusting a table you have not used before') and a routing hint ('Call list_environments to see what exists') for the environment argument. It does not name a sibling alternative for the general case, which keeps it just shy of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_code_assets_using_tableA
Read-onlyIdempotent

Find which Colab notebooks and saved queries reference a table.

The question to ask before changing or dropping a table: list_scheduled_queries says what writes it, this says who reads it.

Unlike the other tools here this one opens every asset it considers, which costs Dataform read quota. It is bounded by max_assets and reports how much of the project it actually covered -- a result is evidence about the assets scanned, never proof that nothing else uses the table.

Args: table: Table name to search for. A bare name matches any qualification; 'dataset.table' or a fully-qualified name narrows it. environment: Which configured environment to read. Omit for the default. asset_type: Restrict to 'sql', 'notebook' or 'data_canvas'. max_assets: Ceiling on how many bodies to read.

ParametersJSON Schema
NameRequiredDescriptionDefault
tableYes
asset_typeNo
max_assetsNo
environmentNo

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only/idempotent/non-destructive safety profile, but the description adds genuinely non-obvious behavior: it opens every candidate asset and consumes Dataform read quota, and its output is a bounded sample rather than proof of completeness. The only missing piece is any note about result shape beyond the coverage caveat, so a 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening line is front-loaded and the Args block is efficient, but the middle paragraph is somewhat convoluted ('Unlike the other tools here this one opens every asset it considers...') and the parenthetical coverage caveat runs long. Every sentence still earns its place, just not with maximal crispness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still tells the agent that results report coverage of the project scanned and must not be read as proof of non-usage, which is the key interpretive context. All four parameters are documented and the quota cost of calling it is disclosed. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the whole burden and does: it explains table-name matching semantics (bare name = any qualification, qualified name narrows), the environment selector, the asset_type filter enumerating 'sql', 'notebook', 'data_canvas' (values absent from the schema), and max_assets as a ceiling on bodies read.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Find which Colab notebooks and saved queries reference a table') and immediately distinguishes itself from siblings via the write-vs-read contrast with ``list_scheduled_queries``. An agent can pick this tool over list_code_assets or get_code_asset without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames the use case ('The question to ask before changing or dropping a table') and names the complementary tool with its role ('list_scheduled_queries says what writes it, this says who reads it'). No inference needed about when this tool is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_code_assetA
Read-onlyIdempotent

Return one Colab notebook or saved query's contents, by name or id.

Notebook outputs are stripped -- across 52 real notebooks they were 77% of the bytes, and none of the logic.

Args: asset: Display name (as shown in BigQuery Studio) or the asset id. environment: Which configured environment to read. Omit for the default.

ParametersJSON Schema
NameRequiredDescriptionDefault
assetYes
environmentNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), so the bar is lower, yet the description still adds real value: it discloses that notebook outputs are stripped from the returned contents, with a quantified rationale. It does not mention pagination or size limits, but the return-content disclosure is genuinely useful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in one sentence, then Args are clearly delimited. The aside about outputs being 77% of bytes across 52 notebooks is slightly chatty but it justifies the stripping behavior rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description would need to explain the return value; it characterizes the return as notebook or saved-query contents with outputs stripped. Combined with the environment scoping, an agent has enough to call it correctly, though the returned structure and error cases are not described.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning, and it does: 'asset' is explained as either a display name (as shown in BigQuery Studio) or an asset id, and 'environment' as the configured environment to read, with the default-bypass behavior stated. Only the accepted format of environment identifiers is left implicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Return') plus precise resource ('one Colab notebook or saved query's contents') and retrieval keys ('by name or id'). An agent can immediately distinguish this single-item fetch from the sibling list_code_assets or find_code_assets_using_table without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the retrieval use case and clarifies that environment can be omitted for the default, but it never states when to prefer this over list_code_assets or find_code_assets_using_table, nor any prerequisites. Usage is inferable but not routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_notebook_scheduleA
Read-onlyIdempotent

Get one scheduled notebook in full: its cron, its notebook, and recent runs.

Call this after list_notebook_schedules to see why a scheduled notebook is failing. The error text for each failed run is included, and notebook_id is a code asset id — pass it to get_code_asset to read the code that failed.

Args: schedule: The schedule's name or its id from list_notebook_schedules. runs: How many recent runs to include, newest first. environment: Which configured environment to read. Omit for the default.

ParametersJSON Schema
NameRequiredDescriptionDefault
runsNo
scheduleYes
environmentNo

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotent/non-destructive, but the description adds real behavioral context: error text for each failed run is included, runs are returned newest first, and notebook_id is revealed to be a code asset id. It stops short of mentioning auth requirements or limits, but for a read tool this is well above the annotation-covered baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-sentence summary, then a short usage paragraph, then a compact Args block. Every sentence supplies either routing or parameter meaning; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly compensates by describing the return contents (cron, notebook, recent runs with error text). Combined with named entry points and a follow-up tool, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage the description carries the full burden and does so for all three params: schedule accepts a name or an id from list_notebook_schedules, runs sets the count and ordering ('newest first'), and environment selects a configured environment with an explicit omit-for-default rule.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get one scheduled notebook in full') and enumerates the payload (cron, notebook, recent runs). It is clearly distinguishable from the sibling list_notebook_schedules, which returns the collection rather than one schedule.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'Call this after list_notebook_schedules to see why a scheduled notebook is failing,' and hands off the next step ('pass it to get_code_asset to read the code that failed'). The trigger condition and the alternative tool chain are both named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_scheduled_queryA
Read-onlyIdempotent

Get one scheduled query in full: its SQL, destination, and recent runs.

Call this after list_scheduled_queries to see why a query is failing, or what SQL actually produces a table.

Args: query: The scheduled query's name, or the id from list_scheduled_queries. runs: How many recent runs to include, newest first. environment: Which configured environment to look in.

ParametersJSON Schema
NameRequiredDescriptionDefault
runsNo
queryYes
environmentNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds useful behavior beyond that: the return payload (SQL, destination, recent runs) and the newest-first ordering of runs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the retrieval scope, then a compact when-to-use sentence, then a terse Args block that earns its place given the empty schema descriptions. No sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully previews what is returned (SQL, destination, recent runs). Combined with parameter coverage and usage routing, it is nearly complete; only defaults and run-detail shape are left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the full burden, and it defines all three parameters: query as name or the id from list_scheduled_queries, runs as the count of recent runs newest-first, and environment as the configured environment to search. It does not note the runs/environment defaults, but the semantics are otherwise well covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get one scheduled query in full') and enumerates what the full retrieval contains: SQL, destination, and recent runs. This cleanly distinguishes it from the sibling list_scheduled_queries, which returns summaries rather than one query's details.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger conditions – call after list_scheduled_queries to diagnose failures or inspect the SQL behind a table – which is clear sequencing guidance. It stops short of stating when not to use it (e.g., for bulk inspection), but the context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_table_schemaA
Read-onlyIdempotent

Get a table's columns, partitioning, size and freshness. Free — scans no data.

Call this before writing a query, for two reasons beyond column names:

  • partitioning says whether a WHERE clause can actually limit the scan. A date-shaped column name does NOT mean the table is partitioned; if this field is null, every query reads the whole table.

  • Nested columns are expanded to dotted paths and flagged repeated, which is what tells you a column needs UNNEST.

Args: dataset_id: The dataset, e.g. "events_raw". table_id: The table or view name. environment: Which configured BigQuery environment to use. Omit to use the default. Call list_environments to see what exists.

ParametersJSON Schema
NameRequiredDescriptionDefault
table_idYes
dataset_idYes
environmentNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the bar is lower, yet the description adds real behavioral context: it is free and scans no data (cost profile), partitioning=null means every query reads the whole table, and nested columns appear as dotted paths flagged `repeated` to signal UNNEST needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-line summary, then two tightly scoped bullets that each teach a distinct query-writing consequence, then an Args block. No filler sentences; every clause changes how the agent behaves.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must convey returns — and it names all four returned facets (columns, partitioning, size, freshness) plus the `repeated` flag semantics. Nothing an agent needs to call and interpret this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden, and it does: dataset_id ('events_raw'), table_id (table or view name), and environment (which configured BigQuery environment, omit for default, see list_environments). Every parameter gets meaning and an example.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource plus the exact payload: 'Get a table's columns, partitioning, size and freshness.' This distinguishes it cleanly from siblings like list_tables (enumerates) and check_table_freshness (freshness only), so an agent can route without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to call it ('Call this before writing a query') and why, and points to list_environments for discovering valid environment values. It lacks an explicit when-not / alternative comparison against check_table_freshness, but the pre-query positioning is clear guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_code_assetsA
Read-onlyIdempotent

List Colab notebooks and saved queries in BigQuery Studio.

Use this for anything the user calls a Colab notebook, Colab Enterprise notebook, "colab script", BigQuery notebook, saved query or data canvas -- BigQuery Studio stores all of them as code assets and this lists them all.

Free -- this reads metadata only and never opens an asset. Bodies are what cost quota, so filter here first and open individual assets afterwards.

Args: environment: Which configured environment to read. Omit for the default. asset_type: Restrict to one of 'sql', 'notebook', 'data_canvas'. Saved queries usually outnumber notebooks by a wide margin, so this is the difference between a readable answer and 600 rows. name_contains: Case-insensitive substring match on the display name. limit: Maximum assets to return.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
asset_typeNo
environmentNo
name_containsNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, but the description adds substantive behavior: it is free, reads metadata only, never opens an asset, and that asset bodies cost quota. The cost/quota disclosure is genuinely new information that shapes agent behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then usage, then cost note, then args in a clean block. The alias enumeration is long but justified for vocabulary mapping; a stray 'this lists them all' is minor redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a free list tool with no output schema, the description covers what is returned (notebooks, saved queries, data canvases) and the cost profile. It stops short of describing the return shape or pagination, but nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and no enums are declared, so the description carries the full burden and does so: it supplies the enum values ('sql', 'notebook', 'data_canvas') absent from the schema, explains name_contains is case-insensitive, and gives the rationale for asset_type (row-volume filtering). This is meaning well beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (Colab notebooks and saved queries / code assets) with explicit scope. It collapses several user-facing terms into the single BigQuery Studio concept, so an agent can map user vocabulary to this tool without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly signals when to use it ('Use this for anything the user calls a Colab notebook...') and sequences it against the open step ('filter here first and open individual assets afterwards'), implicitly routing to get_code_asset. No explicit when-not or named exclusion of siblings, so not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_datasetsA
Read-onlyIdempotent

List the BigQuery datasets available in the data platform project.

Call this first to discover what data exists. Free — scans no data.

Args: environment: Which configured BigQuery environment to use. Omit to use the default. Call list_environments to see what exists.

ParametersJSON Schema
NameRequiredDescriptionDefault
environmentNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent and non-destructive, so the safety profile is covered. The description adds genuinely new behavior: the call is free and scans no data, which is material cost information an agent can use when planning.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences plus a one-parameter Args block; the discovery/cost rationale is front-loaded and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-param, read-only listing tool with no output schema, the description covers purpose, cost, and the one parameter well. It could briefly say what a returned dataset entry looks like, which is the only remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must carry the parameter burden, and it does: it explains that environment selects a configured BigQuery environment, that omitting it uses the default, and points to list_environments for valid values. It doesn't state accepted format/name conventions, so not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the BigQuery datasets') scoped to the data platform project, which cleanly separates it from sibling list_tables/list_code_assets. It does not explicitly contrast with list_tables, so it falls just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context ('Call this first to discover what data exists') and routes the agent to list_environments for discovering valid environment values. No explicit when-not-to-use guidance, but the intended entry-point role is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_environmentsA
Read-onlyIdempotent

List the configured BigQuery environments and which one is the default.

Call this when the user names an environment you have not seen, or when a question could plausibly be about more than one. Free — reads only this server's configuration.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive and openWorldHint=false, so the description only needs to add context. It does: 'Free — reads only this server's configuration' tells the agent there is no API cost and that the read scope is server config, not BigQuery itself.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero waste, with the core action front-loaded and the invocation heuristic immediately after. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless, annotated read-only config listing with no output schema, this is nearly complete. It could optionally hint at the returned fields (name, default flag), but annotations and the description together cover what an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the baseline is 4. There is nothing for the description to disambiguate and it correctly avoids inventing parameter guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (list) and resource (configured BigQuery environments) and adds the distinguishing detail that it also reports which is the default. The resource is clearly distinct from the sibling list_datasets/list_tables tools, so an agent can pick it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete trigger conditions: call it when the user names an unseen environment, or when a question could span more than one. No explicit when-not or named alternative, but the context for invoking it is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_notebook_runsA
Read-onlyIdempotent

List individual scheduled-notebook runs, by default the failed ones.

Use this for "what has been failing?" across every scheduled notebook at once, rather than per schedule. Runs are returned newest first.

Outcome cannot be filtered server-side — the API rejects a jobState filter — so this reads the runs in the window and filters here. That makes lookback_days the cost control: each 100 runs is one API call.

Args: status: 'failed' (default), 'succeeded', 'running', or 'all'. environment: Which configured environment to read. Omit for the default. schedule: Restrict to one schedule, by its name or id. lookback_days: How far back to read. Defaults to 30. limit: Maximum runs to return.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
statusNofailed
scheduleNo
environmentNo
lookback_daysNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds substantial non-obvious behavior: the API rejects a jobState filter so filtering happens client-side, results are newest-first, and lookback_days is the cost lever (100 runs per API call).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose and default behavior, then the rationale, then compact arg notes. Despite the length, every sentence carries actionable information that is absent from the structured fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description covers ordering, filtering, and cost but not the shape of a returned run record. Otherwise complete for a read-only list tool with five undocumented parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the full burden and does: it enumerates the status values, explains environment defaulting, says schedule accepts a name or id, and clarifies lookback_days and limit semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (scheduled-notebook runs) with an explicit default scope (failed). It draws the boundary against the sibling list_notebook_schedules by contrasting 'across every scheduled notebook at once, rather than per schedule.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit triggering question ('what has been failing?') and the condition that favors this tool over the per-schedule alternative. The trade-off against sibling tools is stated rather than inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_notebook_schedulesA
Read-onlyIdempotent

List scheduled Colab notebooks with how many recent runs passed or failed.

This is the health overview for scheduled notebook work: what is scheduled, whether it is active or paused, and — the part that is otherwise invisible — how its actual runs have been going.

Do not read a schedule's own state as health. A schedule reports its last scheduled run as "OK" when it successfully launched the notebook, whether or not the notebook then failed; on this platform every schedule says OK while hundreds of runs have failed. The pass/fail numbers here come from the execution jobs, which is the only place the outcome exists.

Args: environment: Which configured environment to read. Omit for the default. state: Restrict to 'active' or 'paused'. A paused schedule that used to fail is a common find — someone paused it instead of fixing it. name_contains: Case-insensitive substring match on the schedule name. lookback_days: How far back to read runs for the pass/fail counts. Defaults to 30 so monthly schedules show at least one run. Larger windows cost proportionally more (90 days is roughly 23 API pages). Set to 0 to skip run history entirely and just list what exists.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateNo
environmentNo
lookback_daysNo
name_containsNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover readOnly/idempotent/openWorld, so the bar is lower, yet the description adds genuinely new behavioral context: the platform quirk that every schedule reports OK regardless of run failure, where pass/fail data actually comes from (execution jobs), and the cost model of lookback_days (90 days ≈ 23 API pages).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, and the Args block is efficient. The middle paragraph repeats the 'schedule state isn't health' point across two sentences, which is a minor redundancy but justified given how counterintuitive the caveat is.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must convey returns, and it does — list of schedules with active/paused state and pass/fail counts. Combined with full parameter documentation and the semantic caveat, an agent has everything needed to call and interpret this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and none of the four parameters have enums or descriptions in the schema, so the description carries the full burden — and it does: environment default behavior, state values 'active'/'paused' with an interpretation tip, name_contains matching semantics, and lookback_days default rationale plus the 0-to-skip behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List scheduled Colab notebooks') plus the distinguishing payload ('how many recent runs passed or failed'). This separates it from get_notebook_schedule (single schedule detail) and list_notebook_runs (raw run listing) without the agent needing to open any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Frames the tool as 'the health overview' and explicitly warns against the common misinterpretation (schedule self-reported 'OK'). It reads as clear when-to-use context but does not name sibling alternatives or state explicit exclusions by tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_scheduled_queriesA
Read-onlyIdempotent

List scheduled queries: what they write, when they run, and their state.

Use this to answer "what populates this table?" and "why is this table stale?" — a disabled or failing scheduled query is the usual cause, and check_table_freshness can see the staleness but not the reason.

The SQL is not included here; call get_scheduled_query for one of them.

Args: dataset: Only queries writing into this destination dataset. contains: Only queries whose name contains this text. include_disabled: Keep disabled queries in the result. They are the most likely explanation for a table that stopped updating, so this defaults to True. environment: Which configured environment to look in. Scheduled queries are regional, so this must be the environment whose location holds them.

ParametersJSON Schema
NameRequiredDescriptionDefault
datasetNo
containsNo
environmentNo
include_disabledNo

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered. The description adds real context beyond them: the SQL is deliberately excluded, disabled queries are retained because they are the likeliest cause of stalled tables, and scheduled queries are regional so environment must match. Return format and pagination are not described, keeping it just short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The lead sentence and the usage paragraph are front-loaded and dense with decision-relevant information. The Args block is longer than strictly necessary for some entries, but each line adds semantics absent from the schema, so there is little filler to cut.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no enum metadata, the description compensates well by summarizing the shape of the result (write target, schedule, state) and by pointing to get_scheduled_query for the detail it omits. It stops short of describing pagination or result limits, which an agent listing potentially many queries would benefit from knowing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the entire burden and does so for all four parameters. It clarifies that dataset filters on the destination dataset (not the source), that contains matches on query name, that environment is a location constraint driven by regionality, and it explains both the value and default of include_disabled rather than merely restating it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List scheduled queries') and immediately defines the scope of returned data ('what they write, when they run, and their state'). It also separates itself from get_scheduled_query, which is where the SQL lives, so an agent can distinguish the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit use cases in the user's own words ('what populates this table?', 'why is this table stale?') and names the alternative tool check_table_freshness with the reason it is insufficient ('can see the staleness but not the reason'). It also issues a clear redirect: call get_scheduled_query when the SQL itself is needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_tablesA
Read-onlyIdempotent

List tables and views inside a dataset. Free — scans no data.

Args: dataset_id: The dataset to inspect, e.g. "events_raw". environment: Which configured BigQuery environment to use. Omit to use the default. Call list_environments to see what exists.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataset_idYes
environmentNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive, so the safety profile is covered. The description adds genuinely new behavioral context: that the call is free and scans no data — a cost/side-effect trait not present in the annotations. It does not describe result shape or pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the key 'free/no scan' fact are front-loaded in the first two sentences, and the args block is tight with no filler. Every line earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, read-only list tool with no output schema, the definition covers purpose, cost, and both parameters adequately. The only minor gap is not stating what the returned entries contain (e.g., names only vs. metadata), but this is not critical for invoking it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden and largely meets it: it explains dataset_id (with a concrete example) and environment (default behavior plus a pointer to list_environments). Both parameters get meaning beyond their bare names and types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: 'List tables and views inside a dataset.' It's clear what the tool returns at a high level, though it doesn't explicitly differentiate itself from siblings like list_datasets or get_table_schema beyond the word 'inside a dataset'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a cost-based usage signal ('Free — scans no data') and routes the agent to a sibling for the environment parameter ('Call list_environments to see what exists'). No explicit when-not-to-use guidance, but the selection context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_queryA
Read-onlyIdempotent

Run a read-only (SELECT/WITH) SQL query against BigQuery and return rows.

Cost safety: the query is ALWAYS dry-run first to estimate how much data it will scan. If that estimate is above the warning threshold, the query does NOT run — instead this returns status: "confirmation_required" with the estimated size and cost. Stop there, tell the user the estimated scan and cost, and ask. Only re-call with confirm_expensive=true once they have agreed: that flag records the user's decision, not yours. Queries above the hard cap never run, even with confirmation.

Always fully-qualify tables as <project>.<dataset>.<table>, and check get_table_schema first — a WHERE clause only limits the scan on a table that is actually partitioned.

Args: sql: A SELECT (or WITH ... SELECT) query. max_rows: Max rows to return, to keep responses small. 0 (the default) uses the server's configured limit. confirm_expensive: Set True only after the user has agreed to a query previously flagged as costly. Leave False for the first attempt. environment: Which configured BigQuery environment to query. Omit to use the default. Call list_environments to see what exists.

ParametersJSON Schema
NameRequiredDescriptionDefault
sqlYes
max_rowsNo
environmentNo
confirm_expensiveNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint/destructiveHint but don't convey the two-phase dry-run cost gate, the warning-threshold stop condition, the hard cap that never runs even with confirmation, or the confirmation_required status payload. This is exactly the behavioral context annotations cannot express, and the description delivers it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose then cost-safety then qualification advice then args. Every sentence earns its place. Slightly prose-heavy in the cost paragraph, but the detail is load-bearing for safe invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-param query tool with no output schema, the description covers execution semantics, cost gating, the confirmation_required return shape, table qualification, and environment selection. Nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the full burden. It documents all four params: sql (SELECT/WITH), max_rows (0=server default), confirm_expensive (only after user agreement, records user decision not agent's), and environment (omit for default, list_environments to discover). Semantic value well beyond bare schema titles/defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run a read-only SELECT/WITH SQL query against BigQuery and return rows'). The read-only constraint and SQL dialect are explicit, distinguishing it from sibling listing/schema tools that don't execute queries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit workflows: check get_table_schema first before writing WHERE clauses; call list_environments to see environments; the dry-run confirmation protocol states exactly when to stop, what to tell the user, and when to re-call with confirm_expensive. Names siblings (get_table_schema, list_environments) and the condition for each.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 14 tool updatesv0.4.0
    • First observedcheck_table_freshness
    • First observedfind_code_assets_using_table
    • First observedget_code_asset
    • First observedget_notebook_schedule
    • First observedget_scheduled_query
    • First observedget_table_schema
    • First observedlist_code_assets
    • First observedlist_datasets
    • First observedlist_environments
    • First observedlist_notebook_runs
    • First observedlist_notebook_schedules
    • First observedlist_scheduled_queries
    • First observedlist_tables
    • First observedrun_query

TDQS

A4.4/5.0

Scored across 14 tools

Disambiguation5/5

Each tool targets a distinct resource and action: discovery (list_environments/datasets/tables), schema/freshness, code assets, scheduled SQL queries, and scheduled notebooks. The list/get pairs and the freshness-vs-scheduled-query distinction are explicitly spelled out, leaving no real overlap.

Naming Consistency5/5

Consistent snake_case verb_noun pattern throughout (list_datasets, get_table_schema, check_table_freshness, run_query). find_code_assets_using_table is longer but follows the same convention.

Tool Count5/5

14 tools sit in the ideal band and each earns its place, covering distinct phases of a BigQuery workflow (discovery, schema inspection, cost-safe querying, schedule diagnosis) without redundancy.

Completeness4/5

The read-and-query surface is thorough: environments, datasets, tables, schema, freshness, scheduled queries, notebook schedules/runs, code assets, and ad-hoc query execution. Write/DDL operations are absent, which appears intentional for a read-only query server, but leaves a minor gap if mutation were ever expected.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    A Model Context Protocol server that provides access to BigQuery. This server enables LLMs to inspect database schemas and execute queries.
    130
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables LLMs to interact with Google BigQuery by inspecting database schemas, listing tables, and executing SQL queries. This server facilitates seamless data analysis and management through natural language via the Model Context Protocol.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    A read-only BigQuery MCP server with auto-LIMIT injection, dry-run cost guard, and ADC authentication. Allows safe SQL querying of BigQuery by LLMs without risk of data modification or unexpected costs.
    26 PyPI
    1
    MIT