wrds
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@wrdsExport annual revenue and net income for German firms 2010-2022 to CSV"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
WRDS Research for Claude
Ask Claude in plain English whether data exists on WRDS, and have it pull it for you.
"Do we have P&L data for European companies 2004–2022 by country and industry? If so, export it."
Claude searches the WRDS catalog, tells you what's available (source, firms, years, countries), then queries or exports the data — with the WRDS AI policy enforced.
Not affiliated with WRDS or the Wharton School. You need your own WRDS account; you only see the libraries your institution subscribes to.
Before you start: the WRDS AI policy
Row-level data may only be sent to a protected AI tool (vendor contractually doesn't train on it). Check that your Claude plan qualifies (Team / Enterprise / API usually do; ask your WRDS representative if unsure).
Only these sources are AI-allowed: CRSP, S&P (Compustat, Capital IQ), LSEG (Worldscope, Datastream, I/B/E/S), MSCI, Audit Analytics, Revelio, WRDS-created data.
For every other library the server returns metadata and counts only; exports are saved to disk without a preview.
Related MCP server: Rice Stock Data Portal MCP Server
Install
Requires uv (brew install uv on macOS, or see the uv docs).
1. Save your WRDS login (once, in a terminal — the password is typed there, never into Claude):
uvx --from git+https://github.com/hosseinzk/wrds-mcp wrds-mcp-setupIt asks for your WRDS username and password, tests the connection (approve the Duo push if one
arrives) and saves them to ~/.pgpass (Windows: %APPDATA%\postgresql\pgpass.conf).
2a. Claude Code — install the plugin (MCP server + a research skill):
/plugin marketplace add hosseinzk/wrds-mcp
/plugin install wrds@wrds-toolsRestart Claude Code, then ask "Which WRDS databases do I have access to?"
2b. Claude Desktop / other MCP clients — install the server and add it to your MCP config (Claude Desktop: Settings → Developer → Edit config):
uv tool install git+https://github.com/hosseinzk/wrds-mcp{ "mcpServers": { "wrds": { "command": "wrds-mcp" } } }Exports and the query log go to ~/wrds-data (set WRDS_OUTPUT_DIR to change it).
Tools
Tool | What it does |
| Libraries your account can access, vendor, AI-allowed flag, sample/trial flag |
| Search table/column names and descriptions across all libraries |
| Tables, approximate row counts, columns and descriptions |
| Single-row counts / year ranges on any library (numbers only for non-AI-allowed ones) |
| Read-only SELECT, up to 500 rows, AI-allowed sources only |
| Full result to CSV / Parquet / Excel in your output folder |
Every query is logged to <output folder>/logs/queries.log. Connections are read-only.
Examples
"Which databases do I have access to? Which are full and which are just samples?"
"Is quarterly R&D spending available for Tesla 2015–2024?"
"Export annual revenue, EBIT and net income for all German and French listed firms 2010–2022, with GICS sector."
"Median EBITDA margin by country and GICS sector for Europe, 2004–2022, in EUR."
Available Tools
7 toolscheck_availabilityA
Run an aggregate-only availability check (counts, distinct counts, min/max years) on any accessible library, including ones where row-level data is not AI-allowed. The query must return a single summary row, e.g. SELECT count(*), count(distinct gvkey), min(fyear), max(fyear) FROM comp.g_funda WHERE ...
| Name | Required | Description | Default |
|---|---|---|---|
| sql | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and it does disclose a key behavioral constraint: the query must return a single summary row, implying rejection of multi-row results. It does not cover permissions, what happens on violation, or rate/access behavior, so the disclosure is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and constraint come first, followed by a compact illustrative example; every element earns its place. The example sentence is long but directly encodes the required shape of the input rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the description covers purpose, input shape, and the aggregate-only constraint. The main remaining gap is the lack of explicit routing versus query/export_query, but for a one-parameter tool this is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single sql parameter, so the description must compensate, and it does by giving concrete example syntax (count(*), count(distinct gvkey), min/max fyear, FROM comp.g_funda WHERE ...) plus the single-summary-row constraint on the input. This meaningfully clarifies what belongs in the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('run an aggregate-only availability check') and pins the scope ('on any accessible library, including ones where row-level data is not AI-allowed'). This distinguishes it reasonably from the generic query sibling, though it never names the alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the 'aggregate-only' framing and the note about libraries where row-level data is not AI-allowed, which hints this is the way to probe restricted libraries. However, there is no explicit when-not or named alternative (e.g., use query for row-level retrieval), leaving the agent to infer the boundary with query/export_query.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_tableB
Show columns, types and WRDS column descriptions (where available) for a table.
| Name | Required | Description | Default |
|---|---|---|---|
| table | Yes | ||
| library | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Show' implies a read-only metadata lookup with no side effects, and it discloses that column descriptions are returned only 'where available', but it says nothing about permissions, cost, or behavior for missing/invalid tables.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, tightly written sentence with the key output contents front-loaded and no filler. Nothing needs trimming and nothing is buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return-value detail is not required, and the one-liner adequately conveys the tool's intent. However, with 0% parameter coverage and no annotations, the unresolved 'library' semantics leave a real gap for an agent trying to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and neither parameter is documented. The description only implies 'table'; the required 'library' parameter is never explained, nor is the expected naming format (e.g., whether it takes a library code like 'comp').
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Show') and resource ('columns, types and column descriptions for a table'), which is a clear describe/schema-inspection operation. It is naturally differentiated from siblings like list_tables and query, though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus list_tables, search_catalog, or query, nor any mention of prerequisites. Usage is only weakly implied — an agent must infer it is for inspecting a known table's schema before querying.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_queryA
Run a read-only SELECT with no row limit and save it to the output folder (default ~/wrds-data). fmt: csv, parquet or xlsx. Returns the path, row count and columns; a 5-row preview is included only for AI-allowed sources.
| Name | Required | Description | Default |
|---|---|---|---|
| fmt | No | csv | |
| sql | Yes | ||
| filename | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does reasonably well: it discloses read-only semantics, that no row limit is applied, the default output location, accepted formats, and that a 5-row preview is granted only for AI-allowed sources. It omits overwrite behavior, permission requirements, and whether the filename needs an extension.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core action. The 'fmt: csv, parquet or xlsx' fragment is terse but slightly choppy against the prose around it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, yet the description still summarizes them (path, row count, columns) without harm. Remaining gaps are minor operational details such as overwrite behavior and filename format conventions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds value for `fmt` by enumerating csv, parquet, and xlsx, and clarifies `sql` is a read-only SELECT, but `filename` gets no guidance (path vs. name, extension handling) beyond the default folder mention.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: run a read-only SELECT and save the result to a file. It implies the distinction from the sibling `query` tool by emphasizing persistence to the output folder and no row limit, but it never names `query` explicitly, so the differentiation is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied ('no row limit', 'save it to the output folder') which suggests this is for large result sets the agent wants persisted, but it never states when to use this over `query` or `check_availability`. No exclusions or prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_librariesA
List WRDS libraries this account can access, with the vendor and whether data rows may be returned to the AI under the WRDS AI policy. Optional substring filter.
| Name | Required | Description | Default |
|---|---|---|---|
| contains | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry behavioral weight; it discloses that results are account-scoped and include vendor plus an AI data-return policy flag, which is useful context. However, it says nothing about pagination, whether the operation is read-only (implied by 'list'), or rate/auth constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence that names the resource, scope, return fields, and the filter in minimal words. Nothing is redundant or padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value structure is documented elsewhere, and the description still summarizes the notable return fields. The only gap is lack of routing guidance among the six sibling catalog/query tools, which is minor for a simple enumerate operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single 'contains' parameter is undocumented in the schema, but the description rescues it by explaining it as an optional substring filter. That adds the matching semantics ('substring', 'optional') the schema lacks, though no case-sensitivity or field-matched details are given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (list WRDS libraries) plus scope (those this account can access) and the fields returned (vendor, AI data-return policy). It is distinguishable from list_tables/describe_table by resource, but does not explicitly name or contrast with any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, prerequisites, or alternative is given. The mention of an optional substring filter hints at usage but does not explain when to prefer this over search_catalog or when to enumerate libraries before calling list_tables.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tablesB
List tables in a WRDS library with approximate row counts.
| Name | Required | Description | Default |
|---|---|---|---|
| library | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses that row counts are 'approximate' rather than exact, and 'List' implies a read-only operation, but permissions, result size, or pagination behavior are not addressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no wasted words, and the resource and scope are front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers the core purpose plus the approximation caveat. It could say more about how libraries relate to the sibling list_libraries tool, but nothing critical for a simple read is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter, library, has 0% schema description coverage, and the description only alludes to it via 'in a WRDS library'. It adds minimal semantic meaning beyond the bare parameter name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (tables) with the scope qualifier 'in a WRDS library'. It is clear what the tool returns, though it does not explicitly distinguish itself from siblings like describe_table or search_catalog.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no exclusions, and no mention of alternatives such as search_catalog or describe_table. Usage is only weakly implied by the verb 'List'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
queryA
Run a read-only SELECT and return up to limit rows (max 500) as CSV. Only
allowed when every referenced library is AI-allowed under the WRDS policy.
| Name | Required | Description | Default |
|---|---|---|---|
| sql | Yes | ||
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose key traits: read-only, a hard 500-row ceiling, CSV output, and a policy gate on referenced libraries. It omits error behavior (e.g. whether exceeding 500 errors or clamps) and any rate-limit or permission detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the operational facts (read-only, row cap, CSV) are front-loaded ahead of the policy caveat. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained further, and the description already covers the essentials for a query tool. The remaining gap is the undocumented sql parameter contract and ambiguous overflow behavior past the 500-row cap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so both parameters are undocumented in the schema. The description rescues 'limit' by supplying the 500 max that the schema lacks, but says nothing about the required 'sql' parameter's dialect, allowed statements, or format. Partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Run a read-only SELECT'), the return shape ('up to limit rows as CSV'), and the scope constraint (max 500). It is distinguishable from export_query and list_tables at a glance, though it never names a sibling to sharpen the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one hard precondition ('only allowed when every referenced library is AI-allowed under the WRDS policy'), which is genuine usage guidance. However, it never says when to prefer this over export_query, or what to do for larger result sets, leaving the main routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_catalogB
Search table names, column names and column descriptions across all accessible libraries for a keyword (e.g. 'revenue', 'nace', 'sic', 'orbis'). Use this first to answer 'is X available?' questions.
| Name | Required | Description | Default |
|---|---|---|---|
| keyword | Yes | ||
| max_results | No | ||
| library_contains | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the access scope ('all accessible libraries') and that three distinct field types are matched, which is meaningful behavioral context for a read tool. It omits anything about result limits, ordering, or how hits are returned, and there is no explicit read-only statement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core purpose and followed by a directive on when to reach for it. Every clause carries information; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The existence of an output schema means return values need not be described, lowering the bar. Still, for a three-parameter tool with zero schema coverage and no annotations, the description leaves library_contains and max_results unexplained and gives no guidance versus the overlapping check_availability sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it only partially does: the example keywords ('revenue', 'nace', 'sic', 'orbis') clarify what 'keyword' should contain. The other two parameters, library_contains and max_results (default 200), receive no explanation at all, leaving their meaning and effect undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Search) and precisely enumerates the searched resource fields (table names, column names, column descriptions) plus the scope (all accessible libraries), which is far more informative than the bare name. It does not, however, differentiate itself from the sibling check_availability, which the phrasing 'is X available?' arguably overlaps with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this first to answer "is X available?" questions' gives an implied entry-point role and usage context. But it never mentions the sibling check_availability or when that tool should be preferred instead, leaving the agent to guess at overlap between the two.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
check_availability - First observed
describe_table - First observed
export_query - First observed
list_libraries - First observed
list_tables - First observed
query - First observed
search_catalog
TDQS
Scored across 7 tools
The tools are clearly divided into metadata discovery (list_libraries, list_tables, describe_table, search_catalog) and data access (check_availability, query, export_query). Within each group, purposes are distinct, with check_availability explicitly aggregate-only and query/export_query differentiated by output destination and row limits. No two tools appear to do the same thing.
All tool names use snake_case, and most follow a verb_noun pattern (e.g., list_libraries, describe_table, export_query). The tool 'query' is a lone verb without a noun, a minor deviation from the otherwise consistent pattern.
Seven tools is well-scoped for a read-only data access server, covering discovery and retrieval without bloat. Each tool earns its place and none feel redundant.
The set covers the full read-only data lifecycle: discovering libraries, tables, and columns; searching across catalogs; checking availability under AI policy; and running bounded or unbounded queries. No obvious gaps exist for the stated domain.
Maintenance
Related MCP Connectors
Query your warehouse or a CSV with Claude/ChatGPT over MCP, governed by table-level ACL + audit.
Query 40 databases from Claude, ChatGPT, or Cursor — on any device. Read-only, encrypted, audited.
List datasets, schemas, run APL queries, and use prompts for exploration, anomalies, and monitoring.
Connect your ads, shop, analytics, social, CRM and finance platforms once, then let Claude, ChatGPT, Cursor or any MCP client read, join and explain your numbers. Public statistics from the World Bank, IMF, Eurostat, OECD, WHO and SEC filings come as context, searchable and chartable from the same tools. Read-only by design, every number carries its source.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceProvides database access capabilities to Claude with support for both SQLite and SQL Server databases. It enables users to manage schemas, execute SQL queries, and export results through natural language interaction.997 npmMIT
- FlicenseNot gradedqualityDmaintenanceEnables querying Rice Business Stock Market Data Portal using natural language through Claude Desktop, allowing users to ask about financial ratios, stocks, and companies.-
- AlicenseNot gradedqualityBmaintenanceEnables Claude to run SQL queries against Infor Compass (Data Fabric) directly from chat, with automatic export of large result sets to Excel.2MIT
- FlicenseNot gradedqualityDmaintenanceEnables Claude to query, create, update, and delete Salesforce records directly using SOQL and SOSL, with tools for metadata exploration and bulk operations.-