Cryo MCP Server
低温 MCP 🧊
Cryo区块链数据提取工具的模型完成协议 (MCP) 服务器。
Cryo MCP 允许您通过实现 MCP 协议的 API 服务器访问 Cryo 强大的区块链数据提取功能,从而可以轻松地从任何兼容 MCP 的客户端查询区块链数据。
对于 LLM 用户:SQL 查询工作流指南
使用此 MCP 服务器对区块链数据运行 SQL 查询时,请遵循以下工作流程:
使用
query_dataset下载数据:result = query_dataset( dataset="blocks", # or "transactions", "logs", etc. blocks="15000000:15001000", # or use blocks_from_latest=100 output_format="parquet" # important: use parquet for SQL ) files = result.get("files", []) # Get the returned file paths使用
get_sql_table_schema探索模式:# Check what columns are available in the file schema = get_sql_table_schema(files[0]) # Now you can see all columns, data types, and sample data使用
query_sql运行 SQL :# Option 1: Simple table reference (DuckDB will match the table name to file) sql_result = query_sql( query="SELECT block_number, timestamp, gas_used FROM blocks", files=files # Pass the files from step 1 ) # Option 2: Using read_parquet() with explicit file path sql_result = query_sql( query=f"SELECT block_number, timestamp, gas_used FROM read_parquet('{files[0]}')", files=files # Pass the files from step 1 )
或者,使用与query_blockchain_sql组合方法:
# Option 1: Simple table reference
result = query_blockchain_sql(
sql_query="SELECT * FROM blocks",
dataset="blocks",
blocks_from_latest=100
)
# Option 2: Using read_parquet()
result = query_blockchain_sql(
sql_query="SELECT * FROM read_parquet('/path/to/file.parquet')", # Path doesn't matter
dataset="blocks",
blocks_from_latest=100
)有关完整的工作示例,请参阅examples/sql_workflow_example.py 。
Related MCP server: EVM MCP Server
特征
完整 Cryo 数据集访问:通过 API 服务器查询任何 Cryo 数据集
MCP 集成:与 MCP 客户端无缝协作
灵活的查询选项:支持所有主要的Cryo过滤和输出选项
区块范围选项:查询特定区块、最新区块或相对范围
合约过滤:按合约地址过滤数据
最新区块访问:轻松访问最新的以太坊区块数据
多种输出格式:JSON、CSV 和 Parquet 支持
架构信息:获取详细的数据集架构和示例数据
SQL 查询:直接针对下载的区块链数据运行 SQL 查询
安装(可选)
如果您直接使用uvx运行该工具,则不需要这样做。
# install with UV (recommended)
uv tool install cryo-mcp要求
Python 3.8+
紫外线
Cryo的工作安装
访问以太坊 RPC 端点
DuckDB(用于 SQL 查询功能)
快速入门
与 Claude Code 一起使用
运行
claude mcp add以获得交互式提示。输入
uvx作为运行命令。输入
cryo-mcp --rpc-url <ETH_RPC_URL> [--data-dir <DATA_DIR>]作为参数或者,提供
ETH_RPC_URL和CRYO_DATA_DIR作为环境变量。
claude的新实例现在可以访问 cryo,因为它配置为命中您的 RPC 端点并将数据存储在指定的目录中。
可用工具
Cryo MCP 公开了以下 MCP 工具:
list_datasets()
返回所有可用 Cryo 数据集的列表。
例子:
client.list_datasets()query_dataset()
使用各种过滤选项查询 Cryo 数据集。
参数:
dataset(str):要查询的数据集的名称(例如,“blocks”、“transactions”、“logs”)blocks(str,可选):块范围规范(例如,'1000:1010')start_block(int,可选):起始块号(块的替代)end_block(int,可选):结束块号(块的替代)use_latest(bool, 可选): 如果为 True,则查询最新区块blocks_from_latest(int,可选):从最新到包含的块数contract(str,可选):要过滤的合约地址output_format(str,可选):输出格式('json','csv','parquet')include_columns(列表,可选):与默认值一起包含的列exclude_columns(列表,可选):从默认值中排除的列
例子:
# Get transactions from blocks 15M to 15.01M
client.query_dataset('transactions', blocks='15M:15.01M')
# Get logs for a specific contract from the latest 100 blocks
client.query_dataset('logs', blocks_from_latest=100, contract='0x1234...')
# Get just the latest block
client.query_dataset('blocks', use_latest=True)lookup_dataset()
获取有关特定数据集的详细信息,包括模式和示例数据。
参数:
name(str): 要查找的数据集的名称sample_start_block(int,可选):样本数据的起始块sample_end_block(int,可选):样本数据的结束块use_latest_sample(bool,可选):使用最新块作为样本sample_blocks_from_latest(int,可选):样本的最新块数
例子:
client.lookup_dataset('logs')get_latest_ethereum_block()
返回有关最新以太坊区块的信息。
例子:
client.get_latest_ethereum_block()SQL查询工具
Cryo MCP 包含几个用于针对区块链数据运行 SQL 查询的工具:
query_sql()
对下载的区块链数据运行 SQL 查询。
参数:
query(str):要执行的 SQL 查询files(列表,可选):要查询的 parquet 文件路径列表。如果为 None,则使用数据目录中的所有文件。include_schema(bool,可选):是否在结果中包含架构信息
例子:
# Run against all available files
client.query_sql("SELECT * FROM read_parquet('/path/to/blocks.parquet') LIMIT 10")
# Run against specific files
client.query_sql(
"SELECT * FROM read_parquet('/path/to/blocks.parquet') LIMIT 10",
files=['/path/to/blocks.parquet']
)query_blockchain_sql()
使用 SQL 查询区块链数据,自动下载任何所需数据。
参数:
sql_query(str): 要执行的 SQL 查询dataset(str,可选):要查询的数据集(例如,“区块”、“交易”)blocks(str,可选):块范围规范start_block(int,可选):起始块号end_block(int,可选):结束块号use_latest(bool, 可选): 如果为 True,则查询最新区块blocks_from_latest(int,可选):要包含的最新块之前块的数量contract(str,可选):要过滤的合约地址force_refresh(bool,可选):即使存在新数据也强制下载。include_schema(bool,可选):在结果中包含架构信息
例子:
# Automatically downloads blocks data if needed, then runs the SQL query
client.query_blockchain_sql(
sql_query="SELECT block_number, gas_used, timestamp FROM blocks ORDER BY gas_used DESC LIMIT 10",
dataset="blocks",
blocks_from_latest=100
)list_available_sql_tables()
列出所有可用 SQL 查询的表。
例子:
client.list_available_sql_tables()get_sql_table_schema()
获取特定 parquet 文件的架构。
参数:
file_path(str): parquet 文件的路径
例子:
client.get_sql_table_schema("/path/to/blocks.parquet")get_sql_examples()
获取不同区块链数据集的示例 SQL 查询。
例子:
client.get_sql_examples()配置选项
启动 Cryo MCP 服务器时,您可以使用以下命令行选项:
--rpc-url URL:以太坊 RPC URL(覆盖 ETH_RPC_URL 环境变量)--data-dir PATH:存储下载数据的目录(覆盖 CRYO_DATA_DIR 环境变量,默认为 ~/.cryo-mcp/data/)
环境变量
ETH_RPC_URL:未通过命令行指定时使用的默认以太坊 RPC URLCRYO_DATA_DIR:未通过命令行指定时存储下载数据的默认目录
高级用法
针对区块链数据的 SQL 查询
Cryo MCP 允许您对区块链数据运行强大的 SQL 查询,将 SQL 的灵活性与 Cryo 的数据提取功能相结合:
两步 SQL 查询流程
您可以将数据提取和查询分为两个独立的步骤:
# Step 1: Download data and get file paths
download_result = client.query_dataset(
dataset="transactions",
blocks_from_latest=1000,
output_format="parquet"
)
# Step 2: Use the file paths to run SQL queries
file_paths = download_result.get("files", [])
client.query_sql(
query=f"""
SELECT
to_address as contract_address,
COUNT(*) as tx_count,
SUM(gas_used) as total_gas,
AVG(gas_used) as avg_gas
FROM read_parquet('{file_paths[0]}')
WHERE to_address IS NOT NULL
GROUP BY to_address
ORDER BY total_gas DESC
LIMIT 20
""",
files=file_paths
)组合 SQL 查询流
为了方便起见,您还可以使用处理这两个步骤的组合函数:
# Get top gas-consuming contracts
client.query_blockchain_sql(
sql_query="""
SELECT
to_address as contract_address,
COUNT(*) as tx_count,
SUM(gas_used) as total_gas,
AVG(gas_used) as avg_gas
FROM read_parquet('/path/to/transactions.parquet')
WHERE to_address IS NOT NULL
GROUP BY to_address
ORDER BY total_gas DESC
LIMIT 20
""",
dataset="transactions",
blocks_from_latest=1000
)
# Find blocks with the most transactions
client.query_blockchain_sql(
sql_query="""
SELECT
block_number,
COUNT(*) as tx_count
FROM read_parquet('/path/to/transactions.parquet')
GROUP BY block_number
ORDER BY tx_count DESC
LIMIT 10
""",
dataset="transactions",
blocks="15M:16M"
)
# Analyze event logs by topic
client.query_blockchain_sql(
sql_query="""
SELECT
topic0,
COUNT(*) as event_count
FROM read_parquet('/path/to/logs.parquet')
GROUP BY topic0
ORDER BY event_count DESC
LIMIT 20
""",
dataset="logs",
blocks_from_latest=100
)注意:对于 SQL 查询,下载数据时请始终使用output_format="parquet"以确保 DuckDB 获得最佳性能。使用query_blockchain_sql时,应使用read_parquet()函数直接在 SQL 中引用文件路径。
使用块范围查询
Cryo MCP 支持 Cryo 的全部块规范语法:
# Using block numbers
client.query_dataset('transactions', blocks='15000000:15001000')
# Using K/M notation
client.query_dataset('logs', blocks='15M:15.01M')
# Using offsets from latest
client.query_dataset('blocks', blocks_from_latest=100)合同过滤
按合约地址过滤日志和其他数据:
# Get all logs for USDC contract
client.query_dataset('logs',
blocks='16M:16.1M',
contract='0xa0b86991c6218b36c1d19d4a2e9eb0ce3606eb48')列选择
仅包含您需要的列:
# Get just block numbers and timestamps
client.query_dataset('blocks',
blocks='16M:16.1M',
include_columns=['number', 'timestamp'])发展
项目结构
cryo-mcp/
├── cryo_mcp/ # Main package directory
│ ├── __init__.py # Package initialization
│ ├── server.py # Main MCP server implementation
│ ├── sql.py # SQL query functionality
├── tests/ # Test directory
│ ├── test_*.py # Test files
├── pyproject.toml # Project configuration
├── README.md # Project documentation运行测试
uv run pytest
执照
麻省理工学院
致谢
Available Tools
10 toolsget_latest_ethereum_blockB
Get information about the latest Ethereum block
Returns:
Information about the latest block including block number
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the return includes 'block number,' but doesn't specify other behavioral traits such as rate limits, error conditions, data freshness, or whether it's a read-only operation. This leaves significant gaps in understanding how the tool behaves beyond its basic function.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured, with two sentences that efficiently convey the tool's purpose and return value. There's no wasted language, and it's front-loaded with the main function. However, it could be slightly more polished by integrating the return information into the first sentence for better flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of blockchain data retrieval and the lack of annotations and output schema, the description is incomplete. It mentions returning 'information about the latest block including block number,' but doesn't detail other returned fields (e.g., timestamp, transactions), error handling, or performance considerations. This leaves the agent with insufficient context for effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, and the schema description coverage is 100%, so there are no parameters to document. The description doesn't need to add parameter semantics, and it appropriately doesn't mention any. A baseline of 4 is applied since no parameters exist, and the description doesn't introduce confusion about them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get information about the latest Ethereum block.' It specifies the verb ('get') and resource ('latest Ethereum block'), making the function unambiguous. However, it doesn't explicitly differentiate from sibling tools like 'get_transaction_by_hash' or 'query_blockchain_sql,' which prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention when this tool is appropriate (e.g., for real-time block data) or when to use other tools like 'query_blockchain_sql' for more complex queries. This lack of contextual direction leaves the agent without usage instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sql_examplesA
Get example SQL queries for different blockchain datasets with DuckDB
SQL WORKFLOW TIPS:
1. First download data: result = query_dataset('dataset_name', blocks='...', output_format='parquet')
2. Inspect schema: schema = get_sql_table_schema(result['files'][0])
3. Run SQL: query_sql("SELECT * FROM read_parquet('/path/to/file.parquet')", files=result['files'])
OR use the combined approach:
- query_blockchain_sql(sql_query="SELECT * FROM read_parquet('...')", dataset='blocks', blocks='...')
Returns:
Dictionary of example queries categorized by dataset type and workflow patterns
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly describes the return format ('Dictionary of example queries categorized by dataset type and workflow patterns'), which is helpful. However, it doesn't mention potential limitations like rate limits, authentication needs, or whether the examples are static or dynamically generated. The description doesn't contradict any annotations (none exist).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (purpose, workflow tips, combined approach, returns). It's appropriately sized for the complexity, though the workflow tips section is somewhat lengthy for a tool that just returns examples. Every sentence earns its place by providing practical guidance, but it could be slightly more concise by integrating tips more tightly with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (parameterless, no output schema, no annotations), the description is quite complete. It explains the purpose, provides extensive usage guidance, and describes the return format. The only minor gap is lack of explicit behavioral constraints (e.g., whether examples are curated/static), but overall it's well-rounded for this type of helper tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, and schema description coverage is 100% (empty schema). The description appropriately doesn't discuss parameters, which is correct for a parameterless tool. It focuses instead on output semantics and usage context, adding value beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get example SQL queries for different blockchain datasets with DuckDB'. It specifies the exact resource (example SQL queries) and distinguishes from siblings like query_sql (executes SQL) or list_datasets (lists datasets). The verb 'Get' is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance through 'SQL WORKFLOW TIPS' and 'OR use the combined approach', detailing when to use this tool (for learning/example queries) versus alternatives like query_sql or query_blockchain_sql (for actual execution). It names specific sibling tools (query_dataset, get_sql_table_schema, query_sql, query_blockchain_sql) and explains their roles in workflows.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sql_table_schemaA
Get the schema and sample data for a specific parquet file
WORKFLOW NOTE: Use this function to explore the structure of parquet files
before writing SQL queries against them. This will show you:
1. All available columns and their data types
2. Sample data from the file
3. Total row count
Usage example:
1. Get list of files: files = list_available_sql_tables()
2. For a specific file: schema = get_sql_table_schema(files[0]['path'])
3. Use columns in your SQL: query_sql("SELECT column1, column2 FROM read_parquet('/path/to/file.parquet')")
Args:
file_path: Path to the parquet file (from list_available_sql_tables or query_dataset)
Returns:
Table schema information including columns, data types, and sample data
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden and does well by disclosing key behaviors: it returns schema, sample data, and row count; it's for exploration (not modification); and it requires a file path from other tools. It doesn't mention performance characteristics or error handling, but covers the core functionality adequately.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (purpose, workflow note, usage example, args, returns) and every sentence adds value. It's slightly longer than minimal but justified by the comprehensive guidance. The front-loaded purpose statement is excellent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no annotations and no output schema, the description provides excellent context: clear purpose, usage guidelines, parameter explanation, and return value description. It doesn't detail the exact output structure, but given the tool's exploratory nature and the sibling context, this is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate fully. It clearly explains the single parameter's purpose ('Path to the parquet file'), source ('from list_available_sql_tables or query_dataset'), and provides usage examples showing how to obtain and use it, adding significant value beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Get the schema and sample data'), resource ('for a specific parquet file'), and distinguishes it from siblings like list_available_sql_tables (which lists files) and query_sql (which executes queries). The WORKFLOW NOTE further clarifies its exploratory purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('to explore the structure of parquet files before writing SQL queries') and provides a detailed workflow example showing how it integrates with sibling tools (list_available_sql_tables and query_sql). It clearly positions this as a preparatory step for querying.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_transaction_by_hashC
Get detailed information about a transaction by its hash
Args:
tx_hash: The transaction hash to look up
Returns:
Detailed information about the transaction
| Name | Required | Description | Default |
|---|---|---|---|
| tx_hash | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure but offers minimal information. It states what the tool does but doesn't cover important aspects like error handling (what happens with invalid hashes), performance characteristics, rate limits, authentication requirements, or whether this is a read-only operation. The description is functionally correct but lacks operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured with clear sections for purpose, arguments, and returns. Each sentence serves a distinct purpose without redundancy. The formatting with headers makes it easy to parse, though the 'Returns' section could be more specific given there's no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and 0% schema description coverage, the description is insufficiently complete. It covers the basic operation but lacks critical information about what 'detailed information' includes, error conditions, performance expectations, and how this tool relates to the available SQL query alternatives on the server.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explicitly documents the single parameter 'tx_hash' and its purpose ('The transaction hash to look up'), which adds value beyond the schema's 0% description coverage. However, it doesn't provide format details (e.g., hex string, length requirements) or validation rules that would be helpful for proper usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Get') and resource ('detailed information about a transaction'), making it immediately understandable. However, it doesn't differentiate this tool from potential siblings like 'query_blockchain_sql' or 'query_dataset' that might also retrieve transaction data through different mechanisms.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. With siblings like 'query_blockchain_sql' and 'query_dataset' available, there's no indication whether this is the preferred method for transaction lookups, if it's faster for single transactions, or when SQL queries would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_available_sql_tablesA
List all available parquet files that can be queried with SQL
USAGE NOTES:
- This function lists parquet files that have already been downloaded
- Each file can be queried using read_parquet('/path/to/file.parquet') in your SQL
- For each file, this returns the file path, dataset type, and other metadata
- Use these file paths in your SQL queries with query_sql()
Returns:
List of available files and their metadata
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden and does well by explaining what gets returned (file path, dataset type, metadata), how to use the output (with read_parquet() and query_sql()), and the prerequisite that files must be 'already downloaded'. It doesn't mention performance characteristics or error conditions, but provides substantial behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections, front-loads the core purpose, and every sentence adds value. The USAGE NOTES bullet points efficiently convey critical information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no annotations or output schema, the description provides excellent context about what the tool does, how to use its output, and relationships to other tools. It could mention error conditions or performance, but covers the essential usage context thoroughly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema coverage, so the baseline is 4. The description appropriately doesn't discuss parameters since none exist, focusing instead on what the tool does and how to use its output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'list' and the resource 'parquet files that can be queried with SQL', specifying that these are files that have already been downloaded. It distinguishes from siblings like list_datasets by focusing specifically on SQL-queryable parquet files rather than datasets in general.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The USAGE NOTES section explicitly states when to use this tool (to get file paths for SQL queries) and how to use the output with query_sql(). It also distinguishes from alternatives by noting these are 'already downloaded' files, implying list_datasets might show available datasets that aren't yet downloaded.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_datasetsB
Return a list of all available cryo datasets
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It states the tool returns a list but doesn't specify format, pagination, rate limits, authentication needs, or whether it's read-only. For a tool with zero annotation coverage, this leaves significant gaps in understanding its behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states what the tool does. It's front-loaded with the core action and resource, with zero wasted words or redundant information, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, no output schema, no annotations), the description is adequate as a basic read operation. However, it lacks details about return format, data scope, or how it fits with sibling tools, which would help the agent use it more effectively in context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage, so the schema fully documents the lack of inputs. The description doesn't need to add parameter information, and it appropriately doesn't mention any. Baseline for 0 parameters is 4, as the description focuses on the tool's purpose without unnecessary parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Return a list') and resource ('all available cryo datasets'), making the purpose immediately understandable. However, it doesn't differentiate from sibling tools like 'lookup_dataset' or 'list_available_sql_tables', which prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'lookup_dataset' or 'query_dataset'. It lacks any context about prerequisites, timing, or exclusions, leaving the agent to infer usage from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lookup_datasetA
Look up a specific dataset and return detailed information about it. IMPORTANT: Always use this
function before querying a new dataset to understand its required parameters and schema.
The returned information includes:
1. Required parameters for the dataset (IMPORTANT for datasets like 'balances' that need an address)
2. Schema details showing available columns and data types
3. Example queries for the dataset
When the dataset requires specific parameters like 'address' (for 'balances'),
ALWAYS use the 'contract' parameter in query_dataset() to pass these values.
Example:
For 'balances' dataset, lookup_dataset('balances') will show it requires an 'address' parameter.
You should then query it using:
query_dataset('balances', blocks='1000:1010', contract='0x1234...')
Args:
name: The name of the dataset to look up
sample_start_block: Optional start block for sample data (integer)
sample_end_block: Optional end block for sample data (integer)
use_latest_sample: If True, use the latest block for sample data
sample_blocks_from_latest: Number of blocks before the latest to include in sample
Returns:
Detailed information about the dataset including schema and available fields
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| sample_blocks_from_latest | No | ||
| sample_end_block | No | ||
| sample_start_block | No | ||
| use_latest_sample | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden and does well by explaining the tool's behavior: it's a read-only lookup (implied by 'look up' and 'return information'), it provides schema details and required parameters, and it includes sample data generation capabilities. However, it doesn't mention rate limits, authentication needs, or error conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded with the core purpose and important usage guideline. Every sentence adds value, though the example section is somewhat lengthy. The structure flows logically from purpose to usage to parameters to returns.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters, 0% schema coverage, no annotations, and no output schema, the description does an excellent job of explaining what the tool does, when to use it, and what parameters mean. The main gap is lack of information about return format details, though it describes what information will be included.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, but the description fully compensates by explaining all 5 parameters in detail. It clarifies that 'name' is the dataset identifier, and the other 4 parameters control sample data generation (with specific examples like 'sample_blocks_from_latest'). This adds substantial meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('look up', 'return detailed information') and resource ('dataset'), distinguishing it from siblings like list_datasets (which lists datasets) or query_dataset (which queries data). It explicitly explains this is for understanding dataset parameters and schema before querying.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool ('Always use this function before querying a new dataset') and when to use alternatives (query_dataset for actual queries). It includes a concrete example showing the workflow between lookup_dataset and query_dataset, making the usage context clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_blockchain_sqlA
Download blockchain data and run SQL query in a single step
CONVENIENCE FUNCTION: This combines query_dataset and query_sql into one call.
You can write SQL queries using either approach:
1. Simple table references: "SELECT * FROM blocks LIMIT 10"
2. Explicit read_parquet: "SELECT * FROM read_parquet('/path/to/file.parquet') LIMIT 10"
DATASET-SPECIFIC PARAMETERS:
For datasets that require specific address parameters (like 'balances', 'erc20_transfers', etc.),
ALWAYS use the 'contract' parameter to pass ANY Ethereum address. For example:
- For 'balances' dataset: Use contract parameter for the address you want balances for
query_blockchain_sql(
sql_query="SELECT * FROM balances",
dataset="balances",
blocks='1000:1010',
contract='0x123...' # Address you want balances for
)
Examples:
```
# Using simple table name
query_blockchain_sql(
sql_query="SELECT * FROM blocks LIMIT 10",
dataset="blocks",
blocks_from_latest=100
)
# Using read_parquet() (the path will be automatically replaced)
query_blockchain_sql(
sql_query="SELECT * FROM read_parquet('/any/path.parquet') LIMIT 10",
dataset="blocks",
blocks_from_latest=100
)
```
ALTERNATIVE WORKFLOW (more control):
If you need more control, you can separate the steps:
1. Download data: result = query_dataset('blocks', blocks_from_latest=100, output_format='parquet')
2. Inspect schema: schema = get_sql_table_schema(result['files'][0])
3. Run SQL query: query_sql("SELECT * FROM blocks", files=result['files'])
Args:
sql_query: SQL query to execute - using table names or read_parquet()
dataset: The specific dataset to query (e.g., 'transactions', 'logs', 'balances')
If None, will be extracted from the SQL query
blocks: Block range specification as a string (e.g., '1000:1010')
start_block: Start block number (alternative to blocks)
end_block: End block number (alternative to blocks)
use_latest: If True, query the latest block
blocks_from_latest: Number of blocks before the latest to include
contract: Contract address to filter by - IMPORTANT: Use this parameter for ALL address-based filtering
regardless of the parameter name in the native cryo command (address, contract, etc.)
force_refresh: Force download of new data even if it exists
include_schema: Include schema information in the result
Returns:
SQL query results and metadata
| Name | Required | Description | Default |
|---|---|---|---|
| blocks | No | ||
| blocks_from_latest | No | ||
| contract | No | ||
| dataset | No | ||
| end_block | No | ||
| force_refresh | No | ||
| include_schema | No | ||
| sql_query | Yes | ||
| start_block | No | ||
| use_latest | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does well by explaining the two SQL query approaches, dataset-specific parameter requirements, and the return format ('SQL query results and metadata'). However, it doesn't mention performance characteristics, rate limits, or error conditions that would be helpful for a complex tool with 10 parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (purpose, convenience note, SQL approaches, dataset-specific guidance, examples, alternative workflow, and parameter details). While comprehensive, some sections could be more concise - the examples are quite detailed, and the description is lengthy overall for a tool description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 10 parameters, 0% schema coverage, no annotations, and no output schema, the description does an excellent job of providing context. It explains the tool's relationship to siblings, provides usage examples, clarifies parameter semantics, and describes return values. The main gap is lack of information about performance, limits, or error handling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage and 10 parameters, the description compensates excellently. It provides detailed explanations for key parameters like 'contract' with specific examples, clarifies parameter relationships (blocks vs start_block/end_block), and explains default behaviors. The 'Args' section adds crucial semantic context beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Download blockchain data and run SQL query in a single step' and explicitly distinguishes it from siblings by naming 'query_dataset' and 'query_sql' as separate tools that this one combines. It specifies the exact functionality and differentiates from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool vs alternatives: it states this is a 'CONVENIENCE FUNCTION' that combines two other tools, and provides an 'ALTERNATIVE WORKFLOW' section detailing when to use the separate steps for 'more control'. It clearly delineates the trade-offs between convenience and control.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_datasetA
Download blockchain data and return the file paths where the data is stored.
IMPORTANT WORKFLOW NOTE: When running SQL queries, use this function first to download
data, then use the returned file paths with query_sql() to execute SQL on those files.
Example workflow for SQL:
1. First download data: result = query_dataset('transactions', blocks='1000:1010', output_format='parquet')
2. Get file paths: files = result.get('files', [])
3. Run SQL query: query_sql("SELECT * FROM read_parquet('/path/to/file.parquet')", files=files)
DATASET-SPECIFIC PARAMETERS:
For datasets that require specific address parameters (like 'balances', 'erc20_transfers', etc.),
ALWAYS use the 'contract' parameter to pass ANY Ethereum address. For example:
- For 'balances' dataset: Use contract parameter for the address you want balances for
query_dataset('balances', blocks='1000:1010', contract='0x123...')
- For 'logs' or 'erc20_transfers': Use contract parameter for contract address
query_dataset('logs', blocks='1000:1010', contract='0x123...')
To check what parameters a dataset requires, always use lookup_dataset() first:
lookup_dataset('balances') # Will show required parameters
Args:
dataset: The name of the dataset to query (e.g., 'logs', 'transactions', 'balances')
blocks: Block range specification as a string (e.g., '1000:1010')
start_block: Start block number as integer (alternative to blocks)
end_block: End block number as integer (alternative to blocks)
use_latest: If True, query the latest block
blocks_from_latest: Number of blocks before the latest to include (e.g., 10 = latest-10 to latest)
contract: Contract address to filter by - IMPORTANT: Use this parameter for ALL address-based filtering
regardless of the parameter name in the native cryo command (address, contract, etc.)
output_format: Output format (json, csv, parquet) - use 'parquet' for SQL queries
include_columns: Columns to include alongside the defaults
exclude_columns: Columns to exclude from the defaults
Returns:
Dictionary containing file paths where the downloaded data is stored
| Name | Required | Description | Default |
|---|---|---|---|
| blocks | No | ||
| blocks_from_latest | No | ||
| contract | No | ||
| dataset | Yes | ||
| end_block | No | ||
| exclude_columns | No | ||
| include_columns | No | ||
| output_format | No | json | |
| start_block | No | ||
| use_latest | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes the tool's behavior: it downloads data to files, returns a dictionary of file paths, and integrates with query_sql. It mentions dataset-specific requirements and workflow dependencies, though it doesn't cover potential errors, rate limits, or file storage details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (workflow note, dataset-specific parameters, Args, Returns) and uses bullet points for readability. It's appropriately sized for a complex tool but could be slightly more concise by integrating the example workflow more tightly with the parameter explanations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 10 parameters, 0% schema coverage, no annotations, and no output schema, the description provides comprehensive context. It explains the tool's role in a larger workflow, details all parameters, and describes the return value. The main gap is lack of error handling or performance considerations, but it's largely complete given the constraints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Given 0% schema description coverage, the description compensates fully by explaining all 10 parameters in detail. It clarifies dataset-specific usage (e.g., contract parameter for address filtering), provides examples for blocks and output_format, and explains parameter relationships (e.g., blocks vs. start_block/end_block). This adds significant meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Download blockchain data and return the file paths where the data is stored.' It specifies the verb ('download'), resource ('blockchain data'), and output ('file paths'), distinguishing it from siblings like query_sql or list_datasets that don't download data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool vs. alternatives: 'When running SQL queries, use this function first to download data, then use the returned file paths with query_sql() to execute SQL on those files.' It also advises to use lookup_dataset() first to check dataset parameters, offering clear workflow instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_sqlA
Run a SQL query against downloaded blockchain data files
IMPORTANT WORKFLOW: This function should be used after calling query_dataset
to download data. Use the file paths returned by query_dataset as input to this function.
Workflow steps:
1. Download data: result = query_dataset('transactions', blocks='1000:1010', output_format='parquet')
2. Get file paths: files = result.get('files', [])
3. Execute SQL using either:
- Direct table references: query_sql("SELECT * FROM transactions", files=files)
- Or read_parquet(): query_sql("SELECT * FROM read_parquet('/path/to/file.parquet')", files=files)
To see the schema of a file, use get_sql_table_schema(file_path) before writing your query.
DuckDB supports both approaches:
1. Direct table references (simpler): "SELECT * FROM blocks"
2. read_parquet function (explicit): "SELECT * FROM read_parquet('/path/to/file.parquet')"
Args:
query: SQL query to execute - can use simple table names or read_parquet()
files: List of parquet file paths to query (typically from query_dataset results)
include_schema: Whether to include schema information in the result
Returns:
Query results and metadata
| Name | Required | Description | Default |
|---|---|---|---|
| files | No | ||
| include_schema | No | ||
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively explains the tool's behavior: it executes SQL queries against parquet files, supports two query approaches (direct table references or read_parquet()), and mentions DuckDB as the underlying engine. It also notes that results include query results and metadata. The main gap is lack of information about error handling, performance characteristics, or limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (purpose, workflow, examples, parameter explanations, return value). While comprehensive, it could be more concise - some information is repeated (e.g., both workflow steps and DuckDB approaches mention the two query methods). Every sentence earns its place, but some tightening is possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (SQL execution with file dependencies), no annotations, and no output schema, the description provides substantial context. It explains the workflow, parameter usage, and return values. The main gap is lack of output format details - while it mentions 'Query results and metadata', it doesn't specify the structure. For a SQL execution tool, more detail on result format would be helpful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by explaining all three parameters. It clarifies that 'query' is the SQL to execute with syntax examples, 'files' are typically from query_dataset results, and 'include_schema' controls whether schema information is included in results. The description adds substantial value beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as 'Run a SQL query against downloaded blockchain data files', specifying both the action (run SQL query) and the target resource (downloaded blockchain data files). It distinguishes this from sibling tools like query_dataset (which downloads data) and get_sql_table_schema (which shows schema).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit workflow guidance: 'This function should be used after calling query_dataset to download data.' It names the specific prerequisite tool (query_dataset) and explains the sequence of operations. The 'IMPORTANT WORKFLOW' section clearly defines when to use this tool versus alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v1.0.0- First observed
get_latest_ethereum_block - First observed
get_sql_examples - First observed
get_sql_table_schema - First observed
get_transaction_by_hash - First observed
list_available_sql_tables - First observed
list_datasets - First observed
lookup_dataset - First observed
query_blockchain_sql - First observed
query_dataset - First observed
query_sql
TDQS
Scored across 10 tools
Most tools have distinct purposes, but there is some overlap between query_blockchain_sql and the combination of query_dataset + query_sql. The descriptions clarify that query_blockchain_sql is a convenience function combining the two, which helps reduce confusion, but agents might still need to decide between the separate or combined approach.
All tool names follow a consistent verb_noun pattern with snake_case, such as get_latest_ethereum_block, list_available_sql_tables, and query_dataset. This consistency makes the tool set predictable and easy to navigate for agents.
With 10 tools, the server is well-scoped for its purpose of blockchain data querying and SQL analysis. Each tool serves a clear role, from data retrieval (e.g., get_transaction_by_hash) to dataset exploration (e.g., lookup_dataset) and SQL execution (e.g., query_sql), without feeling bloated or sparse.
The tool set provides comprehensive coverage for blockchain data workflows, including data fetching (get_latest_ethereum_block, get_transaction_by_hash), dataset discovery (list_datasets, lookup_dataset), data download (query_dataset), SQL querying (query_sql, query_blockchain_sql), and schema inspection (get_sql_table_schema, list_available_sql_tables). There are no obvious gaps, and the tools support full CRUD-like operations for the domain.
Maintenance
Related MCP Connectors
MCP server for OpenAI API (chat completions, image generation, embeddings) via AceDataCloud
Read-only MCP server for The Quiet Protocol's engines, benchmarks, proof, and business data.
MCP server connecting AI agents to non-custodial staking data across 130+ networks.
Token-free MCP server for structured RevoGrid Core, Pro, and Enterprise knowledge retrieval.
Related MCP Servers
- AlicenseDqualityDmaintenanceA Model Context Protocol server that gives LLMs the ability to interact with Ethereum networks, manage wallets, query blockchain data, and execute smart contract operations through a standardized interface.5411 npm14MIT
- AlicenseBqualityDmaintenanceA Model Context Protocol server that enables AI agents to interact with 30+ Ethereum-compatible blockchain networks, providing services like token transfers, contract interactions, and ENS resolution through a unified interface.28128 npm379MIT
- AlicenseBqualityDmaintenanceAn MCP server that provides access to Etherscan blockchain data APIs, allowing users to query Ethereum blockchain information through natural language.636 npm18MIT
- AlicenseNot gradedqualityCmaintenanceAn MCP server that bridges AI models with Ethereum blockchains via all JSON-RPC calls, enabling natural language queries for block numbers, balances, transactions, and smart contract data.8 npm21MIT