Skip to main content
Glama

hsh-finetune-dataset

Made-to-order, answer-verified datasets for LLM fine-tuning. Describe the task (e.g. 'step-by-step math reasoning', 'SQL generation', 'instruction-following for support replies') and we deliver a clean, HuggingFace-ready dataset in Alpaca schema (instruction/input/output), deduplicated, train/val/test split, with every checkable answer verified in code. Drop the repo straight into Gradients (SN56), TRL, Axolotl, or Unsloth. Verified sample live: huggingface.co/datasets/HSH-Intelligence/verified-math-reasoning-3k. Tier S: 1-2K rows ($75). Tier M: 2-5K rows ($150). Tier L: 5-10K rows ($300). Custom/larger scoped on request.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
domainNoSubject domain (e.g. 'math', 'SQL', 'customer support', 'legal Q&A').
row_countYesNumber of training rows needed (1000-10000 standard; larger scoped on request).
schema_hintNoPreferred schema. Default: Alpaca instruction/input/output (Gradients-ready).
verificationNoHow answers are checked. Programmatic (code-verified ground truth) where the task allows.
target_platformNoWhere you'll train — tunes the delivered format.
task_descriptionYesPlain English: what the model should learn to do (the instruction-following task).

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description covers key behavioral traits: delivered as a cleaned, HuggingFace-ready dataset in Alpaca schema, deduplicated, with train/val/test split and programmatic verification. It also mentions pricing tiers and a sample. However, it omits details on processing time, synchronous/asynchronous behavior, or limitations (e.g., maximum row count beyond stated tiers).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, somewhat lengthy paragraph (~120 words). It front-loads the core purpose (first sentence) but includes extra details like pricing tiers and a sample link that, while useful, could be condensed. It is not overly verbose but could be more structured (e.g., bullet points) to improve scanability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (custom dataset creation with verification) and lack of output schema, the description provides sufficient context: what the dataset contains, the schema, verification method, splitting, and pricing. It also gives a concrete example. However, it could be enhanced by specifying turnaround time or limitations (e.g., data size caps).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All 6 parameters are documented in the input schema (100% coverage), so baseline is 3. The description adds overarching context (e.g., default schema, verification approach, pricing) but does not provide additional per-parameter details beyond what the schema already gives. Thus, the description adds some value but not enough to raise the score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool creates 'made-to-order, answer-verified datasets for LLM fine-tuning', specifying the output format (Alpaca schema), verification method, and sample availability. This distinguishes it from siblings like hsh-custom-dataset by emphasizing verification and tailored dataset construction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool (for fine-tuning tasks needing verified datasets) and includes examples of applicable tasks. However, it does not explicitly state when not to use it or mention alternative tools, which would be helpful for differentiation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

B3.3/5.0
Disambiguation4/5

Most tools have distinct purposes, but some closely related tools (e.g., hsh-b2b-*, hsh-esg-* variants) could cause confusion. Descriptions help differentiate, but an agent might still misselect similar products.

Naming Consistency3/5

Naming convention is mixed: some tools use hyphens (hsh-b2b-contact), others use underscores (hsh_broker_data_request). While mostly readable, the inconsistency could be confusing for agents expecting a uniform pattern.

Tool Count3/5

32 tools is on the high side for a single server, but given its purpose as a data marketplace, the large number reflects a wide catalog. However, it may be overwhelming for agents to navigate.

Completeness3/5

Covers many data domains but has obvious gaps (e.g., weather, social media). The inclusion of custom data request tools (hsh_describe_data_need, hsh_broker_data_request) mitigates these gaps, allowing agents to request missing data.