Skip to main content
Glama

Crawlora MCP

datasets_github_users_facets

Read-only

Returns terms aggregation counts over the GitHub users dataset, honoring the same filters as search. Use alongside the related search tool to inspect filter counts under the same query filters.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
qNoOptional full-text query over login, name, company, bio and location, max 256 characters.
latNoOptional latitude for radius filtering or distance sort, from -90 through 90; supply together with lon.
lonNoOptional longitude for radius filtering or distance sort, from -180 through 180; supply together with lat.
cityNoOptional exact geocoded city filter, max 128 characters.
pageNoResult page number, 1-based, default 1; page times page_size must not exceed 10000.
sortNoOptional sort order. Allowed values: relevance, rank_score_desc, followers_desc, account_age_desc, account_age_asc, distance_asc. Defaults to relevance with q, otherwise rank_score_desc.
facetYesRequired facet to aggregate. Allowed values: influence_tier, type, country, country_code, state, city, domains, company, reachable, has_email, has_twitter, has_blog, active_90d, hireable, is_org, is_bot, is_suspected_automation.
loginNoOptional exact login filter, max 128 characters.
stateNoOptional exact geocoded state/region filter, max 128 characters.
domainNoOptional interest-domain tag filter (e.g. ml-ai, web, devops), max 128 characters.
is_botNoOptional bot filter (the crawl skips bots, so this is normally false).
is_orgNoOptional organization filter (the crawl indexes individuals, so this is normally false).
companyNoOptional exact normalized-company filter (lowercased, @ stripped), max 128 characters.
countryNoOptional exact geocoded country filter, max 128 characters.
has_blogNoOptional filter for a public blog/website URL.
hireableNoOptional filter for the GitHub 'available for hire' flag.
radius_mNoOptional radius in meters, from 1 through 50000; requires lat and lon when supplied.
has_emailNoOptional filter for a public email on the profile.
min_reposNoOptional minimum public repository count, 0 or greater.
page_sizeNoPage size, default 20, max 100; page times page_size must not exceed 10000.
reachableNoOptional filter for any public contact channel (email, Twitter, blog or social account).
active_90dNoOptional filter for activity within the last 90 days.
has_twitterNoOptional filter for a public Twitter/X handle.
country_codeNoOptional exact ISO country-code filter, max 128 characters.
max_followersNoOptional maximum follower count, 0 or greater.
min_followersNoOptional minimum follower count, 0 or greater.
influence_tierNoOptional follower-based tier filter. Allowed values: nano, micro, mid, macro, mega.
min_rank_scoreNoOptional minimum composite rank score, 0 or greater.
max_account_age_yearsNoOptional maximum account age in years, 0 or greater.
min_account_age_yearsNoOptional minimum account age in years, 0 or greater.
is_suspected_automationNoOptional filter for suspected automation (commit-farm / mass-repo bots). Omitted by default these records are hidden; pass true to isolate them or false to force-include the same clean view.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
dataYesThe tool result payload (shape varies per tool; see each tool's docs resource).

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds the useful constraint that it honors the same filters as search, but says nothing about aggregation semantics, whether page/page_size affect facet output, or result limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero padding, and the core purpose is front-loaded before the usage note. Every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the schema fully documents the large filter set. The description covers purpose and usage adequately, though for a 31-parameter tool it could say a bit more about how the facet aggregation relates to the pagination parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 31 parameters including the required 'facet' enum list are documented in the schema itself. The description adds no parameter-level meaning beyond confirming the filter set matches search, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Returns terms aggregation counts over the GitHub users dataset") and its scope ("honoring the same filters as search"), which lets an agent distinguish it from datasets_github_users_search. It does not name the search sibling explicitly, only referring to "the related search tool", so it falls just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly frames the use case: run it alongside the search tool to inspect filter counts under the same query filters. It gives positive guidance but no explicit when-not or exclusion (e.g. that it returns counts rather than records), so it is not fully prescriptive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources