Skip to main content
Glama
Crawlora-org

Crawlora MCP

Official

datasets_github_users_facets

Aggregate GitHub user counts by facet (e.g., country, company, influence tier) with filters for location, activity, and contact presence.

Instructions

Facet the GitHub users dataset. Returns terms aggregation counts for the GitHub users dataset. Facet enum: influence_tier, type, country, country_code, state, city, domains, company, reachable, has_email, has_twitter, has_blog, active_90d, hireable, is_org, is_bot, is_suspected_automation. influence_tier enum: nano, micro, mid, macro, mega. Suspected-automation records are excluded by default unless is_suspected_automation is set.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
qNoFull-text query over login, name, company, bio and location, max 256 characters
latNoLatitude for radius filtering
lonNoLongitude for radius filtering
cityNoExact geocoded city filter, max 128 characters
sortNoSort enum: relevance, rank_score_desc, followers_desc, account_age_desc, account_age_asc, distance_asc
facetYesFacet enum: influence_tier, type, country, country_code, state, city, domains, company, reachable, has_email, has_twitter, has_blog, active_90d, hireable, is_org, is_bot, is_suspected_automation
loginNoExact login filter, max 128 characters
stateNoExact geocoded state filter, max 128 characters
domainNoInterest-domain tag filter, max 128 characters
is_botNoBot filter
is_orgNoOrganization filter
companyNoExact normalized-company filter, max 128 characters
countryNoExact geocoded country filter, max 128 characters
has_blogNoFilter by public blog/website presence
hireableNoFilter by the GitHub available-for-hire flag
radius_mNoRadius in meters, 1 through 50000; requires lat and lon when supplied
has_emailNoFilter by public email presence
min_reposNoMinimum public repository count
reachableNoFilter by any public contact channel
active_90dNoFilter by activity within the last 90 days
has_twitterNoFilter by public Twitter/X handle presence
country_codeNoExact ISO country-code filter, max 128 characters
max_followersNoMaximum follower count
min_followersNoMinimum follower count
influence_tierNoFollower-tier enum: nano, micro, mid, macro, mega
min_rank_scoreNoMinimum composite rank score
max_account_age_yearsNoMaximum account age in years
min_account_age_yearsNoMinimum account age in years
is_suspected_automationNoSuspected automation filter; omitted these are hidden by default

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changedv1.17.5
    • addedInput schema / properties / facet / enum
      Added value: +[
      +  "influence_tier",
      +  "type",
      +  "country",
      +  "country_code",
      +  "state",
      +  "city",
      +  "domains",
      +  "company",
      +  "reachable",
      +  "has_email",
      +  "has_twitter",
      +  "has_blog",
      +  "active_90d",
      +  "hireable",
      +  "is_org",
      +  "is_bot",
      +  "is_suspected_automation"
      +]
    • addedInput schema / properties / influence_tier / enum
      Added value: +[
      +  "nano",
      +  "micro",
      +  "mid",
      +  "macro",
      +  "mega"
      +]
    • addedInput schema / properties / sort / enum
      Added value: +[
      +  "relevance",
      +  "rank_score_desc",
      +  "followers_desc",
      +  "account_age_desc",
      +  "account_age_asc",
      +  "distance_asc"
      +]
  2. Changed2 schema fields changedv1.5.0
    • changedInput schema / properties / facet / description
      Previous value: -"Facet enum: influence_tier, type, country, country_code, state, city, domains, company, reachable, has_email, has_twitter, has_blog, active_90d, is_org, is_bot"New value: +"Facet enum: influence_tier, type, country, country_code, state, city, domains, company, reachable, has_email, has_twitter, has_blog, active_90d, hireable, is_org, is_bot, is_suspected_automation"
    • addedInput schema / properties / is_suspected_automation
      Added value: +{
      +  "description": "Suspected automation filter; omitted these are hidden by default",
      +  "type": "boolean"
      +}
  3. Addedv1.2.0

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full responsibility for behavioral disclosure. It reveals that suspected-automation records are excluded by default unless explicitly filtered, which is useful. However, it omits other key behaviors such as return format (e.g., list of buckets with doc counts), whether results are limited, or how multiple filters interact. The single nuance is insufficient for a 29-parameter tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded. It states the purpose in one sentence, then lists enums and a key default. No fluff. The enum list is necessary and the structure is efficient, though it could benefit from a brief note on output structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 29 parameters and no output schema, the description is incomplete. It does not mention the shape of the response (e.g., aggregation buckets), whether multiple facets can be requested, or how filters like q, country, or follower counts affect the aggregation. Crucial context for an agent to correctly invoke and interpret results is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and every parameter already has a description. The description adds only a redundant note about the is_suspected_automation default (already in the schema) and repeats the facet and influence_tier enums. It provides no additional meaning beyond what the schema already offers, so it fails to add value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool facets the GitHub users dataset and returns terms aggregation counts. This is a specific verb and resource, and the purpose is unambiguous. However, it does not explicitly contrast with sibling search/item tools, though the distinction is implied by 'returns counts'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like datasets_github_users_search or other facet tools. The only usage hint is the default exclusion of suspected-automation records, but no context on when to choose faceting over search or other operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools