Skip to main content
Glama
Crawlora-org

Crawlora MCP

Official

datasets_github_users_search

Search enriched GitHub user profiles using filters for location, influence, activity, and contact details. Retrieve targeted users for outreach or analysis.

Instructions

Search the GitHub users dataset. Searches enriched public GitHub user profiles stored in a search index. influence_tier enum: nano, micro, mid, macro, mega. Sort enum: relevance, rank_score_desc, followers_desc, account_age_desc, account_age_asc, distance_asc.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
qNoFull-text query over login, name, company, bio and location, max 256 characters
latNoLatitude for radius filtering or distance sort
lonNoLongitude for radius filtering or distance sort
cityNoExact geocoded city filter, max 128 characters
pageNoPage number, defaults to 1
sortNoSort enum: relevance, rank_score_desc, followers_desc, account_age_desc, account_age_asc, distance_asc
loginNoExact login filter, max 128 characters
stateNoExact geocoded state filter, max 128 characters
domainNoInterest-domain tag filter (e.g. ml-ai, web, devops), max 128 characters
is_botNoBot filter (normally false; the crawl skips bots)
is_orgNoOrganization filter (normally false; the crawl indexes individuals)
companyNoExact normalized-company filter, max 128 characters
countryNoExact geocoded country filter, max 128 characters
has_blogNoFilter by public blog/website presence
hireableNoFilter by the GitHub available-for-hire flag
radius_mNoRadius in meters, 1 through 50000; requires lat and lon when supplied
has_emailNoFilter by public email presence
min_reposNoMinimum public repository count
page_sizeNoPage size, defaults to 20 and maxes at 100; page * page_size must be <= 10000
reachableNoFilter by any public contact channel
active_90dNoFilter by activity within the last 90 days
has_twitterNoFilter by public Twitter/X handle presence
country_codeNoExact ISO country-code filter, max 128 characters
max_followersNoMaximum follower count
min_followersNoMinimum follower count
influence_tierNoFollower-tier enum: nano, micro, mid, macro, mega
min_rank_scoreNoMinimum composite rank score
max_account_age_yearsNoMaximum account age in years
min_account_age_yearsNoMinimum account age in years
is_suspected_automationNoSuspected automation (commit-farm/mass-repo bots); omitted these are hidden by default, pass true to isolate them

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv1.17.5
    • addedInput schema / properties / influence_tier / enum
      Added value: +[
      +  "nano",
      +  "micro",
      +  "mid",
      +  "macro",
      +  "mega"
      +]
    • addedInput schema / properties / sort / enum
      Added value: +[
      +  "relevance",
      +  "rank_score_desc",
      +  "followers_desc",
      +  "account_age_desc",
      +  "account_age_asc",
      +  "distance_asc"
      +]
  2. Changed1 schema field changedv1.5.0
    • addedInput schema / properties / is_suspected_automation
      Added value: +{
      +  "description": "Suspected automation (commit-farm/mass-repo bots); omitted these are hidden by default, pass true to isolate them",
      +  "type": "boolean"
      +}
  3. Addedv1.2.0

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, yet it only mentions that the tool searches enriched profiles in a search index. It does not disclose pagination behavior, default sort, the fact that all parameters are optional, how filters combine, whether an empty query is allowed, rate limits, or return-format expectations. For a search tool of this complexity, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, with the core purpose in the first sentence. The two enum listings are redundant with the schema and arguably unnecessary, but they are short and don't bloat the description substantially. No wasted filler or vague prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a 30-parameter search tool with no annotations and no output schema, so the description needs to provide substantial context. It does not explain pagination constraints, defaults, how parameters interact, the meaning of 'enriched' profiles, or what the response contains. The schema is rich, but the description leaves an agent without enough context to invoke the tool correctly in edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents all 30 parameters, including the influence_tier and sort enums. The description redundantly restates these two enums without adding new meaning or clarifying relationships between parameters (e.g., radius_m requiring lat/lon). Baseline 3 is appropriate because the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Search the GitHub users dataset' and 'Searches enriched public GitHub user profiles stored in a search index.' This clearly conveys a search operation over a distinct dataset, and the 'enriched' / 'search index' phrasing helps set it apart from raw GitHub API search siblings like github_search_users. It doesn't explicitly name sibling exclusions, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through the word 'Search' and the tool name, so an agent can infer this is for querying GitHub user profiles. However, it provides no explicit guidance on when to use this tool versus related siblings such as datasets_github_users_item (fetching a specific user), datasets_github_users_nearby (geographic search), datasets_github_users_facets (facet discovery), or github_search_users (GitHub API search). No exclusions or alternative-based selection criteria are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools