GitHub User Scraper icon

Lead generation · Developer tools

GitHub User Scraper — Profiles, Repos & Orgs

Fetch GitHub user or organisation metadata via the GitHub API — name, bio, company, location, blog, public repo count, follower count, and optionally the user's pinned/public repositories — export to JSON or CSV. Free REST API, optional token for higher limits.

Free Apify credit covers a first run. No credit card to try.

What this Actor scrapes

Every GitHub user and organisation has a public profile page — and a matching REST API endpoint. This Actor ingests a list of usernames or org slugs, fans them out in parallel, and returns one structured row per profile: display name, bio, company, location, blog URL, public repo count, follower and following counts, hireable flag, account-creation timestamp, and an optional public-repos sub-array. Organisations and personal accounts share the same output shape; the type field tells them apart.

What we handle for you

  • 🛡️ Browser fingerprint rotationcurl-cffi replays real Chrome / Firefox / Safari TLS handshakes so requests look like a browser, not a Python script.
  • 🌐 Proxy rotation via Apify Proxy — fresh session and exit IP on every block; residential pool available on paid plans.
  • 🔁 Retries with exponential backoff — up to 5 attempts per request on 408 / 429 / 5xx, Retry-After honoured.
  • 🧱 Rate-limit-aware pacing — when the target pushes back, we slow down and surface a clear status message; you never get a silent empty dataset.
  • 🧊 Clean, typed rows — Pydantic-validated output, ISO-8601 timestamps, stable login IDs, direct JSON / CSV / Excel export from Apify Console.
  • 💰 Pay-Per-Event pricing — you pay only for rows that reach your dataset. No data, no charge (minus the small actor-start warm-up fee).

Use cases

  • Technical recruiting — find developers whose bio mentions a framework and whose hireable flag is true; enrich a sourcer's shortlist at scale.
  • B2B lead gen — pull GitHub orgs in a target geography, filter by public_repos > N, then enrich with the org's website or LinkedIn.
  • ATS enrichment — pipe candidate GitHub profile URLs through this Actor to append repo count, follower count, and company to your applicant-tracker rows.
  • Developer-OSINT — map a pseudonymous developer's public identity: linked Twitter handle, declared company, registered location, and account-creation date.
  • DevRel contributor intelligence — rank contributors to a project by follower count and organisation, then prioritise outreach.

Input

Paste this into the Apify Console, or send it as the run input over the API. Proxy settings are on by default; you rarely need to touch them.

FieldTypeRequiredWhat it does
usernames array yes GitHub usernames or organisation slugs, one per line. Profile URLs are also accepted — we strip the host.
githubToken string no Personal access token. Without one you get 60 requests/hour; with one, 5 000/hour.
includeRepos boolean no Adds a `repos` array (up to 100 most-recently-updated public repos with name + stars). One extra API call per user.
maxReposPerUser integer no Cap on repos returned per user when includeRepos=true. Hard ceiling 100.
concurrency integer no Parallel API requests.
{
  "usernames": [
    "apify",
    "torvalds"
  ],
  "includeRepos": true,
  "maxReposPerUser": 10,
  "concurrency": 4,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}

Output

One row per result, schema-validated before it is written. Export JSON, CSV, Excel or XML from the run, or read it over the API.

logintypenamecompanybloglocationemailbiotwitter_usernamepublic_repospublic_gistsfollowersfollowinghtml_urlavatar_urlhireable

{
  "login": "apify",
  "type": "Organization",
  "name": "Apify",
  "company": null,
  "blog": "https://apify.com",
  "location": "Prague, Czechia",
  "email": null,
  "bio": "The full-stack web scraping and browser automation platform.",
  "public_repos": 412,
  "followers": 856,
  "hireable": null,
  "html_url": "https://github.com/apify",
  "created_at": "2012-09-21T08:07:25Z",
  "scraped_at": "2025-11-01T10:30:00Z"
}

Pricing

EventPriceWhen
Actor start$0.20Once per run, covers warm-up and proxy session setup.
Result emitted$0.0030Per result written to the dataset.

You pay only for results that land. Cap any run with maxTotalChargeUsd. See pricing & billing for worked examples.

Limitations

  • Only the /users/{login} REST endpoint is called. Contribution graphs, sponsor tiers, starred repos, and commit activity are not included in this Actor.
  • Private-org membership and private repos require OAuth scopes that are outside what a read-only personal token provides — those fields are out of scope.
  • GitHub's unauthenticated rate limit is 60 requests per hour per IP. For bulk runs of 500+ users, supply a GitHub token or enable Apify Proxy to spread requests across IPs.
  • Data freshness reflects whatever GitHub returns at run time. Cached or recently-deleted accounts may return stale or empty rows.

FAQ

Is this a github profile scraper as well as a user scraper?
Yes — the same Actor handles both personal user accounts and organisation profiles. The type field in each row tells you which you got (User vs Organization). If you search for "github profile scraper" this is the tool.
What about the github contributor api — can I get contribution data?
GitHub's public REST API does not expose a user's contribution graph or commit-activity heatmap. This Actor returns what the /users/{login} endpoint publishes: bio, follower count, repo count, hireable flag, and (optionally) the public-repos list. For commit-level data you need the repository contributors endpoint, which is a separate scrape.
Why is `email` null for most profiles?
GitHub hides email addresses by default. We surface exactly what the public API returns; we do not probe commit metadata or decode noreply addresses.
Are GitHub organisations supported?
Yes — org slugs and personal usernames use the same endpoint and produce the same row shape. The type field distinguishes them.
Does this respect GitHub's ToS and GDPR?
We request only publicly published profile fields via the official REST API — no login walls, no scraping of private content. Whether your downstream use of the data complies with GDPR or GitHub's Acceptable Use Policy is your legal call; we surface the data, you own the compliance obligation.
Can I run this at scale for recruiting?
With a GitHub token you get 5 000 API requests per hour, enough for a few thousand profiles per run. We handle rate-limit responses and back off automatically; you set the concurrency and we do the rest.

Ready to run it?

Open the listing on Apify, paste the input above, and watch rows land. If it ever breaks, it is our problem before it is yours.

Related Actors

Teams that run this also run