Lead generation · Developer tools
GitHub User Scraper — Profiles, Repos & Orgs
Fetch GitHub user or organisation metadata via the GitHub API — name, bio, company, location, blog, public repo count, follower count, and optionally the user's pinned/public repositories — export to JSON or CSV. Free REST API, optional token for higher limits.
Free Apify credit covers a first run. No credit card to try.
What this Actor scrapes
Every GitHub user and organisation has a public profile page — and a matching REST API endpoint. This Actor ingests a list of usernames or org slugs, fans them out in parallel, and returns one structured row per profile: display name, bio, company, location, blog URL, public repo count, follower and following counts, hireable flag, account-creation timestamp, and an optional public-repos sub-array. Organisations and personal accounts share the same output shape; the type field tells them apart.
What we handle for you
- 🛡️ Browser fingerprint rotation —
curl-cffireplays real Chrome / Firefox / Safari TLS handshakes so requests look like a browser, not a Python script. - 🌐 Proxy rotation via Apify Proxy — fresh session and exit IP on every block; residential pool available on paid plans.
- 🔁 Retries with exponential backoff — up to 5 attempts per request on
408 / 429 / 5xx,Retry-Afterhonoured. - 🧱 Rate-limit-aware pacing — when the target pushes back, we slow down and surface a clear status message; you never get a silent empty dataset.
- 🧊 Clean, typed rows — Pydantic-validated output, ISO-8601 timestamps, stable
loginIDs, direct JSON / CSV / Excel export from Apify Console. - 💰 Pay-Per-Event pricing — you pay only for rows that reach your dataset. No data, no charge (minus the small
actor-startwarm-up fee).
Use cases
- Technical recruiting — find developers whose bio mentions a framework and whose
hireableflag istrue; enrich a sourcer's shortlist at scale. - B2B lead gen — pull GitHub orgs in a target geography, filter by
public_repos > N, then enrich with the org's website or LinkedIn. - ATS enrichment — pipe candidate GitHub profile URLs through this Actor to append repo count, follower count, and company to your applicant-tracker rows.
- Developer-OSINT — map a pseudonymous developer's public identity: linked Twitter handle, declared company, registered location, and account-creation date.
- DevRel contributor intelligence — rank contributors to a project by follower count and organisation, then prioritise outreach.
Input
Paste this into the Apify Console, or send it as the run input over the API. Proxy settings are on by default; you rarely need to touch them.
| Field | Type | Required | What it does |
|---|---|---|---|
usernames | array | yes | GitHub usernames or organisation slugs, one per line. Profile URLs are also accepted — we strip the host. |
githubToken | string | no | Personal access token. Without one you get 60 requests/hour; with one, 5 000/hour. |
includeRepos | boolean | no | Adds a `repos` array (up to 100 most-recently-updated public repos with name + stars). One extra API call per user. |
maxReposPerUser | integer | no | Cap on repos returned per user when includeRepos=true. Hard ceiling 100. |
concurrency | integer | no | Parallel API requests. |
{
"usernames": [
"apify",
"torvalds"
],
"includeRepos": true,
"maxReposPerUser": 10,
"concurrency": 4,
"proxyConfiguration": {
"useApifyProxy": false
}
} Output
One row per result, schema-validated before it is written. Export JSON, CSV, Excel or XML from the run, or read it over the API.
logintypenamecompanybloglocationemailbiotwitter_usernamepublic_repospublic_gistsfollowersfollowinghtml_urlavatar_urlhireable
{
"login": "apify",
"type": "Organization",
"name": "Apify",
"company": null,
"blog": "https://apify.com",
"location": "Prague, Czechia",
"email": null,
"bio": "The full-stack web scraping and browser automation platform.",
"public_repos": 412,
"followers": 856,
"hireable": null,
"html_url": "https://github.com/apify",
"created_at": "2012-09-21T08:07:25Z",
"scraped_at": "2025-11-01T10:30:00Z"
} Pricing
| Event | Price | When |
|---|---|---|
| Actor start | $0.20 | Once per run, covers warm-up and proxy session setup. |
| Result emitted | $0.0030 | Per result written to the dataset. |
You pay only for results that land. Cap any run with maxTotalChargeUsd. See pricing & billing for worked examples.
Limitations
- Only the
/users/{login}REST endpoint is called. Contribution graphs, sponsor tiers, starred repos, and commit activity are not included in this Actor. - Private-org membership and private repos require OAuth scopes that are outside what a read-only personal token provides — those fields are out of scope.
- GitHub's unauthenticated rate limit is 60 requests per hour per IP. For bulk runs of 500+ users, supply a GitHub token or enable Apify Proxy to spread requests across IPs.
- Data freshness reflects whatever GitHub returns at run time. Cached or recently-deleted accounts may return stale or empty rows.
FAQ
Is this a github profile scraper as well as a user scraper?
type field in each row tells you which you got (User vs Organization). If you search for "github profile scraper" this is the tool.What about the github contributor api — can I get contribution data?
/users/{login} endpoint publishes: bio, follower count, repo count, hireable flag, and (optionally) the public-repos list. For commit-level data you need the repository contributors endpoint, which is a separate scrape.Why is `email` null for most profiles?
noreply addresses.Are GitHub organisations supported?
type field distinguishes them.Does this respect GitHub's ToS and GDPR?
Can I run this at scale for recruiting?
Ready to run it?
Open the listing on Apify, paste the input above, and watch rows land. If it ever breaks, it is our problem before it is yours.
Related Actors