Developer tools · Lead generation
GitHub Repo Scraper
Fetch full GitHub repository metadata for one or many repos in one call — stars, forks, languages, topics, license, default branch, latest release, contributor count — export to JSON or CSV. A GitHub repo API wrapper; optional token for higher rate limits.
Free Apify credit covers a first run. No credit card to try.
What this Actor scrapes
GitHub exposes public repository data through its REST API, but turning a list of repos into a reliable dataset is messier than it looks: secondary-rate-limit errors kick in at burst speeds, the languages and release endpoints are separate calls, and unauthenticated requests cap out at 60 per hour. This GitHub repo scraper fans out requests in parallel, handles the retry dance automatically, and delivers one richly-typed row per repository — covering everything from stargazers_count through latest_release_tag to scraped_at.
Give it a list of owner/repo slugs or full GitHub URLs. It writes clean, Pydantic-validated rows straight into your Apify dataset. Use it for competitor benchmarking, OSS health checks, DevRel dashboards, AI/RAG corpus building, or any workflow that needs bulk GitHub repository data on demand.
What we handle for you
- 🛡️ Browser fingerprint rotation —
curl-cffiimpersonates real Chrome / Firefox / Safari TLS handshakes so requests look like a browser, not a Python script. - 🌐 Residential proxy rotation via Apify Proxy — fresh session and exit IP whenever the target pushes back.
- 🔁 Retries with exponential backoff on
408 / 429 / 5xx— up to 5 attempts per request,Retry-Afterheaders honoured. - 🧱 Rate-limit-aware pacing — when GitHub's secondary rate limit kicks in, we slow down and wait rather than hammering until banned.
- 🧊 Clean, typed dataset rows — Pydantic-validated, ISO-8601 timestamps, stable field names, JSON / CSV / Excel export straight from Apify Console.
- 💰 Pay-Per-Event pricing — you pay only for results that land in your dataset. No data, no charge.
Use cases
- Competitor OSS benchmarking — track stars and forks across rival projects week-over-week and pipe deltas to Slack or a BI tool.
- Dependency health monitoring — feed your stack's transitive repo list and flag anything archived, disabled, or unmaintained.
- RAG corpus building — pull language breakdowns and README metadata for a curated set of repos to seed a vector store.
- Hiring and M&A research — quantify the open-source surface area of a target company or candidate's personal GitHub activity.
- Newsletter automation — ingest a curated list weekly, diff the star counts, surface the fastest movers.
- DevRel dashboards — track your own org's repos alongside ecosystem repos in one unified dataset.
Input
Paste this into the Apify Console, or send it as the run input over the API. Proxy settings are on by default; you rarely need to touch them.
| Field | Type | Required | What it does |
|---|---|---|---|
repos | array | yes | List of repos in owner/repo form (e.g. apify/apify-sdk-python) or as full GitHub URLs. Each input becomes one dataset row. |
githubToken | string | no | Personal access token. Without one you get 60 requests/hour; with one, 5 000/hour. Read-only access to public repos is sufficient (no scopes needed). |
includeLanguages | boolean | no | Adds a `languages` map (language → bytes) per repo. One extra API call per repo. |
includeLatestRelease | boolean | no | Adds `latest_release_tag` and `latest_release_published_at`. One extra API call per repo. |
concurrency | integer | no | Parallel API requests. 8 is comfortable with a token; 2-3 without. |
{
"repos": [
"apify/apify-sdk-python",
"apify/crawlee-python"
],
"githubToken": "",
"includeLanguages": true,
"includeLatestRelease": true,
"concurrency": 4,
"proxyConfiguration": {
"useApifyProxy": false
}
} Output
One row per result, schema-validated before it is written. Export JSON, CSV, Excel or XML from the run, or read it over the API.
ownernamefull_namehtml_urldescriptionforkarchiveddisabledstargazers_countforks_countwatchers_countopen_issues_countsize_kblanguagelanguagestopics
{
"owner": "apify",
"name": "apify-sdk-python",
"full_name": "apify/apify-sdk-python",
"html_url": "https://github.com/apify/apify-sdk-python",
"description": "The Apify SDK for Python.",
"fork": false,
"archived": false,
"stargazers_count": 415,
"forks_count": 41,
"watchers_count": 415,
"open_issues_count": 12,
"size_kb": 2048,
"language": "Python",
"languages": {
"Python": 198432,
"Shell": 1024
},
"topics": [
"apify",
"scraping",
"sdk"
],
"license": "Apache-2.0",
"default_branch": "main",
"homepage": "https://docs.apify.com/sdk/python",
"created_at": "2022-08-01T10:00:00Z",
"updated_at": "2026-05-30T14:22:00Z",
"pushed_at": "2026-05-29T08:11:00Z",
"latest_release_tag": "v3.4.0",
"latest_release_published_at": "2026-05-20T12:00:00Z",
"scraped_at": "2026-06-01T09:00:00Z"
} Pricing
| Event | Price | When |
|---|---|---|
| Actor start | $0.20 | Once per run, covers warm-up and proxy session setup. |
| Result emitted | $0.0020 | Per result written to the dataset. |
You pay only for results that land. Cap any run with maxTotalChargeUsd. See pricing & billing for worked examples.
Limitations
- Private repos require a token with the matching scopes — this Actor only processes repos the token can read. Do not reuse production tokens here.
- README content, raw code, commit graphs, and pull request history are outside the scope of this Actor. Use GitHub's search API or a dedicated commits scraper for those.
- Large orgs with thousands of repos will hit the 5 000 req/hr authenticated ceiling on long runs. Plan batches or spread runs over multiple hours.
- GitHub caches some counts (stars, forks) for a few minutes. Compare runs at least 5 minutes apart to catch real movement.
FAQ
Do I need a GitHub token?
public_repo read-only scope is all you need.Is this a GitHub REST API alternative or replacement?
Can I use this to fetch github repo metadata api-style for hundreds of repos at once?
owner/repo slugs, set your token, and the Actor fans them out in parallel while respecting GitHub's rate limits. Output lands in a clean dataset ready to export or query.How is this different from calling the GitHub REST API myself?
What if a repo doesn't exist or has been deleted?
Can I scrape private repositories?
Ready to run it?
Open the listing on Apify, paste the input above, and watch rows land. If it ever breaks, it is our problem before it is yours.
Related Actors