AI & LLM data · Developer tools
arXiv Scraper — Search & Export Paper Metadata
Search arXiv by query, category, or author and export structured paper metadata — title, authors, abstract, primary category, DOI, PDF URL, submitted and updated timestamps — to JSON or CSV. An arXiv API wrapper that handles pagination, retries, and rate-limit pacing for your pipeline.
Free Apify credit covers a first run. No credit card to try.
What this Actor scrapes
arXiv's Atom feed at export.arxiv.org/api/query is the canonical source for preprint paper metadata. It is also paginated, rate-limited, and quick to push back on anything that looks like aggressive bulk access. This Actor wraps it with a polished input schema, paces requests to stay within arXiv's courtesy guidelines, paginates automatically across large result sets, and writes one structured row per paper. We absorb the transient errors and pushback; you get a dataset that drops cleanly into research dashboards, citation-tracking tools, RAG pipelines, or ML training corpora.
Looking to download arXiv papers metadata across an entire category slice (for example, all cs.AI submissions from 2025)? A single sweep like that can exceed 30 000 records — hours of hand-rolled pagination that we handle end-to-end.
What we handle for you
- 🛡️ Browser fingerprint rotation —
curl-cffiimpersonates real Chrome / Firefox / Safari TLS handshakes so the upstream sees a browser, not a Python script. - 🌐 Residential proxy rotation via Apify Proxy — fresh session ID and exit IP whenever the upstream pushes back.
- 🔁 Retries with exponential backoff on
408 / 429 / 5xx— up to 5 attempts per page,Retry-Afterheader honoured. - 🧱 Rate-limit-aware pacing — we slow down rather than accumulate blocks; partial progress is always surfaced, never silently dropped.
- 🧊 Clean, typed dataset rows — Pydantic-validated fields, ISO-8601 timestamps, stable IDs. Export as JSON, CSV, or Excel directly from Apify Console.
- 💰 Pay-Per-Event pricing — you pay only for results that land in your dataset. No data, no charge beyond the small start fee.
Use cases
- RAG corpus building — pull every
cs.AI/cs.LG/cs.CLabstract from the past year and load it straight into ChromaDB, Pinecone, or Weaviate for semantic search over papers. - Citation tracking — schedule weekly runs for
au:<your-name>and diff to detect new citations of your work. - Trend monitoring — daily pull from a specific category to feed a research digest or newsletter.
- Dataset curation — extract all papers matching a topic + date range to seed a systematic literature review or benchmark evaluation.
- Notification pipeline — stream new results into Slack or Discord when a paper matches a saved query.
- VC / competitive intelligence — map research output by lab, author, or topic over time to surface emerging areas.
Input
Paste this into the Apify Console, or send it as the run input over the API. Proxy settings are on by default; you rarely need to touch them.
| Field | Type | Required | What it does |
|---|---|---|---|
searchQuery | string | yes | arXiv search query string. Use field prefixes like ti: (title), au: (author), cat: (category). Examples: cat:cs.AI, ti:transformer AND au:vaswani. |
sortBy | string | no | Field used to order results. |
sortOrder | string | no | Ascending or descending. |
maxResults | integer | no | Total papers to fetch across pages. arXiv recommends ≤30000 per query. Default 50. |
pageSize | integer | no | Papers per API call. arXiv caps page size at 2000; default 50. |
{
"searchQuery": "cat:cs.AI",
"sortBy": "submittedDate",
"sortOrder": "descending",
"maxResults": 3,
"pageSize": 3,
"proxyConfiguration": {
"useApifyProxy": false
}
} Output
One row per result, schema-validated before it is written. Export JSON, CSV, Excel or XML from the run, or read it over the API.
arxiv_idurlpdf_urltitlesummaryauthorsprimary_categorycategoriesdoijournal_refcommentpublishedupdatedscraped_at
{
"arxiv_id": "2401.12345v2",
"url": "https://arxiv.org/abs/2401.12345v2",
"pdf_url": "https://arxiv.org/pdf/2401.12345v2",
"title": "Scaling Laws for Sparse Mixture-of-Experts Language Models",
"authors": [
"Alex Doe",
"Jamie Smith"
],
"primary_category": "cs.CL",
"categories": [
"cs.CL",
"cs.LG"
],
"doi": null,
"journal_ref": null,
"comment": "Accepted at NeurIPS 2025",
"published": "2026-04-12T16:00:00+00:00",
"updated": "2026-04-14T09:00:00+00:00",
"scraped_at": "2026-06-01T10:00:00+00:00"
} Pricing
| Event | Price | When |
|---|---|---|
| Actor start | $0.20 | Once per run, covers warm-up and proxy session setup. |
| Result emitted | $0.0015 | Per result written to the dataset. |
You pay only for results that land. Cap any run with maxTotalChargeUsd. See pricing & billing for worked examples.
Limitations
- Metadata only — the Actor uses the Atom API. Full-text search over PDF content is not supported; queries operate on arXiv metadata fields (title, abstract, authors, categories).
- Author disambiguation — arXiv does not expose canonical author IDs in the public API. Resolving name collisions across similar author strings is left to the caller.
- 30 000-record soft ceiling — arXiv's own documentation recommends keeping single queries under 30 000 results. The Actor enforces a polite inter-request delay to stay within arXiv's rate-limit guidance; very large sweeps will take proportionally longer.
- Preprint freshness window — newly submitted papers typically appear in the feed within 1–2 hours of arXiv ingest, but that window is not guaranteed.
FAQ
Is this legal?
User-Agent per their documentation.What is the arXiv API and can I use it directly?
export.arxiv.org) is a free public endpoint for querying paper metadata. You can query it directly with Python using the arxiv PyPI library or raw HTTP calls — but at scale, you will hit pagination complexity, rate-limit pushback, and XML parsing overhead. This Actor handles all of that and writes clean structured rows without you touching a single namespace.I already know Python — why not just write an arXiv API wrapper?
Can I download PDFs?
pdf_url field for every paper. You can pass that URL to a follow-up Actor or a curl loop to fetch the actual files.Why do some records have a null DOI?
null for those entries so your pipeline can handle them gracefully.How do I target a specific date range?
submittedDate:[YYYYMMDD TO YYYYMMDD]. Include that expression in your searchQuery field, e.g. cat:cs.AI AND submittedDate:[20250101 TO 20251231].Ready to run it?
Open the listing on Apify, paste the input above, and watch rows land. If it ever breaks, it is our problem before it is yours.
Related Actors