Data collection

Web-scale corpus collection that stays responsible

Build training and retrieval corpora without breaking the web.

$0.29

per GB at scale

38 ms

datacenter latency

10 Gbps

uplinks throughout

The problem

Training and retrieval-augmented systems need web data at a scale that will absolutely take down small sites if collected carelessly. The reputational and legal consequences of that are now material, and publishers are watching.

At the same time, single-address collection at corpus scale is simply not possible — you will be blocked before you finish a domain.

The solution

Distributed exits make the volume achievable, and per-domain concurrency caps make it responsible. Both matter, and most providers only give you the first.

Datacenter proxies carry the bulk of the load cheaply; residential handles the minority of sources that filter hosting ASNs.

Mechanics

How proxies solve it

1

Corpus-scale throughput

Datacenter bandwidth at scale lands near 29 cents per gigabyte, which is what makes web-scale collection affordable.

2

Per-domain rate control

Cap concurrency and requests per second per host so a long tail domain is never overwhelmed.

3

Escalate only where needed

Route the small share of ASN-filtering sources to residential automatically rather than paying residential rates for everything.

4

Provenance for every document

Exit country, timestamp and status stored per fetch, which is increasingly a dataset governance requirement.

Workflow

How we would build it

  1. 1

    Read robots.txt and mean it

    Honour crawl-delay and disallow rules. It costs you very little and it is the difference between a crawler and a nuisance.

  2. 2

    Tier your sources

    Datacenter by default, residential only for domains that demonstrably reject it. The cost delta is roughly tenfold.

  3. 3

    Cap per-domain load

    Set concurrency and RPS limits per host in the dashboard, not just globally.

  4. 4

    Record provenance per document

    URL, fetch time, exit country, status and content hash. Governance reviews will ask for all five.

FAQ

AI & training data questions

Yes, subject to the Acceptable Use Policy. You must honour robots.txt, respect crawl delays and not collect personal data without a lawful basis. We will suspend accounts generating substantiated publisher complaints.
Set per-domain concurrency and requests-per-second caps in the dashboard. They are enforced at the gateway, so a bug in your crawler cannot bypass them.
Get started

Ready to start ai & training data?

Your first gigabyte is free, which is normally enough to validate the approach against your real target before you commit to anything.