crawlora-net/dead-web-commoncrawl
Viewer β’ Updated β’ 173M β’ 62
None defined yet.
Crawlora is a hosted web-data API: structured, normalized JSON from the public web β
search/SERP, e-commerce, social, finance, app stores, media, and reviews β instead of
HTML to parse. We also publish the open datasets behind our own data-journalism studies,
CC BY 4.0, hosted here for anyone to load with datasets, pandas, or DuckDB.
| Dataset | Study |
|---|---|
| dead-web-index | Dead-Web Index β is a domain alive, redirecting, blocked, or dead? |
| anti-bot-adoption-index | Anti-Bot Adoption Index β which WAF/anti-bot vendor protects the top web |
| ai-crawler-blocking-index | AI-Crawler Blocking Index β who blocks GPTBot, ClaudeBot, CCBot |
| steam-catalog | State of Steam 2026 β top 10,000 Steam games |
| song-length-index | Are Songs Really Getting Shorter? β 6,618 Hot 100 hits, 1959β2025 |
Browse all 11+ datasets on this org or the full catalog at crawlora.net/datasets.
The same public API backs SDKs (Go, TypeScript, Python, Ruby, Java, PHP), a hosted MCP server with 778 tools, Agent Skills, and n8n/Zapier/Make integrations β see the GitHub org for all of them.