Universe: the Tranco top-1,000,000 domain list. For each domain,
Crawlora's fingerprinting pipeline fetched the homepage over plain HTTP(S) — no browser
rendering — and ran the response (final HTML body + response headers) against a signature
database of known technology fingerprints, the same category of detection BuiltWith and
Wappalyzer use. Detector signals: <script src> URLs, <link href> URLs, <meta name="generator"> values, inline markup markers, preconnect hints, and response headers (with
regex version extraction where available).
- Scope is HTTP-only. A site that only exposes a technology via client-side JavaScript execution (no static marker in the initial HTML/headers) will not be detected by this pass. This dataset does not run a browser-rendering enrichment pass.
- Scope is Tranco-seeded only. Crawlora separately maintains an
internal-harvestcorpus (domains sourced from other product surfaces, not the Tranco list) — that corpus carries no Trancorankand is entirely out of scope for this release. Every table here is scoped torun_id=top1m-2026-08-25-full-v2,seed_source=tranco. - Detector version 2026.10.47. This specific run is the first full-scale pass carrying a
batch of fixes and additions: failure-diagnostics classification, an
is_infrastructureheuristic (not published in this release — see below), a plain-HTTPwww.fallback, 10 new Chinese-market technology signatures, WordPress/Shopify false-positive fixes, and a probe status-precedence fix. Comparing raw counts against an older run's detector version will conflate real adoption change with newly-added detection — don't do it without controlling for the signature-set change.
reachable=true: the homepage HTTP fetch succeeded and produced a body usable for detection.reachable=false: the probe attempt failed — connection error, timeout, a block/challenge page, or a bad HTTP status. This is not the same as "no technology detected." Areachable=falserecord is never written as "200, zero technologies" — it's excluded entirely from every adoption-rate denominator in this release. Of the 1,000,000-domain universe, 847,491 (84.7%) are reachable; the remaining 152,509 (15.3%) failed to fetch for reasons this release does not enumerate (the underlyingfailure_reason/probe_error/block_matcheddiagnostic fields are scan-quality telemetry, not tech-stack content, and are not published here).
The detector tags each individual technology match with a confidence level: high (direct
signature match — the default), or low (an implied technology added because a directly
detected one implies it, e.g. detecting Next.js implies React; added only if not already
directly detected). A medium tier is defined in the detector's tier list but this release does
not assert how often it's actually emitted in practice — treat any specific frequency claim
about medium as unverified. Per-technology confidence is not published as a queryable field
in this release (see data/sample-schema.md) — the underlying technologies[] detail is
stored but not indexed in the source system, so a confidence-tier breakdown would require a full
document scan rather than a simple aggregation, and hasn't been computed for this release.
Rank bands (Top 1,000 / 1K–10K / 10K–100K / 100K–1M) are computed by each domain's Tranco rank
at scan time — a pinned snapshot, not a live-queryable cut. The public
/datasets/techstack/* REST API supports sorting by rank but has no rank-range filter, so this
breakdown is derived directly from the underlying index rather than the API the rest of this
dataset is documented against. It will be refreshed alongside future full-census runs, on the
cadence described in CHANGELOG.md — not automatically.
The .cn-vs-rest-of-universe numbers in data/cn-visibility-gap.csv measure this detector's
ability to see a technology stack, not the sophistication, quality, or modernity of the
underlying sites. .cn domains are disproportionately built on infrastructure this detector's
signature database was not originally built to recognize — self-hosted stacks in the
Baidu/Alibaba/Tencent ecosystem, CDN and analytics vendors with limited Western adoption, and
markup conventions the signature set doesn't yet cover. A domain showing "zero technologies
detected" or a low tech count is a statement about signature coverage, not about what
technology the site actually runs. This release's detector version added 10 new Chinese-market
signatures in this exact refresh specifically to narrow this gap, and the gap narrowed measurably
as a result (see CHANGELOG.md) — treat the current gap as a coverage floor that continues to
shrink with future signature work, not a fixed property of the Chinese web.
data/sample-domains-10k.csv is a stratified systematic sample, not a random draw or a
naive top-N: within each of the 4 rank bands used elsewhere in this release, domains are drawn
at a fixed stride by Tranco rank so all four bands are represented roughly proportionally to
their role in the headline chart, not proportionally to their raw share of the 1,000,000-domain
universe (which would make the Top-1,000 band nearly invisible at ~0.1% of a naive random
sample). Exact method and row counts per band: data/sample-schema.md.
data/captcha-crosscheck.csv cites a CAPTCHA-adoption figure from Crawlora's separately-run
Anti-Bot Adoption Index — a different dataset, using a different detection method (live
challenge-response probing vs. this dataset's static HTML/header signature matching). It is
cited only as an external sanity check that the two independently-built pipelines land in the
same range; it is not folded into any total in this release.