Skip to content

Latest commit

 

History

History
91 lines (76 loc) · 6.03 KB

File metadata and controls

91 lines (76 loc) · 6.03 KB

Methodology

What was crawled and how

Universe: the Tranco top-1,000,000 domain list. For each domain, Crawlora's fingerprinting pipeline fetched the homepage over plain HTTP(S) — no browser rendering — and ran the response (final HTML body + response headers) against a signature database of known technology fingerprints, the same category of detection BuiltWith and Wappalyzer use. Detector signals: <script src> URLs, <link href> URLs, <meta name="generator"> values, inline markup markers, preconnect hints, and response headers (with regex version extraction where available).

  • Scope is HTTP-only. A site that only exposes a technology via client-side JavaScript execution (no static marker in the initial HTML/headers) will not be detected by this pass. This dataset does not run a browser-rendering enrichment pass.
  • Scope is Tranco-seeded only. Crawlora separately maintains an internal-harvest corpus (domains sourced from other product surfaces, not the Tranco list) — that corpus carries no Tranco rank and is entirely out of scope for this release. Every table here is scoped to run_id=top1m-2026-08-25-full-v2, seed_source=tranco.
  • Detector version 2026.10.47. This specific run is the first full-scale pass carrying a batch of fixes and additions: failure-diagnostics classification, an is_infrastructure heuristic (not published in this release — see below), a plain-HTTP www. fallback, 10 new Chinese-market technology signatures, WordPress/Shopify false-positive fixes, and a probe status-precedence fix. Comparing raw counts against an older run's detector version will conflate real adoption change with newly-added detection — don't do it without controlling for the signature-set change.

reachable — what it does and does not mean

  • reachable=true: the homepage HTTP fetch succeeded and produced a body usable for detection.
  • reachable=false: the probe attempt failed — connection error, timeout, a block/challenge page, or a bad HTTP status. This is not the same as "no technology detected." A reachable=false record is never written as "200, zero technologies" — it's excluded entirely from every adoption-rate denominator in this release. Of the 1,000,000-domain universe, 847,491 (84.7%) are reachable; the remaining 152,509 (15.3%) failed to fetch for reasons this release does not enumerate (the underlying failure_reason/probe_error/block_matched diagnostic fields are scan-quality telemetry, not tech-stack content, and are not published here).

Confidence tiers

The detector tags each individual technology match with a confidence level: high (direct signature match — the default), or low (an implied technology added because a directly detected one implies it, e.g. detecting Next.js implies React; added only if not already directly detected). A medium tier is defined in the detector's tier list but this release does not assert how often it's actually emitted in practice — treat any specific frequency claim about medium as unverified. Per-technology confidence is not published as a queryable field in this release (see data/sample-schema.md) — the underlying technologies[] detail is stored but not indexed in the source system, so a confidence-tier breakdown would require a full document scan rather than a simple aggregation, and hasn't been computed for this release.

Rank-band methodology

Rank bands (Top 1,000 / 1K–10K / 10K–100K / 100K–1M) are computed by each domain's Tranco rank at scan time — a pinned snapshot, not a live-queryable cut. The public /datasets/techstack/* REST API supports sorting by rank but has no rank-range filter, so this breakdown is derived directly from the underlying index rather than the API the rest of this dataset is documented against. It will be refreshed alongside future full-census runs, on the cadence described in CHANGELOG.md — not automatically.

What this dataset is NOT (read before citing the .cn figures)

The .cn-vs-rest-of-universe numbers in data/cn-visibility-gap.csv measure this detector's ability to see a technology stack, not the sophistication, quality, or modernity of the underlying sites. .cn domains are disproportionately built on infrastructure this detector's signature database was not originally built to recognize — self-hosted stacks in the Baidu/Alibaba/Tencent ecosystem, CDN and analytics vendors with limited Western adoption, and markup conventions the signature set doesn't yet cover. A domain showing "zero technologies detected" or a low tech count is a statement about signature coverage, not about what technology the site actually runs. This release's detector version added 10 new Chinese-market signatures in this exact refresh specifically to narrow this gap, and the gap narrowed measurably as a result (see CHANGELOG.md) — treat the current gap as a coverage floor that continues to shrink with future signature work, not a fixed property of the Chinese web.

Sampling method (row-level sample)

data/sample-domains-10k.csv is a stratified systematic sample, not a random draw or a naive top-N: within each of the 4 rank bands used elsewhere in this release, domains are drawn at a fixed stride by Tranco rank so all four bands are represented roughly proportionally to their role in the headline chart, not proportionally to their raw share of the 1,000,000-domain universe (which would make the Top-1,000 band nearly invisible at ~0.1% of a naive random sample). Exact method and row counts per band: data/sample-schema.md.

Independent cross-validation

data/captcha-crosscheck.csv cites a CAPTCHA-adoption figure from Crawlora's separately-run Anti-Bot Adoption Index — a different dataset, using a different detection method (live challenge-response probing vs. this dataset's static HTML/header signature matching). It is cited only as an external sanity check that the two independently-built pipelines land in the same range; it is not folded into any total in this release.