# Two files, and why they do not join

`eshop-audit-anonymised.csv` carries what each measured site is built on and how it is
configured. `eshop-weight-anonymised.csv` carries how heavy it is.

**They are deliberately not joinable.** The weight file is shuffled independently under a
recorded seed and its `row` ids correspond to nothing in the other file.

The reason is in section 07 of the paper. In short: image count, script count, HTML bytes and
wire bytes identified 726 of 738 reachable rows uniquely, the sample frame is public and named
in the crawler, so joining the two halves was a naming exercise for anyone willing to re-crawl.
Splitting them keeps every published aggregate exactly and removes the line between a byte
count and a security posture.

**What you can still do:** reproduce every figure in the paper. Weight and alt-coverage
percentiles come from the weight file, everything else from the main file, and the compound
failure cross-tab runs on the main file alone because the alt measure it needs is kept there as
a boolean.

**What you cannot do:** rebuild a per-site row. That is the point.

Rows: 900 measured, 738 reachable. Licence: CC BY 4.0, per `licence.txt`.
