In 2003, Roy Rosenzweig framed digital preservation around a paradox of scarcity and abundance. Two decades later, the same problem operates at a vastly greater scale and under more concentrated infrastructural conditions. Global data volume was projected to reach approximately 181 zettabytes in 2025, enough, if stored on 25 GB Blu-ray discs, to produce a stack extending some twenty-one times the distance from the Earth to the Moon. The vulnerability of digital records that concerned Rosenzweig has only become more consequential at this scale. Today's big data are distributed across servers, data centres and networks even as ownership and computational capacity have become increasingly concentrated since the Web 1.0 era. Replication, migration, modification and deletion belong to the ordinary operation of these systems. The abundance of the present may leave future historians with little more than scattered remnants of an information environment whose original scale was astronomical.
how the twenty-one is arrived at ↗
global data volume, 2025
≈181 zettabytes = 1.81 × 10²³ bytes — projected, and it counts copies
384,400 km, the mean; the separation varies by about 10 per cent
result
≈21 mean Earth–Moon distances. The binary reading is used throughout; the decimal reading (25 × 10⁹ bytes a disc) would give ≈7.2 trillion discs and ≈23 distances.
status
derived from the figures above — an illustration of scale, not a finding
Existing preservation efforts reveal the magnitude of the disparity. Between 1996 and 2020, the Internet Archive's Wayback Machine accumulated some 733 billion web objects occupying approximately 70 petabytes, an extraordinary scholarly resource that nevertheless amounted to about 0.00012 per cent of the 59 zettabytes of data produced globally in 2020. The quantities are not directly equivalent, but their difference in scale is instructive. More important, much of the contemporary record does not exist as a bounded object awaiting preservation. Social-media records, sensor outputs, user-behaviour data, audiovisual streams and machine-generated materials are sharded, replicated, cached, recombined, migrated and purged as part of the systems in which they operate. There may be no singular authoritative copy to deposit in an archive, while the interfaces, metadata, software and computational contexts required to interpret a surviving copy may disappear before the copy itself.
Archives of the Future studies this period before the archive. Preservation is one part of a larger problem encompassing the production, transmission, migration, authentication, ownership, accessibility and disappearance of digital evidence. A record documented as destroyed, one that was never captured, one known to exist but presently inaccessible, and one whose fate cannot be established represent distinct evidentiary conditions; only the first establishes loss. BDSL documents these processes while the systems that produce them can still be observed, using field research, interviews, technical and legal records, and direct investigation of contemporary information infrastructure. The aim is not to preserve the digital present at anything approaching its original scale. It is to leave a record of how that information was created, circulated, transformed and lost, so that the fragments that remain can be understood in relation to the systems that produced them.
figures cited above ↗
Rosenzweig 2003
Roy Rosenzweig, “Scarcity or Abundance? Preserving the Past in a Digital Era”, American Historical Review 108:3 (2003), 735–762.
181 zettabytes, 2025
A projection of data created, captured, copied and consumed worldwide (IDC Global DataSphere, widely republished). A projection, not a measurement, and it counts copies.
733 billion web objects · ~70 petabytes
Wayback Machine holdings for 1996–2020, as reported by the Internet Archive.
59 zettabytes, 2020
Global data volume for 2020 (IDC).
0.00012 per cent
Derived from the two figures above. It sets bytes stored against bytes produced, which are not equivalent quantities; it is offered as an order of magnitude, not a rate of coverage.
What existed, faintly; what survives, solid. Absence is uneven, and the surviving fragment does not disclose the history of what disappeared around it.
Examines large digital collections that became available as historical evidence by routes their creators rarely intended, from rescued web archives and data dumps to materials disclosed by leaks, hacks, and legal proceedings. The project asks how their histories of survival, custody, and access shape what historians can know from them.