Big Data Studies LabThe Humanities at Infrastructural Scale
Area
01 of 05
Question
What will survive of the digital present?
Projects
Accidental Archives

Research · 01

Archives of the Future

What will survive of the digital present?

In 2003, Roy Rosenzweig framed digital preservation around a paradox of scarcity and abundance. Two decades later, the same problem operates at a vastly greater scale and under more concentrated infrastructural conditions. Global data volume was projected to reach approximately 181 zettabytes in 2025, enough, if stored on 25 GB Blu-ray discs, to produce a stack extending some twenty-one times the distance from the Earth to the Moon. The vulnerability of digital records that concerned Rosenzweig has only become more consequential at this scale. Today's big data are distributed across servers, data centres and networks even as ownership and computational capacity have become increasingly concentrated since the Web 1.0 era. Replication, migration, modification and deletion belong to the ordinary operation of these systems. The abundance of the present may leave future historians with little more than scattered remnants of an information environment whose original scale was astronomical.

181 ZBglobal data volume2025, projected÷one Blu-ray disc25 GB= ≈6.7trilliondiscs, on edge1.2 mm eachstacked×1×5×10×15×20×21the Moon384,400 km≈8.1 million km21 Earth–Moon distancesa unit is a distance,not a disc
how the twenty-one is arrived at ↗
global data volume, 2025
≈181 zettabytes = 1.81 × 10²³ bytes — projected, and it counts copies
disc capacity
25 GB single-layer Blu-ray, read in binary units: 25 × 1024³ bytes = 2.68 × 10¹⁰ bytes
discs required
1.81 × 10²³ ÷ 2.68 × 10¹⁰ ≈ 6.7 × 10¹² discs
disc thickness
1.2 mm, from the Blu-ray specification
stack height
6.7 × 10¹² × 1.2 mm ≈ 8.1 × 10⁶ km
Earth–Moon distance
384,400 km, the mean; the separation varies by about 10 per cent
result
≈21 mean Earth–Moon distances. The binary reading is used throughout; the decimal reading (25 × 10⁹ bytes a disc) would give ≈7.2 trillion discs and ≈23 distances.
status
derived from the figures above — an illustration of scale, not a finding

Existing preservation efforts reveal the magnitude of the disparity. Between 1996 and 2020, the Internet Archive's Wayback Machine accumulated some 733 billion web objects occupying approximately 70 petabytes, an extraordinary scholarly resource that nevertheless amounted to about 0.00012 per cent of the 59 zettabytes of data produced globally in 2020. The quantities are not directly equivalent, but their difference in scale is instructive. More important, much of the contemporary record does not exist as a bounded object awaiting preservation. Social-media records, sensor outputs, user-behaviour data, audiovisual streams and machine-generated materials are sharded, replicated, cached, recombined, migrated and purged as part of the systems in which they operate. There may be no singular authoritative copy to deposit in an archive, while the interfaces, metadata, software and computational contexts required to interpret a surviving copy may disappear before the copy itself.

Archives of the Future studies this period before the archive. Preservation is one part of a larger problem encompassing the production, transmission, migration, authentication, ownership, accessibility and disappearance of digital evidence. A record documented as destroyed, one that was never captured, one known to exist but presently inaccessible, and one whose fate cannot be established represent distinct evidentiary conditions; only the first establishes loss. BDSL documents these processes while the systems that produce them can still be observed, using field research, interviews, technical and legal records, and direct investigation of contemporary information infrastructure. The aim is not to preserve the digital present at anything approaching its original scale. It is to leave a record of how that information was created, circulated, transformed and lost, so that the fragments that remain can be understood in relation to the systems that produced them.

figures cited above ↗
Rosenzweig 2003
Roy Rosenzweig, “Scarcity or Abundance? Preserving the Past in a Digital Era”, American Historical Review 108:3 (2003), 735–762.
181 zettabytes, 2025
A projection of data created, captured, copied and consumed worldwide (IDC Global DataSphere, widely republished). A projection, not a measurement, and it counts copies.
733 billion web objects · ~70 petabytes
Wayback Machine holdings for 1996–2020, as reported by the Internet Archive.
59 zettabytes, 2020
Global data volume for 2020 (IDC).
0.00012 per cent
Derived from the two figures above. It sets bytes stored against bytes produced, which are not equivalent quantities; it is offered as an order of magnitude, not a rate of coverage.
What existed, faintly; what survives, solid. Absence is uneven, and the surviving fragment does not disclose the history of what disappeared around it.

Projects

Accidental Archives

Examines large digital collections that became available as historical evidence by routes their creators rarely intended, from rescued web archives and data dumps to materials disclosed by leaks, hacks, and legal proceedings. The project asks how their histories of survival, custody, and access shape what historians can know from them.

ActiveArchives of the Future Machine-Assisted Reading

Publications

  • 2023

    Big Data Studies: The Humanities in Uncharted Waters

    Cha, Javier · Korean Studies 47: 274–299

  • 2020

    Pik teit'ŏ wa inmunhak ŭi mirae [Big data and the future of humanities]

    Cha, Javier · Munmyŏng kwa kyŏnggye 3: 43–77