Big Data Studies LabThe Humanities at Infrastructural Scale

Big Data Studies Lab

The Humanities at Infrastructural Scale

Founded in Seoul in 2019 and based at the University of Hong Kong since 2022, the Big Data Studies Lab (BDSL) brings together historians, media theorists, area specialists, data scientists, engineers, entrepreneurs, and policymakers to investigate how contemporary information systems are changing the sources, methods, and material conditions of humanities research. BDSL takes the familiar 3Vs of big data (volume, velocity, and variety) as a point of departure for examining the Zettabyte era historically.

BDSL follows digital information at scales far beyond individual observation as it is generated, distributed, transformed, preserved, and lost. Computational experiments, field research, interviews, software archaeology, and close reading of technical and legal records provide the empirical basis for interpretation and, where warranted, broader theorization.

Founded
Seoul National University · 2019
Based at
University of Hong Kong · 2022–
Research areas
5
Projects
4
Team
5
01
  • preservation
  • provenance
  • transmission
  • obsolescence
  • access
  • loss

Archives of the Future

What will survive of the digital present?

The abundance of the digital present may become the scarcity of the historical future. Much of today’s digital record depends on privately operated systems engineered for immediate access, where replication, migration, modification, and deletion are routine. Only a minute fraction is likely to remain accessible to future historians. Archives of the Future studies the conditions under which digital records persist, disappear, or become inaccessible, much of this history unfolding before archival preservation can even begin.

02
  • data centres
  • energy
  • water
  • semiconductors
  • critical minerals
  • supply chains
  • geopolitics

Material Infrastructures

What does big data require in order to exist?

Big data depends on data centres, electrical grids, cooling and water systems, processors, semiconductors, fibre networks, critical minerals, and the institutions that finance and govern them. Material Infrastructures studies these dependencies as part of the history of information itself, connecting the physical systems that sustain computation with the corporate, political, and geopolitical arrangements that make them possible.

03
  • velocity
  • distribution
  • latency
  • synchronization
  • real time
  • jurisdiction

Information in Motion

How does big data circulate across space, systems, and jurisdictions?

Contemporary data are distributed and dynamic, sharded across storage systems, replicated for availability, cached near anticipated demand, and recombined when requested. Information in Motion studies the temporal and geographic conditions of their circulation, asking how distributed systems approximate “real time” across physical distance, network latency, and jurisdictional boundaries.

04
  • personalization
  • engagement
  • attention
  • inference
  • recommendation
  • prediction

Data Doubles

How do machines come to know us?

Personalization increasingly depends on intimate computational representations of the people who use digital systems. Data Doubles studies how behavioural traces and personal disclosures become inferences and predictions, how different models construct different versions of the same person, and how those representations shape what systems subsequently show, recommend, and say to their users.

05
  • scale
  • semantic search
  • representation
  • comparison
  • multimodality
  • provenance

Machine-Assisted Reading

How can humanists investigate evidence beyond individual inspection?

Keyword search helps researchers find what they know enough to seek. Machine-Assisted Reading explores how digital collections might recover something of the serendipity of walking through the stacks, revealing relationships one did not know to look for across bodies of evidence too large and varied for individual inspection. Semantic, visual, relational, and multimodal methods extend this field of discovery while preserving a route from computational representations back to their sources.

Projects

Accidental Archives

Examines large digital collections that became available as historical evidence by routes their creators rarely intended, from rescued web archives and data dumps to materials disclosed by leaks, hacks, and legal proceedings. The project asks how their histories of survival, custody, and access shape what historians can know from them.

ActiveArchives of the Future Machine-Assisted Reading

Making Real Time

Examines how distributed information systems produce the conditions experienced as “real time,” tracing the relationship among physical distance, network latency, computation, caching, synchronization, and human perception.

On holdInformation in Motion

What Computing Takes

Follows the material demands of large-scale computing from electricity, water, and data centres to semiconductors, critical minerals, and global supply chains. The project examines how these infrastructures intersect with the lives and places that sustain them, and how their expansion connects resource security, industrial policy, and geopolitics.

ActiveMaterial Infrastructures

Methods

  • Field Research

    Investigate the physical sites and local conditions of large-scale computing, combining direct observation, interviews, and other evidence to understand how digital infrastructure is built, operated, and experienced.

  • Documentary Research

    Recover the workings of systems that are only partially public by reading technical papers, standards, patents, white papers, and legal proceedings for what they reveal, omit, or disclose indirectly.

  • Software Archaeology

    Reconstruct changes in software from successive releases, reading decompiled code, interfaces, and version histories against the surrounding documentary record.

  • Computational Experiments

    Design controlled experiments to measure system behaviour, test technical claims, and establish what specific operations do under defined conditions.

  • Evidence Modelling

    Develop computational representations of incomplete or heterogeneous evidence in a manner that preserves provenance, uncertainty, and distinctions required by the research question.

  • Machine Learning

    Assemble machine-learning systems for research questions and bodies of evidence, combining and adapting models, representations, and analytical procedures as needed to retrieve, compare, classify, cluster, and identify candidates for humanistic interpretation.