Big Data Studies LabThe Humanities at Infrastructural Scale
Area
04 of 05
Question
How do machines come to know us?
Projects
Accidental Archives
The Personalized Web

Research · 04

Data Doubles

How do machines come to know us?

At the turn of the millennium, Kevin Haggerty and Richard Ericson described an emerging surveillant assemblage in which information collected about people in different settings could be abstracted, recombined, and acted upon as data doubles. Their formulation belonged to the world of Web 1.0, CCTV, transactional databases, identification systems, and comparatively dispersed networks of surveillance. Web 2.0 altered both the scale and organization of this assemblage. Search histories, clicks, pauses, locations, social relationships, images, facial features, and other behavioural traces became inputs to increasingly centralized systems designed to distinguish among users and anticipate what they might want or do next. Convenience and surveillance developed from much of the same machinery: recommendation required finer differentiation, engagement optimization closer observation of attention, and personalization more extensive inference. Shoshana Zuboff’s account of surveillance capitalism captured the economic significance of turning human experience into behavioural data and prediction. The underlying machinery has since become more automated, concentrated, and deeply embedded in everyday life.

Generative AI intensifies these conditions further. Earlier systems principally observed behaviour and derived information from the traces users left behind; conversational systems increasingly invite people to disclose personal histories, relationships, ambitions, anxieties, preferences, uncertainties, and unfinished thoughts directly to machines. Users may also turn to these systems for advice, companionship, emotional support, and consequential decisions, creating sustained exchanges from which computational representations can be revised. A data double under these conditions is a changing, model-dependent representation assembled from behavioural observations, inferred characteristics, predictions, and direct disclosure. It is also relational, since systems locate individuals among patterns learned from other people. Different models can construct different doubles from the same evidence, while subsequent interactions provide further information from which those representations may change.

BDSL investigates this history by reconstructing systems whose most consequential operations remain only partially observable. Successive application releases, decompiled code, interfaces, and version histories reveal changes in client-side collection and design; algorithm papers, technical documentation, and legal proceedings provide different forms of evidence about recommendation, engagement, facial modification, personalization, and server-side inference. BDSL distinguishes what systems observe from what they infer and predict, and examines how those predictions shape what users subsequently encounter. An inferred preference can reorder a feed, recommend a video, alter an image, or shape a conversational response. These actions change the environment in which the next behaviour or disclosure occurs, generating new evidence from which the computational representation can be revised. Data Doubles studies this recursive relationship between people and machines, extending the surveillant assemblage into an age of increasingly centralized systems that observe their users, solicit their disclosures, construct predictions about them, and act on what they infer.

A data double has no stable outline. Behavioural traces form overlapping and changing clusters as models derive, infer, and predict from them. New observations alter those configurations, while different models can construct different representations from the same traces.

Projects

Relevant Publications