Big Data Studies LabThe Humanities at Infrastructural Scale
Area
05 of 05
Question
How can humanists investigate evidence beyond individual inspection?
Projects
Accidental Archives

Research · 05

Machine-Assisted Reading

How can humanists investigate evidence beyond individual inspection?

Search begins with something a researcher knows what and how to ask. Big data presents a different problem when potentially relevant evidence is dispersed across millions of texts, images, audio, video, structured records, and far larger quantities of unstructured binary data that cannot be inspected individually. The serendipity associated with walking through library stacks suggests an algorithmic analogue, where proximity, similarity, and difference can bring into view something the researcher did not know to seek. Machine-Assisted Reading investigates how computation can enlarge this exploratory field, making otherwise inaccessible relationships and candidates available for inquiry while keeping interpretation anchored in the evidence from which they emerge.

The research question guides how evidence is represented computationally, making representation itself part of the hermeneutic process. The same collection can be decomposed into passages for semantic retrieval, entities and relationships for graph analysis, vectors for textual or visual comparison, or clusters for examining recurrent patterns. Each representation makes some features legible and leaves others less visible; none stands for the collection as a whole. BDSL assembles semantic search, embeddings, graph decomposition, clustering, multimodal models, knowledge distillation, and other machine-learning operations according to the problem under investigation. Projects such as Accidental Archives place this approach under demanding conditions, where rescued websites, data dumps, leaks, and documentary disclosures combine enormous scale with uneven provenance and heterogeneous forms. Computation extends what can be observed and compared; researchers determine what those observations mean.

Provenance makes this movement across scale historically usable. A retrieved passage, cluster, graph, semantic neighbour, or machine-generated proposition acquires evidentiary value only when the researcher can establish how it was produced and inspect the records from which it derives. BDSL therefore develops research environments in which curated evidence remains distinguishable from the computational representations and models used to organize, retrieve, compare, or express it. Machine-Assisted Reading seeks an adjustable historical resolution, allowing researchers to move from individual records to patterns visible across much larger collections and then return to the particular evidence from which those patterns emerged. Computation enlarges the field of observation; interpretation remains accountable to the historical record.

The question determines the decomposition. The same collection can be reorganized for semantic, relational, visual, or multimodal comparison, with each representation revealing different candidates for investigation. Every route remains traceable to the evidence from which it was constructed.

Projects

Relevant Publications