How can humanists investigate evidence beyond individual inspection?
Historical research has long depended on instruments that make large bodies of evidence navigable. Catalogues, concordances, indexes, databases, and keyword search help historians find what they know enough to seek. Walking through the stacks offers a different kind of discovery, where an unfamiliar title or unexpected neighbour can suggest a question that no catalogue search would have anticipated. Big data makes such exploratory reading increasingly difficult when collections extend across millions of texts, images, audio, video, structured records, and other digital objects. Machine-Assisted Reading asks how computation can recover some of this exploratory freedom at scales no individual researcher can survey, while preserving the close attention to sources on which historical interpretation depends. The objective is neither distant reading nor the automation of close reading, but an expanded field of comparison in which historians can move between individual records and patterns visible only at larger scales.
The historical question determines how that movement proceeds. The same collection may be represented as passages for semantic search, entities and relationships for graph analysis, vectors for textual or visual comparison, or clusters and distilled propositions for examining patterns across a corpus. Each representation makes certain comparisons possible while leaving others obscure, and none provides a definitive account of the collection. Semantic search, embeddings, graph decomposition, clustering, multimodal models, and knowledge distillation can identify similarities, structures, anomalies, and candidates for closer investigation. Their value lies in making comparisons observable that would otherwise remain beyond practical inspection; their historical significance emerges only when researchers examine those patterns against sources and context. The same principle extends earlier BDSL work on digital rereading, where graph representation allowed movement between individual records and larger historical formations, into semantic, stylistic, visual, and multimodal comparison.
Provenance makes this movement across scale historically usable. A retrieved passage, cluster, graph, semantic neighbour, or machine-generated proposition acquires evidentiary value only when the researcher can establish how it was produced and inspect the records from which it derives. BDSL therefore develops research environments in which curated evidence remains distinguishable from the computational representations and models used to organize, retrieve, compare, or express it. Machine-Assisted Reading seeks an adjustable historical resolution, allowing researchers to move from individual records to patterns visible across much larger collections and then return to the particular evidence from which those patterns emerged. Computation enlarges the field of observation; interpretation remains accountable to the historical record.
The question determines the decomposition. The same collection can be reorganized for semantic, relational, visual, or multimodal comparison, with each representation revealing different candidates for investigation. Every route remains traceable to the evidence from which it was constructed.
Examines large digital collections that became available as historical evidence by routes their creators rarely intended, from rescued web archives and data dumps to materials disclosed by leaks, hacks, and legal proceedings. The project asks how their histories of survival, custody, and access shape what historians can know from them.
Traces the history of personalization from Web 2.0 recommendation and engagement systems to generative AI, reconstructing how digital systems have learned to observe, classify, predict, and act on their users.