AI Giants Overlook History at Their Peril

Benjamin Breen tested the newest models on stubborn puzzles from centuries past. He fed GPT-6 Sol and Opus 5.5 fragments of 17th-century letters and alchemical texts. The results surprised him. What once required years of archival digging now yielded fresh connections in hours.

Breen, a historian of science, laid out his findings yesterday in a Substack essay. He argued major AI developers must begin underwriting historical scholarship. Not as charity. As strategic necessity.

Frontier models already draw strength from digitized primary sources. More of those sources, cleaned and accessible, would sharpen their reasoning. Historians gain new tools to crack open forgotten debates. Both sides win. Yet the labs have mostly stood apart.

That separation narrowed this week. Schmidt Sciences, founded by Eric and Wendy Schmidt, announced $11 million for up to 23 research teams applying AI to archaeology, history, literature and allied fields. Wendy Schmidt said the effort would “shed light on our oldest truths, on all that makes us human—from the origins of civilization to the peaks of philosophical thought.” The grants arrived just as Breen’s piece circulated.

Google DeepMind moved earlier. In 2025 it released Aeneas, a model built to restore, date and contextualize fragmentary ancient inscriptions. Created with the University of Nottingham and partners at Oxford, Warwick and Athens, the system compares weathered texts against vast networks of known examples. Historians who tested it reported meaningful gains in accuracy and speed. DeepMind made an interactive version available at predictingthepast.com.

These projects hint at momentum. They remain scattered. Breen wants systematic investment. He proposes three concrete steps.

First, collaborate with libraries and archives to digitize materials still locked away. In his experiments the biggest barrier was access, not model intelligence. Second, fund teams of historians working alongside AI researchers on specific, tractable problems. Mathematics showed the way; targeted challenges produce rapid progress. Third, treat historical data as core training material rather than afterthought.

Breen’s own tests illustrate the potential. Using the latest models he traced threads of alchemical knowledge across decades. He decoded cryptic correspondence that had resisted conventional analysis. The models did not replace expertise. They amplified it. Historians set the questions, verified outputs and spotted nuances machines missed.

Others have reached similar conclusions. A New York Times magazine examination last year described historians loading sources into tools like NotebookLM to surface patterns across thousands of documents. One researcher sifted fur-trade records from early Canada. Another mapped connections in 1970s New York art circles. The technology did not invent interpretations. It made previously impractical questions realistic.

Paul Dilley at the University of Iowa received $500,000 from Schmidt Sciences to build AI that reconstructs damaged ancient manuscripts. His team targets Coptic codices and Herculaneum scrolls. The goal is software that fills gaps while letting scholars steer the process. Todd Presner at UCLA leads a $308,000 project examining how generative models handle Holocaust history, both as aid and source of distortion. These grants show foundations stepping where labs have not.

Stanford historian Giovanna Ceserani directs another funded effort. Her team designs AI architectures modeled on actual historian workflows. Test cases include 19th-century Senegalese slave-liberation records, 18th-century Ottoman probate inventories and Venetian port logs. The approach rejects generic chatbots in favor of systems shaped by disciplinary practice.

Yet obstacles persist. Many archives hold material only in physical form or low-quality scans. Models trained predominantly on modern English struggle with shorthand, regional dialects or multilingual texts. Ambiguity that historians navigate instinctively can derail automated analysis. And then there is the deeper risk: models that confidently fabricate details when sources run thin.

Breen acknowledges the pitfalls. He has written before about them. What changed, he says, is model capability. GPT-6 and Opus 5.5 crossed a threshold. They handle complex historical inference better than predecessors. The bottleneck shifted from raw intelligence to data and collaboration.

Private labs possess resources to move the field. They command talent, compute and cash. They also stand to benefit. Better historical grounding could reduce hallucinations on factual queries. It might improve performance on long-context reasoning. And it would supply the training signal that public data alone cannot match.

Some efforts already point the direction. The AI Historian project at Oxford, backed by Schmidt Sciences’ AI2050 initiative, aims to build an open-source AI collaborator that mirrors historian methods. Researchers there work directly with domain experts. A separate group at Zurich trains time-locked models on historical texts only, creating windows into specific eras without later contamination.

DeepMind’s Aeneas demonstrated another path. Twenty-three historians evaluated the model on real tasks. The system helped restore texts, assign dates and identify origins. Its creators emphasized partnership. The model augments human judgment rather than replacing it.

But scale remains limited. Most funding still flows through philanthropic channels rather than the balance sheets of OpenAI, Anthropic or Google. Breen believes that must change. AI labs should treat historical research as infrastructure. Digitization campaigns, shared benchmarks, joint workshops. The investment would pay dividends in model quality and in genuine advances in human knowledge.

History offers something models desperately need: dense, contextual, contested data. It rewards careful cross-referencing and source criticism. These habits, if built into AI development, could produce systems less prone to confident error. Historians, meanwhile, gain computational power to tackle questions once beyond any single career.

The timing feels urgent. New models arrive monthly. Archives continue to deteriorate. Without deliberate coordination the opportunity narrows. Breen’s call is measured. He does not promise miracles. He points to concrete experiments that already work and asks for structured support to expand them.

Schmidt Sciences’ latest round suggests some in Silicon Valley hear the message. Their $11 million will test dozens of approaches across geographies and eras. Results will inform what comes next. Yet the major labs retain unmatched capacity to set standards and accelerate progress.

Whether they choose to engage will shape more than academic output. It will influence what future models understand about how societies change, how knowledge travels and how humans grapple with their past. Ignore history and the machines will inherit its gaps. Fund it deliberately and both fields stand to gain. The choice, for now, rests with the labs.


Discover more from Web and IT News

Subscribe to get the latest posts sent to your email.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top

Discover more from Web and IT News

Subscribe now to keep reading and get access to the full archive.

Continue reading