Insights

Why won’t AI fix it?

Off-the-shelf AI doesn’t ask what makes a collection unique: its history, its community, what needs care.

Duane Yarnell stands with one boot on the step of a Kubota tractor in a hayfield, a wagon of square bales behind him
Duane Yarnell haying at the corner of Africa and County Line roads, Westerville, Ohio. Photograph by Gary Gardiner, courtesy of the Westerville History Museum. When we asked AI systems to describe this photograph, most began with the tractor. We’ll say more in a coming post.

General-purpose AI is remarkable. Show a general model a photograph and the model can tell you a great deal: that the building is brick, that the photograph shows cars dating to the 1950s, that the sign in the window is in Portuguese. A model can transcribe handwriting that would take a person hours. For collections that have waited decades for anyone to look at them, that capability is real progress.

But naming what’s in a picture isn’t the same as archival description.

What a general model doesn’t know

Description depends on knowledge that lives in a museum’s files and a town’s memory. Who is in the photograph. Why the building mattered. What the street was called before it was renamed. What the community calls the event, and what people there would rather not see said.

A general model doesn’t have that local knowledge, and the model doesn’t ask for it. Asked about a small town, a general model may answer with the history of the nearest big city, because the city’s history is the one the model has read most.

Computer scientists have begun to map these gaps, and the map should interest every historian. Nikhil Kandpal’s team found that models know far less about rare things, and small collections are full of rare things. Yan Tao’s team found that models lean toward the values of English-speaking and Western European cultures, and Shibani Santurkar’s team found them leaning toward the elites within those cultures. A Google team led by Bahare Fatemi found that models stumble over time: over sequence, over duration, over what was true at a particular moment. For historical description, that last gap matters most. A photograph from 1952 has to be read in its own time, with today’s knowledge kept apart.

What a general model doesn’t ask

Curators ask questions of a record before they describe it. A general model, left to itself, asks none of them.

Who made this record, and for whom? Whose point of view does the record carry?

Whose names are here, and whose are missing? Names enter the record for good reasons and bad ones. A family is named in its portrait, while the woman who worked in its kitchen is not. A ledger names people as property. A newspaper prints a name because of an arrest, and the name follows that person forever.

Was this record meant to be seen? Private letters, sacred objects, and photographs of people who never agreed to be seen ask for care before they ask for description.

What did a word mean then, and what does the word do now? A term that was ordinary in 1920 may be a slur today, and still belong in the record as evidence.

What does the community call the people and places in the record, and what would the community rather leave unsaid?

What must never be guessed? A person’s identity from a face, an ancestry from appearance, a provenance from a hunch.

These are curatorial decisions, and every collection makes them differently. Off-the-shelf AI doesn’t know those decisions are there to be made.

Couldn’t we just give the model more context?

Yes, and giving context is the right instinct. Context changes what a model does. Tell a model what a collection holds, whose knowledge counts, and what deserves caution, and the model’s descriptions change. Tao’s team found the same with culture: naming a cultural setting in the instructions brings a model’s answers closer to that culture.

But context has limits, and the limits matter.

Context helps only so far. In Santurkar’s study, models stayed out of step with many groups’ views even after the researchers told them explicitly whose views to reflect.

Much of what makes a description good has never been written down. That knowledge lives in a curator’s memory, a donor’s stories, a neighbor who was there. Someone has to draw the knowledge out before any model can use it.

Choosing context is itself a curatorial act. Which sources count? Whose account comes first? An old catalog card may carry a mistake, or a term no one should repeat. Handing a model everything isn’t the same as handing the model the right things.

And someone has to check what the model did with the context it was given, and answer for the result.

Deciding, checking, and answering for the result isn’t prompting. That work is curatorial authority, written down so it can be shared, revised, and held to account. In ARCADE, that written authority is the charter, and the charter stays with the collection’s keepers.

The web is reading itself

There’s a wider problem, too. AI learns largely from the open web, and more and more of the web is now written by AI.

About half of new English-language articles on the web are now primarily written by AI, according to an industry study by the growth agency Graphite. It is not peer-reviewed, but its method is published and its figures have held steady across two rounds.

The sample came from Common Crawl, a web archive that is a key source for training AI models.

Source: Graphite, “AI Now Writes as Many Online Articles as Humans,” May 2026

Researchers have begun to show what can follow. Ilia Shumailov and colleagues showed in Nature that when models are trained on text that earlier models wrote, the rare and unusual parts of what they learned, the tails, begin to disappear. In a separate experiment, Anil Doshi and Oliver Hauser found that stories written with AI help were judged better one by one, but the stories resembled one another more than stories written without AI. Writers have named the fear more broadly. Ted Chiang called ChatGPT “a blurry JPEG of the web.” Kyle Chayka has argued that recommendation algorithms were flattening culture before generative AI arrived.

Of course, models don’t simply copy each other, and the field is working to address these problems. Nonetheless, this flattening is instructive for public historians, and the risk is clear. The well-described past gets described again and again, in more voices that sound alike. The rare and the local thin out. The same past is discovered and rediscovered. Even worse, the particular knowledge a small or obscure collection could offer isn’t there to counter the flattening, because those collections have never been made discoverable. That only compounds the problem.

What would help

Models are improving at remarkable speed, and their strengths are worth stating plainly. Models have an impressive range and depth of knowledge.

But models aren’t trained to be curators. Curation is a practice: habits of evidence, uncertainty, care, and authority, worked out over generations and applied collection by collection.

To do curatorial work, a model needs two things it doesn’t come with: a curator’s process, and a practice fitted to each collection.

Teaching models begins with the people who know a collection best: what the collection holds, what it needs, what it must protect. The practice continues by letting those people judge what the models return. ARCADE is building that practice.

Sources

  • Chayka, Kyle. Filterworld: How Algorithms Flattened Culture. New York: Doubleday, 2024.
  • Chiang, Ted. “ChatGPT Is a Blurry JPEG of the Web.” The New Yorker, February 9, 2023.
  • Doshi, Anil R., and Oliver P. Hauser. “Generative AI Enhances Individual Creativity but Reduces the Collective Diversity of Novel Content.” Science Advances 10, no. 28 (2024): eadn5290.
  • Fatemi, Bahare, Mehran Kazemi, Anton Tsitsulin, Karishma Malkan, Jinyeong Yim, John Palowitch, Sungyong Seo, Jonathan Halcrow, and Bryan Perozzi. “Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning.” arXiv:2406.09170, 2024. https://arxiv.org/abs/2406.09170.
  • Kandpal, Nikhil, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. “Large Language Models Struggle to Learn Long-Tail Knowledge.” In Proceedings of the 40th International Conference on Machine Learning, PMLR 202 (2023): 15696–707.
  • Paredes, Jose Luis, Gregory Druck, Bevin Benson, and Ethan Smith. “AI Now Writes as Many Online Articles as Humans.” Graphite, May 2026. https://graphite.io/five-percent/research/ai-now-writes-as-many-online-articles-as-humans-do.
  • Santurkar, Shibani, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. “Whose Opinions Do Language Models Reflect?” In Proceedings of the 40th International Conference on Machine Learning, PMLR 202 (2023): 29971–30004.
  • Shumailov, Ilia, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. “AI Models Collapse When Trained on Recursively Generated Data.” Nature 631 (2024): 755–59.
  • Tao, Yan, Olga Viberg, Ryan S. Baker, and René F. Kizilcec. “Cultural Bias and Cultural Alignment of Large Language Models.” PNAS Nexus 3, no. 9 (2024): pgae346.

A note on authorship: ARCADE is built through AI-inflected practice. The team asked Claude to use ARCADE’s testing and design record to draft this post. It was rewritten by Mark Tebeau and Claude Opus 5.5. Effective use of AI announces its provenance and use.