Insights

Why can’t it be found?

Much of the world’s history sits in small institutions that can’t afford to describe it. In an age of AI search, what isn’t described can’t be found.

Bundles of police files tied with string, stacked from the floor nearly to the ceiling in the concrete rooms of the Historic Archives of the National Police in Guatemala City
Historic Archives of the National Police, Guatemala City, March 20, 2009. Photograph by James Rodríguez (MiMundo.org), licensed. Some 75 million pages of police records were found in a warehouse there in 2005; about 20 million have been digitized. But as Kirsten Weld writes, “digitization alone ensures neither public access nor survival for posterity.” (“No Democracy Without Archives,” Boston Review)

Walk into a small museum, a parish archive, or the back room of a town library almost anywhere in the world, and you’ll find the past in boxes. Photographs, letters, ledgers, maps, recordings. Someone kept them. Someone cared enough to keep them dry. Much of this material has never been described, and much of the rest only barely.

Kept is not the same as found

A collection moves through stages. It is preserved. It may be digitized. It may be described: given names, context, relationships, and an account of what remains uncertain. Only then is it discoverable, by people and by machines.

Most collections stall somewhere along the way, and not because anyone failed. Archivists learned long ago to describe collections as wholes rather than item by item, so that more of them could be opened at all. In 2005, the archivists Mark A. Greene and Dennis Meissner called this approach “More Product, Less Process.” A finding aid that says “Photographs, 1950–1990, three boxes” makes a collection usable. That finding aid doesn’t make the thousands of photographs inside the boxes findable. The photographs are both described and underdescribed.

Digitization isn’t the answer on its own

It’s tempting to think that putting everything online solves the problem. It doesn’t. A scanned photograph with a file name and nothing else is still dark. Once the photograph is on the open web, it may also be gathered up to train AI models, without permission and without anyone describing it. The photograph enters the great compendium of machine-readable knowledge, but not as itself. It is absorbed, not found.

Description is scarce

The historical record is abundant. The labor to describe it is not. Most of the world’s history organizations are small. Many run on little money. They don’t lack collections or care. They lack time.

UNESCO estimated that about 104,000 museums operated worldwide in 2020. Reliable global figures on their size are scarce; the United States keeps the most detailed count.

In the United States alone, 21,588 history organizations operate. More than 60 percent of the private nonprofit ones report annual revenue under $50,000. More than 80 percent report under $200,000.

21,588 dots, one for each history organization in the United States counted by the AASLH 2022 census, ordered from those with no revenue data through the smallest budgets to the largest. The largest block, 9,190 dots, is nonprofits reporting under $50,000 a year; 143 report $10 million or more.
Each dot is one of the 21,588 history organizations in the census, ordered from those with no revenue data, through the smallest budgets, to the largest. Among the 14,444 nonprofits with financial records, 9,190 report revenue under $50,000 a year; 143 report $10 million or more. Gray dots are organizations the census could not match to revenue figures, including government agencies and organizations housed within larger institutions. Source: AASLH, 2022 History Census, Table 1.

Sources: UNESCO, Museums around the World in the Face of COVID-19, 2020; AASLH, 2022 National Census of History Organizations

Search has changed, and it hasn’t

For generations, finding the past meant printed catalogs and browsing shelves. In the digital age, finding the past meant search engines, and the advanced search box on an archive’s website. Even then, much of the record was hard to discover. Material that hadn’t been digitized, or whose finding aid said little, stayed hidden. People asked questions and got lists and hints. Then they went to look, often in person, often with an archivist’s help.

That difficulty hasn’t changed. What AI changes is what comes back when you ask. Instead of a list of leads that invites you to keep looking, you get a finished answer that seems complete. The answer can draw only on what has been described somewhere. The risk isn’t that AI will invent the past. It’s that AI will keep finding the same past, and make it harder to notice what’s missing. Nothing is being destroyed. Much is being left out.

Why this is exciting

Archives and libraries have changed their tools before. In 1966, the Library of Congress began sending out catalog records that computers could read, in a format called MARC, developed under Henriette Avram. Over the decades that followed, MARC let libraries share their catalogs and changed how the world found books. But MARC taught machines to read descriptions. People still had to write them.

What’s new is that machines can now read the objects themselves: a photograph, a handwritten ledger, a faded sign. Machines can do that reading consistently and at scale, which is exactly the work there has never been time for. The same technology that threatens to leave this material out may also help bring it in.

Reading is not yet describing, though. A good description carries what a collection’s keepers know: who is in the picture, why it matters, what should stay private. The opportunity is to put the machines’ capacity in the service of the curators’ knowledge, not in its place.

AI can’t give curators more time. But it can do something else. AI can help them to describe at a scale not previously imagined, which answers the shortage of time another way.

The harder question is whether curators can describe at that scale on their own terms. That question is the heart of ARCADE. We think the practice we’re developing helps them do just that.

Sources

  • American Association for State and Local History. 2022 National Census of History Organizations. AASLH, 2022. https://aaslh.org/2022-census/.
  • Avram, Henriette D. The MARC Pilot Project: Final Report on a Project Sponsored by the Council on Library Resources, Inc. Washington, DC: Library of Congress, 1968.
  • Greene, Mark A., and Dennis Meissner. “More Product, Less Process: Revamping Traditional Archival Processing.” The American Archivist 68, no. 2 (2005): 208–63.
  • UNESCO. Museums around the World in the Face of COVID-19. Paris: UNESCO, 2020.
  • Weld, Kirsten. “No Democracy Without Archives.” Boston Review, July 9, 2020. https://www.bostonreview.net/articles/kirsten-weld-recovering-democracy-through-archives/.

A note on authorship: ARCADE is built through AI-inflected practice. The team asked Claude to use ARCADE’s testing and design record to draft this post. It was rewritten by Mark Tebeau and Claude Opus 5.5. Effective use of AI announces its provenance and use.