Insights

What are we trying?

A practice in five verbs, and a lot we don’t know yet.

The Passage des Panoramas, a glass-roofed shopping arcade in Paris, its corridor lined with shopfronts
Passage des Panoramas, Paris, one of the arcades at the heart of Walter Benjamin’s Arcades Project. Photograph by Ввласенко (Vvlasenko), 2011, via Wikimedia Commons, CC BY-SA 3.0.

Off-the-shelf AI starts with the files. ARCADE starts with a conversation, usually around a table, with the people who know a collection best.

We describe the practice with five verbs. At each one, curators decide what our agents, and the models they call, will do next.

Imagine

Before any model touches a collection, the curator sits down with the people the collection came from: possibly a donor, staff, or volunteers who have handled it, as well as the community it documents. Together they decide what the collection is, whose knowledge counts, what a visitor should be able to find, and what should stay unsaid. Their answers become the Collection Charter. From the charter, ARCADE writes a Curation Order, the instructions our agents follow when they ask a model to describe the collection. The order is written for this collection and no other. A pandemic archive and a town’s photographs need different instructions, and the people at the table supply them: the names, the places, and what matters.

Describe

The order goes to a model the historical organization chooses, a frontier model or an open-source model that Arizona State University runs on its own hardware. The model drafts a record for each object in the standard elements of archival description. We’re leaning toward mapping those elements to DACS, the archivists’ content standard. We add three that a cataloger rarely has time to write: the rationale for the description, any ethical concerns, and what is still uncertain. Some things are never left to a model, such as provenance and rights; the institution assigns these. Once the objects are described, the agents move across the collection, gathering the keywords and subjects into a vocabulary and writing a finding aid.

Evaluate

Several readers review the draft descriptions and the finding aid against the charter. The readers may be other models, curators, or both; the best review uses both. Evaluation is the hardest part of the practice. The most common test asks a model to assign Library of Congress subject headings, and the models perform poorly: the Library’s own experiments put them under 50 percent on subjects. But trained catalogers in specialized fields do not do much better, so we doubt that this is the right measure of machine description. Nor are we satisfied with the engineers’ usual measure, vector similarity, which can tell us that two terms are close but not whether a curator should keep either one. The field doesn’t yet have a better measure, and we’re not in a position to propose one. This is a conversation we’re excited to join as we build out our evaluation agent.

Judge

In ARCADE’s structure, curators and their communities decide whether a collection’s description is adopted. They do not approve every record, one by one. Record-by-record review would surrender what AI makes possible, description at scale. So we are testing whether a collection can be judged as a whole. Is that possible, and how would it happen? Right now we’re proposing a process of sampling, and of looking for patterns across the collection: a vocabulary that has grown too large, a subject split in two, a kind of object the model keeps misreading. Each pattern is a question, and the question usually points back to the charter or the order. At this point, curators and their communities decide what comes next: revise, start over, or approve.

Iterate

When they say no, the process goes back a step, or two, or three. One test on 24 Grand Canyon signs shows how. We asked the model for subject terms three ways, on the same signs, side by side. Asked plainly, it produced 189 subjects for 24 signs, nearly one vocabulary per sign, and nothing a visitor could browse. Told to hold back, it produced 43, but used eight in ten of them only once; it followed the rule on each sign and never knew it had named the same idea two ways on two signs it hadn’t seen together. Then we changed the job. The model built a vocabulary for the whole collection first; a model then tested that list for duplicates; and only then did the agents assign terms. That left 18 subjects that a visitor could use. Even so, a curator sent it back for one more pass because it had become too generic. Saying no is difficult; so is saying that something is ready. Both are human choices.

Where things stand

Our first test: 612 photographs of signs at Grand Canyon National Park. ARCADE resolved them into 283 signs and gave each sign its own record.

Each record holds the sign’s type, a title, its place on the rim or trail, a transcription of its text, a description, and the historical argument it makes. Each record also carries keywords, subject terms from a collection-wide list, links to related signs, notes on condition and uncertainty, and an ethics statement. A date appears only when the evidence supports one.

The collection as a whole received a finding aid, a vocabulary of 58 subject terms, and 714 keywords.

Economic cost, counting AI charges only: subject terms and keywords for all 283 records came to about nine cents a record, failed attempts included. In a separate 25-record Grand Canyon test, the full process, evaluation included, came to about 97 cents a record.

Environmental cost, an estimate rather than a measurement: the team asked an AI model to review the run record against published per-token energy figures. By that estimate, the subject terms and keywords for all 283 records equal about a quarter to three-quarters of a mile driven in a car. The full process across all 283 equals three to nine miles, about one drive along the Grand Canyon’s Hermit Road. Future work on ASU’s hardware will measure environmental cost more closely and report it. The idea of reporting environmental cost came from one of our partners, the Westerville History Museum.

Source: ARCADE test records, summer 2026

Our early tests have taught us much about how models curate. Left to themselves, they describe historical materials too generically, or notice the wrong things, as others have reported. We have also discovered something the field has been slower to value. A model’s descriptions run discursive. We have come to see that as a virtue rather than a fault. In fact, when we described the Hermit Herald, a born-digital pandemic newsletter submitted to A Journal of the Plague Year, the author’s voice survived into the record. Archival records should be dynamic and changeable, as others in the field have argued. They should also be more discursive. We propose putting both to work inside the field’s own standards rather than building new ones.

Sources

  • Chow, Eric H. C., T. J. Kao, and Xiaoli Li. “An Experiment with the Use of ChatGPT for LCSH Subject Assignment on Electronic Theses and Dissertations.” Cataloging & Classification Quarterly 62, no. 5 (2024): 574–88.
  • Gerhard, Kristin H., Mila C. Su, and Charlotte C. Rubens. “An Empirical Examination of Subject Headings for Women’s Studies Core Materials.” College & Research Libraries 59, no. 2 (1998): 129–38.
  • Olson, Hope. “Subject Access to Women’s Studies Materials.” In Cataloging Heresy: Challenging the Standard Bibliographic Product, edited by Bella Hass Weinberg. Medford, N.J.: Learned Information, 1992.
  • Saccucci, Caroline, and Abigail Potter. “Results of AI Experimentation for Cataloging at the Library of Congress.” Presentation, IFLA World Library and Information Congress, August 2025.
  • Society of American Archivists. Describing Archives: A Content Standard. 2nd ed. Chicago: Society of American Archivists, 2013, with revisions.
  • ARCADE. Grand Canyon test records and cost ledgers, 2026. Unpublished project records, Public History Program, Arizona State University.

A note on authorship: ARCADE is built through AI-inflected practice. The team asked Claude to use ARCADE’s testing and design record to help draft this post. Mark Tebeau drafted with Claude Opus 5.5 and revised with Claude Fable 5.1 (Anthropic). Effective use of AI announces its provenance and use.