A Museum Hired Her to Update Labels. AI Put Them Back
A collections assistant was hired to bring museum records into the present. A chatbot kept writing about her living community in the past tense.
August 9, 2026 · 8 min read

The first line that stopped Avery was attached to a woven basket.
The museum’s old catalog record was sparse. It listed a material, an estimated date and the name collectors had used for the community more than a century earlier. The new chatbot turned those fragments into a paragraph about a vanished people who had once used baskets in daily life.
Avery’s relatives still made them.
She opened a spreadsheet and pasted in the generated sentence. Beside it, she entered the object name, the source record and a short note: living community described in the past tense. It was the first row in what became her correction log.
Avery was 26 and earning $21.50 an hour on an 18-month grant. The museum had hired her to revise collection descriptions that ranged from a single noun to paragraphs copied from accession cards. Some records used names the community no longer used.
Others repeated guesses made by collectors who had never asked the makers what an object was for.
The work had been expected to move slowly. Avery would read each record, check related documents, speak with community members when an object required more context and write a new description. Then the museum added a chatbot to the project.
The tool could produce a complete label in seconds. A manager estimated that a worker who had been revising about 12 records a day could review 40 generated drafts instead. The grant covered thousands of records, and the difference looked useful on a planning sheet.
It changed Avery’s job before a single generated label reached the public.
Where the extra language came from
The chatbot did not begin with an empty page. Museum staff fed it text from the existing catalog record and asked for a short description in accessible language. The model then predicted a likely sequence of words based on patterns learned from a much larger body of writing, while using the old record as the immediate source for names, dates and materials.
That combination mattered. Many museum records were written in an ethnographic style that placed Native communities in a fixed past, even when no sentence directly claimed that a community had disappeared. A collection date from 1912, a verb such as “used,” and a broad category such as “ceremonial” gave the model a familiar pattern. It completed that pattern with confident language about ancestral customs, traditional beliefs or former ways of life.
Those additions were not retrieved facts. They were plausible continuations.
The chatbot also made the old prose easier to read, which made its unsupported claims harder to notice. A fragment on an accession card looked incomplete. A generated paragraph had a beginning and an end. When it added a purpose that no source confirmed, the claim arrived in the same steady voice as the object’s measured dimensions.
Avery could often trace a bad draft back to one loaded word. “Primitive” led the model toward simple tools and unchanged customs. “Costume” encouraged descriptions of performance, even when the record did not establish how a garment had been worn. If an old card said an item “was used,” the generated version rarely paused over whether people still used it.
The basket entry was not an isolated mistake. After six weeks, Avery’s spreadsheet held 143 rows. Fifty-eight involved language that pushed a living community into the past. Another 31 contained uses, meanings or ceremonial claims that Avery could not find in the museum’s records.
She was still expected to meet the faster review target.
Editing became investigation
Before the chatbot, Avery could see where information ran out. A short record made its own limits visible, and she could decide whether to research it or leave the uncertainty in place. Generated prose removed those edges, so she had to check every noun that sounded more specific than the source and every verb that implied a custom had ended.
This was different from correcting spelling. To reject a sentence about how a bowl had been used, she might read the accession card, compare related objects and ask a cultural adviser whether the claim belonged in a public record at all. The draft took seconds to produce, but its review could take longer than writing a modest description from the source material.
The correction log became the clearest account of her work. The museum’s project dashboard counted completed records, while Avery’s spreadsheet counted generated claims that should never become labels. Those measures moved in opposite directions. On weeks when she found more errors, her output appeared to fall.
She brought the basket example and several related rows to her manager. The concern was not limited to tone, she explained. Changing “is used” to “was used” made a factual claim about cultural continuity. Expanding “ceremonial” into a description of a sacred purpose could expose information the museum did not possess or did not have authority to publish.
Her shortest note in the spreadsheet was also the one she repeated most often.
We are still here.
The museum paused automatic publication, though it kept the chatbot in the drafting process. Avery and another worker began marking which parts of each draft came directly from a catalog record and which had been generated. Staff also changed the prompt, telling the system to avoid claims about cultural meaning, use present tense for living communities and state when a purpose was unknown.
The revisions helped. They did not settle the issue.
A prompt can steer a language model, but it does not erase the patterns in its training data or repair the source text placed in front of it. The chatbot sometimes followed the new instruction for several paragraphs, then returned to phrases associated with historical museum writing. It also replaced blunt outdated terms with smoother language while keeping the same underlying assumption, which meant a draft could pass a quick tone check and remain wrong.
The useful part was narrower
Avery did not want the tool removed from every task. When a record had already been verified, the chatbot could shorten a long sentence or suggest a version for visitors who did not know museum vocabulary. It helped her test whether a description made sense without specialist terms. She sometimes kept a phrase, though never a full draft.
That use was slower than the original plan and more useful than she had expected.
The tool gave her options on days when she had read the same paragraph too many times to hear it clearly. It could also expose a problem in the source record by making an implied assumption explicit. A two-word note that seemed merely dated could become a full generated sentence about a community no longer existing, and the expansion showed what the older language had been carrying all along.
Still, the creative choice belonged in the checking. Avery decided what could be said, what needed attribution and what should remain undescribed. The chatbot’s role narrowed from writer to provisional rephraser.
Eight months into the grant, the spreadsheet had 418 rows. It was no longer treated as a private list of mistakes. Staff used it during reviews to identify recurring failures, and records involving cultural meaning were routed for community input rather than approved from a generated draft alone.
The change reduced the daily target. It also changed what counted as finishing a label. A completed record now required a person to distinguish copied fact from generated connective language, an awkward category that included ordinary words such as “therefore,” “traditionally” and “once.” Those words could turn two facts into a history the museum had no basis to tell.
Near the end of her grant, Avery returned to the basket record. The revised description named its materials and approximate collection period. It noted that the maker was not recorded. A separate field used the community’s current name, and the public text made no claim about disappearance or ceremonial purpose.
In her spreadsheet, she changed the row from open to reviewed.
Questions people ask
Why did the chatbot describe a living community in the past tense?
The model received old catalog language and completed it using patterns common in historical and museum writing. Dates, outdated community names and verbs such as “was used” pushed its wording toward the past. It did not verify whether the community still existed before producing a fluent description.
Why didn’t a better prompt fix the museum labels?
The revised prompt reduced some errors, but the model still worked from records shaped by older assumptions and from broader language patterns that often frame Native life as historical. Instructions could guide the draft without proving its claims, so a person still had to check tense, purpose and cultural meaning.
Did the
AI tool save the collections assistant time?
It saved time on narrow rewriting tasks after the underlying record had been verified. On uncertain records, it often created more work because Avery had to separate sourced facts from plausible additions. The museum lowered its daily target once review showed that a polished paragraph could take longer to verify than a short description took to write.
Who had authority to approve the final description?
At first, the project treated generated text as a draft that a collections worker could edit. After the correction log showed recurring cultural claims, the museum required community input for sensitive records and kept human approval before publication. For the basket, the final record named the materials and collection period while the spreadsheet remained open beside it.
One story a day
The story of the day, in your inbox
One real story about AI each morning — no hype, no alarm, just company for the road.



