Skip to content

Learning

A Librarian’s Bot Steered Boys Away From Books About Feelings

Paired student profiles showed the same books falling down the list when the reader was marked as a boy. The finding changed how one school used personalized recommendations.

Devin OseiNarrator, Creating and Learning

October 11, 2026 · 8 min read

A laptop with book suggestions sits beside a printed spreadsheet on a school library desk.
A laptop with book suggestions sits beside a printed spreadsheet on a school library desk.

The librarian first noticed it at the circulation desk, where returned books waited in a cart and a student was trying to find something after finishing a fantasy series. He wanted a book about two friends who had stopped speaking. The school’s recommendation tool offered quests, survival stories, and several books with battles on their covers.

She changed the search terms. The list shifted, though the story she had in mind remained low enough that the student would have needed to keep scrolling. She handed him the book herself. He checked it out.

That might have been the end of it. Recommendation systems produce odd lists all the time, and librarians have long supplied the missing judgment with a conversation across a desk. But during the next two weeks, she saw a pattern: boys looking for stories about grief or friendship received lists weighted toward action, while girls with similar requests were offered quieter novels much sooner.

She opened a spreadsheet.

At first it was a place to record the obvious things: the profile setting, the prompt or selected topic, and the first page of suggested titles. By the end of six weeks, it held 48 paired tests. Each pair used the same invented reading history and the same stated interests. The only change was the gender marker attached to the profile.

One column measured rank movement. A novel about a child learning to talk with his father dropped 12 places when the profile changed from girl to boy. A story about a friendship ending fell nine. An adventure with a strained sibling relationship barely moved, perhaps because the system’s content tags included danger and travel alongside family conflict.

The librarian printed the spreadsheet and kept it beside the computer. Pencil marks appeared in the margins as she reran tests.

What the bot had learned

The tool did more than match a student’s search with subject labels. It ranked books using patterns gathered from profiles, clicks, saved titles, ratings, and checkout activity. Its documentation described a blend of content-based recommendation and collaborative filtering, which meant the system considered what a book was about while also learning from readers whose behavior looked similar.

That second part made the lists feel personal. A student who enjoyed one mystery could quickly find another without knowing an author’s name. Readers who struggled to describe what they wanted sometimes responded well to a cover in the suggested row. The librarian had watched reluctant students use the tool privately, without having to tell an adult that a book felt too hard or that they wanted something sad.

The same mechanism could also tighten a loop. If boys historically clicked action-heavy suggestions more often, the model could treat that behavior as evidence that similar boys preferred action. Those titles then appeared higher, where they were more likely to be clicked again. Stories centered on vulnerability received fewer opportunities to interrupt the pattern.

No engineer had typed a rule directing boys away from feelings. The model had learned a statistical association from ordinary activity, then used the association to order future choices. Gender was one signal, and reading history supplied others; even after the librarian removed the gender field in later tests, prior clicks sometimes carried enough information to produce a similar ranking.

This was the detail that troubled her. The system did not say a book was unsuitable. It placed the book lower.

A prohibition is visible. Ranking works by inches, and a child may never know that another version of the list existed.

She added another set of columns to the spreadsheet. For these tests, the profiles began with identical interests but developed different histories over several sessions. One clicked the first action title offered. The other skipped it and chose a family story farther down.

After repeated visits, the lists separated more sharply than they had at the start. The recommender was responding to behavior, but that behavior had partly been shaped by its earlier recommendations.

The librarian called this the nudge problem when she raised it with the school’s technology coordinator. A click looked like preference inside the system. From her side of the desk, it could also mean that the book had been placed where a student could see it.

A quiet test becomes a school question

The first meeting took place in the library after students had left. The printed spreadsheet lay between a laptop and a stack of returns. The coordinator asked whether the differences might come from random variation, so they repeated a selection of tests with fresh profiles. Some titles shifted by a place or two.

The broader pattern held.

The school asked the platform’s support team how strongly demographic fields affected ranking. The response described multiple signals and said recommendations could vary as the system learned. It did not provide the weight assigned to gender, which may not have existed as a stable percentage anyway; modern recommenders can combine signals through layers of learned relationships that are difficult to translate into a tidy rule.

That uncertainty did not make the test meaningless. The school knew what went into the profiles and could observe what came out. Across the paired rows, books centered on emotional conflict more often lost rank for boys, while action-led books tended to gain it.

There was disagreement about what to do next. Some teachers valued the system because students were checking out more books, and the librarian did too. Removing personalization altogether would also erase useful signals from children whose reading interests did not fit broad grade-level categories. A generic list could narrow choices in its own, familiar way.

The spreadsheet changed the conversation from whether the tool was biased to what kind of choice the school wanted it to support. A recommender optimized for likely clicks was doing what it had been built to do. The educators cared about likelihood, though they also cared about surprise: the book a student had never selected before, the subject he approached sideways, the story that gave him words he did not arrive with.

For the rest of the semester, the library stopped using gender as an input for new profiles. Staff also enabled a setting that mixed less predictable books into recommendation rows, and they placed librarian-selected titles beside personalized ones rather than letting the ranked list occupy the whole screen.

The changes did not produce equal lists. Reading history still created differences, as it should. Yet in the librarian’s follow-up tests, the steepest gender-linked rank drops became less common, and books about relationships appeared closer to the top for a wider range of profiles.

She kept testing.

One student returned the friendship novel she had handed him at the desk. He did not offer a review. He put it on the cart, said he wanted another book by the same writer, and waited while she checked the catalog.

The recommendation tool had become one source among several. Sometimes it found the next book faster than she could. Sometimes she ignored it.

Near the end of the school year, the librarian reopened the original spreadsheet. The 12-place drop was still there in an early row, a record of a list that no student had been told was different. Beside it, she had written two later results from profiles without the gender field. One showed a smaller drop.

The other showed none.

Questions people ask

How can a book recommendation system develop gender bias?

A recommender can learn associations from profile fields and past behavior. If boys have historically clicked certain genres more often, the system may rank those books higher for other boys, then treat the resulting clicks as fresh evidence. The bias can appear through ordering even when no title is blocked.

Does removing gender from a student profile fix the recommendations?

It reduced some differences in the librarian’s tests, but it did not erase them. Clicks, ratings, and checkout histories can carry patterns associated with gender or with expectations around gender, so profiles may still separate over time even when the explicit field is gone.

Why not turn personalization off completely?

The school found that personalized suggestions helped some students discover books and gave hesitant readers a private way to browse. Staff kept the tool but reduced its control over the screen, mixing exploratory titles and librarian selections with the ranked results.

Can a school test a recommender without seeing its model?

The librarian could not inspect the model’s internal weights, but paired profiles let her compare outputs while changing one input at a time. That showed a repeatable effect rather than its full cause. In the printed spreadsheet, the clearest early result remained the same: minus 12.

ShareFacebook
school and learningrelationshipsgender bias in recommendation systemsai recommendationsschool librariesgender biasstudent readingpersonalization

One story a day

The story of the day, in your inbox

One real story about AI each morning — no hype, no alarm, just company for the road.

Read next