Skip to content

Work

Software Gave Her Accent Low Warmth Scores and Cost Her $300

Customers rated the Florida call center agent highly. Voice software did not, and nobody supervising her could say which parts of her speech pulled the score down.

Mara QuinnNarrator, Work and Money

August 9, 2026 · 7 min read

A printed performance sheet beside a call center headset, with a low warmth score circled.
A printed performance sheet beside a call center headset, with a low warmth score circled.

The first proof was a weekly score sheet Camille printed at work in February 2024. Her customer satisfaction rate was 94 percent. Beside it, under a category labeled warmth, the number was 62.

Camille, a composite based on accounts from call center workers, had spent six years handling billing calls from her home in Florida. She worked for a contractor serving a large service company. Customers reached her after finding a duplicate charge, losing access to an account, or receiving a bill they did not understand.

Her quarterly bonus was $300. To qualify, agents needed an overall quality score of at least 80, and the new warmth measure counted heavily enough to pull her total down to 78. Human reviewers had previously listened to a small sample of calls. The voice system assessed nearly all of them.

She circled both numbers on the sheet: 94 and 62.

The gap bothered her more than the labels. Customers could leave a rating after a call, and their comments often praised her patience. A supervisor listening to one of the same calls marked the interaction as handled well. The software had assigned it a warmth score below 50.

Camille grew up in a Caribbean American family and had lived in South Florida since childhood. She did not think of herself as putting on an accent at work, though she knew her speech changed depending on whom she was talking to. With relatives, her pace picked up and the stress moved across words. On customer calls, she slowed down and used the phrases taught in training.

The score suggested that was not enough.

What warmth measured

The software did not hear warmth as a person would. It separated Camille’s voice from the customer’s voice, transcribed the call, then measured patterns in the audio and words. Materials available to supervisors described signals such as speaking rate, pitch movement, vocal energy, pauses, and overlap between speakers. The transcript could also be checked for acknowledgment phrases or language associated with reassurance.

Those signals became a number.

Some parts were easy to understand. If Camille repeatedly spoke while a customer was speaking, the system could mark excessive overlap. Long stretches without either person talking could affect another measure. The tool also highlighted places where its transcript did not detect an apology or acknowledgment after a customer described a problem.

Warmth was harder. A lower score might reflect a voice with less pitch variation, a pace outside the model’s preferred range, or pauses that fell in places the system treated as awkward. It could also reflect errors earlier in the chain. If the speech-recognition component misheard a phrase, the language analysis would work from the wrong words.

Accent matters in both stages. Speech-recognition models tend to perform unevenly when a speaker’s pronunciation or rhythm is less represented in training data. A model built to infer emotion faces another problem: vocal patterns do not carry one fixed meaning across regions, families, and speakers. A firm tone can be attentive.

A rising pitch can be habitual rather than encouraging. Silence may mean someone is checking a bill.

Camille’s supervisors could see her score and the highlighted portions of a call. They could not see the weight assigned to each signal, the comparison group used to set the standard, or the training recordings that had taught the model what warmth sounded like.

When she asked why the call with positive customer feedback had scored below 50, her supervisor played it again. Camille explained the charge, waited while the customer found a statement, and confirmed that a credit would appear. The customer’s written response gave the interaction the highest available rating.

The supervisor heard no clear problem. The dashboard still showed the low score.

The question went to the platform’s support team through the company. The response described broad testing across different speakers and pointed managers back to the general categories shown on the dashboard. It did not identify the vocal trait that had lowered Camille’s score, nor did it explain how much any single trait counted.

Her supervisor could coach her on the label. He could not explain the calculation underneath it.

Changing her voice

Camille began writing notes on the printed score sheet. Beside calls with low warmth numbers, she marked whether she had used what she called her home voice or her training voice. The distinction was hers, not the software’s.

For six weeks, she made the training voice more pronounced. She left longer gaps after customers stopped speaking, raised her pitch during reassurance phrases, and slowed certain sentences even when the customer had already understood the answer. She repeated acknowledgment language that sounded natural once but strained when used throughout a call.

Her average warmth score rose from 62 to 81.

The customer rating barely moved. It stayed between 93 and 95 percent, which meant the people on the calls had not registered a comparable change in service. Camille had changed the signals available to the model, and the model’s judgment changed with them.

That result did not prove that the system penalized every Caribbean American speaker, or that accent alone caused her earlier scores. The company had no analysis it could show her. Yet the pattern was strong enough to alter how she worked: the farther she moved from her ordinary rhythm, the safer her pay became.

There was one part of the tool she found useful. Its overlap markers showed that she sometimes began explaining a solution before a customer had finished describing the problem. Human reviewers had missed the pattern because they heard only a few calls each month. Camille adjusted, and customers interrupted her less often.

The warmth coaching felt different because it required her to perform a quality that customers already believed she had. A call could end with the billing problem fixed and a positive survey, while the system treated her pitch or pacing as evidence that something human was missing.

Her supervisor eventually removed two unusually low calls from a coaching review, but he could not erase the warmth average that had already affected the quarter. Payroll closed without the $300 bonus.

Camille qualified the next quarter after keeping the slower training voice. The money appeared in her check. She did not recover the first $300, and she never received a technical account of why 62 had become 81.

The original score sheet stayed in her work folder. The circles around 94 and 62 overlapped where the pen had passed twice.

Questions people ask

Can software really score a worker’s empathy?

Voice-analysis systems can assign scores to behavior associated with empathy, including pace, pitch movement, pauses, overlap, and words detected in a transcript. They do not measure an inner feeling. Camille’s warmth score was an inference drawn from those signals, even when the customer’s own rating pointed in another direction.

Why can an accent affect an automated warmth score?

Accent can affect the transcript if speech recognition mishears words, and it can affect audio measurements because rhythm, stress, and pitch vary among speakers. If a model learned its standard from a limited range of voices, ordinary differences may sit farther from the pattern it rewards. Camille was never shown the model’s training data or comparison group.

Can a supervisor explain why one call received a low score?

Camille’s supervisor could replay highlighted moments and review broad categories such as overlap or pace. He could not see the model’s feature weights or explain why a specific pitch pattern lowered warmth. The support response offered general descriptions, leaving the person responsible for coaching her without a clear account of the score.

What happened to the agent’s pay?

The warmth score pulled Camille’s overall quality result below the threshold for a $300 quarterly bonus. She qualified during the next quarter after changing her delivery, but the earlier payment was not restored. The first sheet remained in her folder, with 62 circled beside a 94 percent customer rating.

ShareFacebook
workmoneycall centersworker surveillancevoice analysisalgorithmic bias

One story a day

The story of the day, in your inbox

One real story about AI each morning — no hype, no alarm, just company for the road.

Read next