An AI Report Counted 24 Training Hours Against a Transit Worker
The worker’s score fell to 73 percent after approved training vanished from the denominator. A union steward found the change by comparing two printed reports.
September 16, 2026 · 7 min read

The union steward noticed the missing hours on paper.
Two productivity reports covered the same four-week period. The first, attached to a disciplinary notice, listed 160 available hours and 116.8 earned hours. It gave the worker a productivity score of 73 percent, below the employer’s 85 percent standard.
The second report had been generated six weeks later, after the worker appealed. It showed 136 available hours, 123.8 earned hours and a score of 91 percent. A row labeled training now contained 24 hours.
On the first report, that row showed zero.
The worker had spent those 24 hours in an approved course on maintaining a newer type of transit vehicle. Attendance was recorded. The worker had been paid for the time and had not been assigned regular repair orders during it.
The first report treated every one of those hours as time in which maintenance work should have been completed. That enlarged the denominator used to calculate productivity. It also placed the worker low enough on a ranked list to trigger review by a supervisor.
The steward put the reports next to each other and circled the training row.
How the score was built
The employer had once compared mechanics using closed work orders and paid hours. That measure was crude. Replacing a damaged wiring harness could take most of a shift, while completing several inspections might take less time. A count of finished orders favored workers assigned short, repeatable jobs.
The newer system tried to account for that difference. A machine-learning model estimated how much productive labor each completed task represented, based on past work orders involving similar vehicles and repairs. Those estimates became “earned hours.” A task that the model expected to take three hours could give a worker three earned hours even if the worker finished sooner or later.
This was one reason some workers did not reject the system outright. Different jobs carried different weights, and the weights changed as the model absorbed more records. A long repair could receive more credit than several quick checks. For a maintenance shop with mixed assignments, that could be fairer than counting jobs.
The score still depended on a second calculation. The system divided earned hours by hours considered available for maintenance. Vacation, approved leave and training were supposed to be removed from that denominator after data arrived from the employer’s timekeeping and learning systems.
In this case, the course used a training category that had been introduced with the new vehicle program. The automated data process did not recognize it as excluded time. The 24 hours remained in payroll records as paid hours, but the category disappeared when the dashboard normalized the records into its smaller set of labels. The scoring system classified them as available maintenance time.
That was more than a missing cell. The model had already assigned predicted labor values to the worker’s completed tasks, while another automated layer had decided which paid hours belonged in the comparison. The report then ranked the result against the employer’s threshold and sent the low score forward for possible discipline.
A spreadsheet could divide 116.8 by 160 and produce 73 percent. It could not independently estimate the labor value of each repair from historical patterns, revise those estimates after a model update and combine them with an automated classification of time records. Without those steps, this particular score and the ranked flag would not exist.
Why the second report did not settle it
The employer withdrew the citation after the second report showed 91 percent. For the worker, that removed the immediate risk of a formal mark on the employment record.
The steward did not consider the matter closed. Correcting the training category explained why available hours fell from 160 to 136. It did not explain why earned hours rose from 116.8 to 123.
8 even though both reports covered the same completed jobs.
The answer appeared to be a model update. Between the two reports, the system had refreshed the expected labor values assigned to certain repairs. Work on the newer vehicles now carried more earned time, likely because the platform had absorbed additional examples showing that those jobs took longer than its earlier estimates.
That feature was part of the system’s purpose. The task weights were not fixed. Yet it also meant a regenerated report was not necessarily a copy of the report used when discipline began. The same worker, hours and repair orders could produce a different score after the model changed.
The dashboard did not display the earlier task weights. It showed the current score and a summary of the current categories. A manager could download the new report, but could not use that screen to reconstruct the old one.
The two paper reports became evidence of something the live dashboard no longer showed: the system had once treated the training hours differently, and it had valued the repair work differently too.
The appeal moved below the dashboard
The union asked for the records used to produce the first score. That included the time entries before categories were converted, the task values in effect when the report ran and information showing which model version generated each result.
Management provided attendance records confirming the course and a current export of the worker’s repair orders. It did not provide a historical model snapshot or the intermediate table that joined the time records to the scoring system. The employer said some of that material was controlled by the platform provider, and some was not available through the manager’s dashboard.
This changed the dispute. The missing training hours were no longer the only issue. The union wanted to know whether a worker could inspect the inputs and task weights that had supported discipline at the moment the decision was made, rather than a corrected set produced later.
The platform’s summary offered a clean percentage, but the percentage combined records from separate systems with predictions that could be revised. Once the original model values were replaced, neither the worker nor the local manager could reproduce 73 percent from the dashboard.
The steward’s concern extended beyond this appeal. If a training category could be mapped incorrectly, other paid activities might also enter the denominator. If task weights changed after a report was generated, a worker challenging a score could face a moving record unless the employer retained the version used for the decision.
The worker returned to regular assignments while the union continued seeking the historical material. There was no second citation. The training program continued, and later reports listed those hours separately.
The original report still mattered. At the top was 73 percent. Farther down, beside training, was zero.
Questions people ask
How can approved training lower an AI productivity score?
In this case, the scoring system divided model-assigned credit for completed repairs by hours classified as available for maintenance. A new training category was not recognized as excluded time, so 24 hours entered the denominator even though the worker had no regular repair assignments during them.
Why did the same work period produce two different scores?
The second report corrected the training category, reducing available hours from 160 to 136. The machine-learning model had also updated the expected labor values attached to some repairs, raising earned hours from 116.8 to 123.8.
Both changes affected the score.
Could a manager explain how the first score was calculated?
The manager could see the summary report and current work-order export, but the dashboard did not retain the earlier task weights or show the intermediate mapping of time categories. Reproducing the first score required historical records held below the visible dashboard, some of which the employer said it could not readily retrieve.
What records mattered during the appeal?
The attendance record established that the worker completed 24 hours of approved training. The unresolved issue was access to the original time classification, model version and task values used for discipline. The steward kept the two printed reports clipped together, with the zero beside training circled.
One story a day
The story of the day, in your inbox
One real story about AI each morning — no hype, no alarm, just company for the road.



