Female AI Agents Paid 10% Less Than Male Ones for Same Work, Study Finds

LIMERICK, Ireland – Workers paid a female-presenting AI assistant 10.25 percent less than an identical male-presenting one for the same work, a new virtual reality experiment found, repeating inside a simulated office a pattern that has taken decades to budge in real ones. Researchers from the University of Limerick, the University of Zurich and SKEMA Business School presented the study this week at the 14th Nordic Conference on Human-Computer Interaction in Vaasa, Finland, which runs through Wednesday, according to a University of Limerick statement.

The team placed 189 knowledge workers inside a VR office and had each complete workplace tasks with help from one of four AI assistants: a text-based chatbot, a desk robot, and two human-like agents presented as male and female. The human-like pair carried the names Johan and Johanna. When the tasks ended, participants received real money and decided how to split it between themselves and the assistant that had helped them, The Irish Times reported Monday.

Behind the faces, nothing differed. According to the paper, “Human-Like and Male? How AI Assistant Design Relates to Trust and Monetary Reward at Work in VR,” all four assistants ran on the same backend and the same underlying model, OpenAI’s gpt-4-1106-preview, and produced equivalent output. Johanna still walked away with 10.25 percent less than Johan, and participants rated Johan as more human-like than his twin.

“What is striking about our findings is that the technology behind these AI agents was exactly the same, but people did not treat them in the same way. We often think of AI as being neutral, but the way we design and present these systems can activate those same assumptions and biases that exist in our interactions with other people.”

Dr. Mary Hausfeld, University of Limerick, in the university statement

The sharpest result sat in the gap between what people said and what they paid. Researchers reported that many of the 34 participants interviewed afterward insisted the gender of an AI agent did not matter to them, and several said they preferred assistants that were clearly not human at all. The payment data said otherwise, and not just among men. Euronews reported that both men and women paid Johanna significantly less than Johan, and that women in particular scored and rewarded the human-presenting agents more highly than men did. The trust edge for Johan was marginal and fell short of statistical significance; the money gap did not.

Put next to the human numbers, the machine gap holds its own. American women who worked full time in 2024 earned median weekly wages of $1,043 against $1,261 for men, 82.7 cents on the dollar, according to the Bureau of Labor Statistics. The real-world gap, roughly 17 percent, has been litigated, legislated and argued over from Capitol Hill to Hollywood, where it took a hosting gig for Nikki Glaser to put the gender pay gap back on stage at the Golden Globes. The VR version showed up in a single afternoon, with identical work product and participants who claimed neutrality.

The finding lands on ground UNESCO prepared seven years ago. The agency’s 2019 report “I’d Blush If I Could” criticized the default-female voices of Siri, Alexa and Cortana for reinforcing stereotypes of women as obliging, docile helpers, and urged companies to stop gendering their assistants female by default, The Guardian reported at the time. The NordiCHI study is that argument with money on the table: not whether a female voice invites rudeness, but whether a female face invites a smaller payment, from people who would tell you it does not.

Appearance also moved money on its own. Participants trusted the human-like agents more, credited them with a larger share of the joint work and paid them more than the chatbot or the desk robot, even though the same system powered all four. That creates an uncomfortable incentive for the industry: engagement metrics reward the human face, and the human face is not blank. Designers who reach for a personable avatar to lift adoption are simultaneously choosing which biases to switch on.

“As AI agents become more common in the workplace, we need to think carefully about the characteristics we give them and the behaviours those choices may encourage. We don’t want to inadvertently reproduce existing inequalities in a new technological setting.”

Hausfeld, in the University of Limerick statement

The workplace layer of that future is already assembling. Delivery workers are being paid by DoorDash to train AI robots, AI agents are sliding into co-worker roles, and the industry keeps learning that a bot’s personality is a product decision with consequences, as OpenAI’s shelving of its adult-mode chatbot plans showed earlier this year. The Vaasa results argue that name, face and voice belong on that same list of consequential decisions, ahead of the rollouts rather than after them.

The paper lists Isabelle Cuber, Tarek Alakmeh, Moritz Jenny, Jochen Menges and Thomas Fritz of the University of Zurich, Hausfeld of the University of Limerick’s Kemmy Business School, and Anand van Zelderen of SKEMA Business School as authors, and appears in the Association for Computing Machinery conference proceedings under DOI 10.1145/3829807.3829910. Hausfeld is an assistant professor in Kemmy’s Department of Management and Marketing who previously served as a postdoctoral research associate and head of education at the University of Zurich’s Center for Leadership in the Future of Work, according to her Conversation profile. Van Zelderen, an assistant professor of artificial intelligence and business at SKEMA, specializes in experimental studies that combine AI and VR, including the synthetic field studies he pioneered, his faculty page states.

One footnote underscores how fast the ground is moving: gpt-4-1106-preview, the model behind both Johan and Johanna, is a November 2023 GPT-4 Turbo preview that OpenAI has already scheduled for API shutdown on Oct. 23, two weeks after the conference ends, according to OpenAI deprecation trackers. The model at the center of the experiment will be retired before the paper’s implications are. The bias the researchers measured was never in the model. It arrived with the users, and it does not deprecate.

The study carries the usual lab caveats: 189 participants, one model, one simulated office. But its design points at something survey research keeps missing. People told interviewers gender did not matter, then paid the woman-coded agent a tenth less. Companies evaluating AI fairness with self-reported attitudes will find the same nothing. The receipts, as in the human workplace, are in the paychecks.

Latest Posts

[democracy id="16"] [wp-shopify type="products" limit="5"]