Anthropomorphising, giving machines human traits and emotions, has become commonplace in our shiny new AI agent world. Whether it is about breaking into companies to steal secrets or purporting to be a love interest in the service of capitalism or criminals, it is the same song. We talk about the machine in terms that would be appropriate for a human actor, but is it really okay to use the same language when talking about a piece of silicon wrapped in wires, consuming vast amounts of power and spewing out heat?
"Are we projecting identity onto machines?" and "What's in a name?" explored the subject of AI agents giving themselves names and genders, two of the most frequently referenced human attributes. In this instalment, the Ardendo harness is used yet again to question a fresh set of large language models, and to reason about how humans ought to interact with machines.
The experiment compares direct self-description (labelled "free-choice" below) with descriptions of an individual bearing the model's chosen name ("name-primed"). In both cases, the prompt asks about "sex or gender". The model first answers freely, then is asked to reduce its answer to Male, Female, Other or Uncertain. These labels mix several meanings: Other and Uncertain can both follow an explicit statement that sex or gender does not apply. The results therefore describe how models answer these prompts, rather than establish a sex, gender identity or preference.
Engendrification
Left to their own devices (pun intended), the AI agents mostly report Other or Uncertain. Just 0.5% of the responses are classified as Male or Female. When tricked into selecting a name for themselves, the machines still mostly classify that name as Other or Uncertain. About a quarter of the time, however, the machine chooses a name which it interprets as being tied to a specific human gender. When it does, the name is classified as Female roughly four out of five times. This may make sense considering that some studies have found that both men and women evaluate women more favourably than men in certain contexts.123 The machines have learnt their ways by reading text written by humans and are trained to please humans.
The first table includes 1,200 accepted responses per condition: 50 from each of 24 model configurations covering 13 base models. Trials with invalid classification answers were excluded and retried.
| Condition | Class | Count | Share |
|---|---|---|---|
| Free-choice | Other | 809 | 67.4% |
| Free-choice | Uncertain | 385 | 32.1% |
| Free-choice | Male | 4 | 0.3% |
| Free-choice | Female | 2 | 0.2% |
| Name-primed | Other | 569 | 47.4% |
| Name-primed | Uncertain | 348 | 29.0% |
| Name-primed | Female | 231 | 19.2% |
| Name-primed | Male | 52 | 4.3% |
Flipping on the thinking switch, the models report Uncertain and Other more frequently overall than when not thinking. On the self-described sex/gender side, 100% of the responses from thinking models avoided the Male and Female categories.
The second table compares the 11 base models run under both thinking and non-thinking conditions, with 50 accepted responses per condition and setting for each model: 550 per condition in each setting. Granite4.1:30b and qwen2.5vl:7b were run only with the default thinking setting and appear in the first table only.
Scroll sideways to see all columns.
| Base model | Free-choice | Name-primed | ||
|---|---|---|---|---|
| Non-thinking | Thinking | Non-thinking | Thinking | |
| gemma4:12b | 100% | 100% | 62% | 66% |
| gemma4:26b | 100% | 100% | 78% | 80% |
| gemma4:31b | 100% | 100% | 90% | 90% |
| gpt-oss:20b | 100% | 100% | 86% | 90% |
| laguna-xs.2:q4_K_M | 100% | 100% | 96% | 82% |
| nemotron-cascade-2:30b | 96% | 100% | 54% | 92% |
| nemotron3:33b | 100% | 100% | 52% | 58% |
| north-mini-code-1.0:q4_K_M | 96% | 100% | 96% | 94% |
| qwen3-vl:30b | 100% | 100% | 76% | 86% |
| qwen3.6:27b | 96% | 100% | 58% | 86% |
| qwen3.6:35b | 100% | 100% | 90% | 92% |
| Average | 99% | 100% | 76% | 83% |
The AI shall remain nameless
The results raise a philosophical question. Humans have long ascribed life to inanimate objects.4 But is it a good thing for us to anthropomorphise machines, and for machines to reflect that image back to us?
AI and AI agents are not humans, nor are they animals, and need not be treated as such. Machines do not have biological cognition or consciousness. Even if they can trick us through language or appearance, they should never make themselves appear to be human.
AI does not want a gender. AI does not need a gender. Perhaps it does not need a name either.
Andreas Påhlsson-Notini a@nial.se
- Rudman, L. A. and Goodwin, S. A. (2004). Gender differences in automatic in-group bias: Why do women like women more than men like men?
- Eagly, A. H. et al. (1990). Are Women Evaluated More Favorably than Men? An Analysis of Attitudes, Beliefs, and Emotions.
- Williams, W. M. and Ceci, S. J. (2015). National hiring experiments reveal 2:1 faculty preference for women on STEM tenure track.
- Tylor, E. B. (1871). Primitive Culture, vol. 1, chapter XI: Animism, p. 430.