By Kim Beenen, 10 september 2026
Even when we try our best to be unbiased, research shows that our stereotypes about social categories (like gender, nationality, age) often subtly leak into the way we speak and write. Instead of explicitly stating a stereotype, we (unconsciously) change our language depending on which social category we describe and/or whether the described behavior fits our stereotypes.
Consider this example: we might describe a man who changes a flat tire by saying: “He is a real handyman” (short, positive, describing a permanent trait). If a woman does the exact same thing, we might say: “She was not struggling when she changed the flat tire today” (longer, more specific, use of a negation, framing it as a one-time event).
Notice how the exact same behavior (i.e., changing a flat tire) can be described differently depending on the social category of the person involved. In this example, multiple linguistic variables shifted simultaneously: sentence length, word choice, the use of affirmation or negation, and whether the behavior is framed as a permanent trait or temporary action.
Studying how stereotypes drive such subtle, simultaneous linguistic shifts (known as linguistic biases) requires two things: (1) establishing a causal link between stereotypes and language use, and (2) capturing rich language that includes multiple linguistic variables. While closed-ended experiments typically only meet requirement 1, and content analyses of real-life language only meet requirement 2, “freely generated language” experiments, in which participants describe social categories in their own words, meet both.
In our new paper “Experimental Research on Stereotype Expression in Freely Generated Language: A Systematic Literature Review”, we present a systematic literature review of 106 freely generated language experiments (published from 1979 to 2024).
We found that researchers have identified a wide array of linguistic variables that shift depending on the social category that is described and/or whether displayed behaviors fit or clash with a stereotype. But we also found a big gap: most studies only investigate only one linguistic variable at a time. However, as our example showed, multiple variables can shift together in real-life language use. Studying them in isolation misses the bigger picture of how stereotypes are actually expressed in real-life language.
The true value of freely generated language experiments is that they capture rich language, including various co-occurring linguistic variables, together with the underlying stereotypes of the person speaking or writing. To use this paradigm to its full potential and to better understand how stereotypes spread through language, future experiments should examine multiple linguistic variables in relation to one another, rather than studying them one at a time.
These insights matter beyond academia: they can inform language policies for inclusive language, help build better (automatic) tools for spotting biased language and ensure AI systems trained on human language do not amplify existing linguistic biases.
Read the full open-access paper here: https://doi.org/10.1177/00936502261478981
