AI agents are irrational, but too similar
A research note on behavioral fidelity, population collapse, and what we are testing at Aetherya
There is a strange failure hiding inside synthetic research.
Ask a language model to act like a customer and it may hesitate, anchor on a price, prefer a fair split to a profitable one, or change its answer when the wording changes. Good. Those are recognizable human failures of judgment.
Now create one hundred such customers.
They may have different names, ages, jobs, and biographies. Yet underneath the costume, the population often behaves like one person copied one hundred times. The answers cluster. The prose converges. Everyone becomes articulate, cooperative, moderately cautious, and suspiciously capable of explaining why they feel what they feel.
That gives us a population with human-looking irrationality and machine-shaped uniformity.
For behavioral simulation, this is a more serious problem than an occasional factual error. A study can tolerate an imperfect respondent. A panel whose apparent disagreement is cosmetic teaches us very little.
Two kinds of realism
Before I use the label "human-like," I have to say what sort of likeness I mean.
The first is individual realism. Does one agent show the sorts of limits, biases, emotions, and inconsistencies found in a person?
The second is population realism. Does a group of agents reproduce meaningful variation between people, including rare views, contradictory preferences, uneven confidence, different thresholds, and different ways of changing their minds?
These properties can move independently.
A model can reproduce loss aversion while giving nearly every simulated respondent the same degree of loss aversion. It can imitate a fairness preference while making the whole population too fair. It can produce five distinct explanations that all lead to the same choice. In each case, the individual answer looks plausible while the sample distribution is wrong.
This distinction matters because most commercial synthetic research is judged by reading a few outputs. If the writing sounds natural and the personas disagree occasionally, the simulation feels convincing. A readable transcript proves that the transcript reads well. It says very little about the population behind it.
Language models behave irrationally
Experimental evidence quickly breaks the image of AI agents as cold utility maximizers.
Marcel Binz and Eric Schulz treated GPT-3 as a participant in classic cognitive psychology experiments. The model performed well on several tasks, sometimes better than people. It also failed badly after small changes to a problem, showed no sign of directed exploration, and struggled with causal reasoning. Competence and brittleness appeared in the same system.
Economic games show a similar pattern. Brookins and DeBacker put GPT-3.5 into the dictator game and the prisoner's dilemma. It leaned toward fair splits. Its cooperation rate in the prisoner's dilemma was about 65 percent. The human comparison was 37 percent. Both results sit far from a narrow game-theoretic optimizer.
In 2026, researchers tested several model families on a larger set of decision problems. Their results showed unstable, context-dependent irrationality. Behavior shifted with the task, prompt, model family, scale, and framing. Some biases weakened as models improved. Others persisted. A prompt asking for rational analysis could itself move the answer.
So "irrational" needs care here. The model never had to make payroll. It has never regretted trusting a brand. Calling its output biased describes a pattern in its answers. Any lived experience behind that pattern belongs to the people whose language shaped the training data. The system learned from human text and was tuned to produce acceptable responses.
That still matters. If an agent is used to model decisions, the pattern of its deviations becomes part of the instrument.
One model can become an entire crowd
The second problem appears when researchers sample a population from the same underlying model.
Peter Park, Philipp Schoenegger, and Chongyang Zhu attempted to replicate fourteen social science studies with GPT-3.5. In six studies, the model produced so little variation that the planned analysis became impossible. They called this the "correct answer" effect. Demographic changes in the prompt often failed to break it.
Four models. Sixteen identity groups. A comparison set built from 3,200 people. That was the setup. The published result was flattening. An identity label may change the average response while still erasing the spread of views held by real members of that group.
This is the key statistical point. Two samples can share a mean while holding opposing distributions.
Suppose real buyers rate a new offer on a ten-point scale. Half give it a two because they distrust the category. Half give it an eight because the product solves an urgent problem. The mean is five. A synthetic panel in which everyone gives it a five has the same mean and none of the same market structure.
That panel will miss both the rejection risk and the enthusiastic niche.
Creative work shows the same population effect from another angle. In a controlled experiment with short-story writers, access to ideas from generative AI improved average quality, especially for less creative writers. The AI-assisted stories also became more similar to one another. Individual performance rose while collective novelty narrowed.
The agent version is easy to imagine. Each respondent sounds coherent. The report looks polished. The population has quietly lost its edges.
Demographics are weak behavioral instructions
The standard fix is to add persona detail.
One agent becomes a 28-year-old designer in Berlin. Another becomes a 54-year-old procurement director in Manchester. A third becomes a price-sensitive parent in Bucharest. More fields should create more difference.
Sometimes they do. Lisa Argyle and her coauthors showed that GPT-3, conditioned on detailed backstories from real survey respondents, could reproduce several patterns found in human subgroups. They called this "algorithmic fidelity." Their work gives synthetic samples a serious empirical case.
The result is conditional. Fidelity must be tested for the group and task in question. Each new domain and population has to establish it again.
Research built around a network of 64 human beliefs makes the weakness of demographic role-play clearer. Demographic information alone produced weak alignment between agent and human opinions. Giving the agent one relevant belief improved alignment on connected topics. Unrelated topics stayed largely unchanged.
That is closer to how people work. Age offers a weak clue about a position on climate policy or a reaction to a landing page. Beliefs connect to other beliefs. Trust has a history. A price feels different depending on reference points, financial pressure, category knowledge, and who is asking.
A demographic profile describes a person from the outside. A behavioral model needs some account of what changes inside.
The missing unit is a distribution of minds
Synthetic research needs to move past asking whether an agent can impersonate a person. The useful question concerns the system's ability to generate a distribution of decision processes.
That requires at least four tests.
First, marginal fidelity. Does each variable have a plausible distribution? A center-heavy result across risk tolerance, trust, price sensitivity, confidence, and purchase intent should trigger suspicion.
Second, joint fidelity. Do variables relate to one another in believable ways? Financial stress may increase price attention, with the size of that effect shaped by loyalty and prior trust. A pile of independently randomized traits creates incoherent variety.
Third, response fidelity. When exposed to the same stimulus, do agents diverge for reasons connected to their state, history, and beliefs? Different wording only shows lexical variety.
Fourth, temporal fidelity. Do agents carry consequences forward? If evidence raises trust in one round, that change should affect later choices. If every call starts from a clean prompt, the persona has no life, only a costume.
These tests also reveal why temperature is an incomplete fix. More sampling noise can increase lexical variety while decision rules stay fixed. A panel can say the same thing in one hundred colorful ways.
What we are researching at Aetherya
This tension sits at the center of our work on Aetherya.
Our goal is structured bias and measurable diversity. Random contradiction is easy and theatrically human. It tells us almost nothing.
Our current approach separates three layers.
The audience layer defines the population before individual personas are generated. We model weighted segments, attach source evidence, record unsupported assumptions, and expose confidence. This is meant to prevent a fluent model from inventing a neat audience and presenting it as fact.
The persona layer carries behavioral variables beyond demographics. These include decision style, risk tolerance, cognitive load threshold, trust history, emotional triggers, technology affinity, goals, and constraints. Coherent downstream differences are the real test of those fields.
The state layer governs the reaction itself. THYMOS is our work on modeling hesitation, bias, emotion, fatigue, objection, and changes in trust across a decision. An agent's response should follow from its current state. A generic instruction to "act human" gives us theater.
We are making distributions the unit of evaluation. Screenshots of convincing dialogue reveal too little. A useful benchmark needs to compare synthetic and human response distributions, segment calibration, within-group variance, between-group distance, test-retest stability, and sensitivity to framing. It should track whether repeated agents collapse toward the same answer and whether a persona remains coherent over several turns.
The harder standard is predictive validity. Given a stimulus and information available before launch, can the simulation identify the objections, segment splits, and directional changes that later appear in human interviews, experiments, analytics, or market outcomes?
That work should be reported with uncertainty. A small synthetic sample deserves an interval around its estimate. Weak source evidence should lower confidence. A thin segment stays thin, however eloquent its prose becomes.
This remains unsolved, here and across the field.
The useful claim today is narrower. A synthetic panel can reveal a hypothesis worth chasing or an assumption worth testing. Human validation still has to carry the weight. One hundred calls to the same model remain one hundred calls, even after each receives a different name.
Irrationality alone tells us little
Early criticism often painted AI decision systems as overly rational and unusually clean versions of human judgment.
Now the stranger risk is visible. They can inherit our biases while their populations remain narrow. They can reproduce the shape of a human mistake while smoothing away the people who reach it for different reasons, along with those who avoid it.
A good behavioral simulation needs agents that are wrong in different, structured, testable ways.
Some should anchor on the old price. Some should ignore price and worry about switching cost. Some should trust a recommendation from a peer and distrust the same claim from a brand. Some should change their mind after evidence. Some should become more resistant because the evidence feels like pressure. A few should react in ways the average model considers unlikely.
That variation carries signal. It is often where the decision lives.
The strongest synthetic research system will explain, measure, and validate why its respondents diverge. The real work begins once the dialogue sounds human.