What if we no longer needed to interview real customers?

A street artist spray-paints a rising curve climbed by small figures, the last one floating away with a red balloon
User research
Date 2 June 2026
Reading time 9 minutes
Author Sami Lini

Intro

The AI tsunami is coming, and the whole digital industry is changing radically. AI replacing designers, researchers and project managers. And of course, the question of replacing the users themselves in studies came up very quickly. AI used to simulate users taking tests. But does it work if we try to replace real people with AI, when we put them in front of a real questionnaire? Spoiler: “meh, not really”.

When NVIDIA announced Nemotron, a database of “synthetic personas”, we figured our clients would fairly quickly come and ask us whether this technology, or another one, could replace real users. Clients, in very different sectors, are already asking us. They’d like to use the modelling they have, or hope to get produced, to simulate behaviours with an AI. Actually, from the very first versions of ChatGPT, around 2023, there were tests where it acted out behaviours. The first reaction of most of the clients we showed them to was: “Oh that’s great, we won’t need to go and interview users, consumers, customers, whatever you call them, anymore. We’ll just have to ask ChatGPT to play the part and put our questions to it directly.” Our competitors are looking into it too, we saw it again recently, and they’re checking whether it would be feasible to simulate one-on-one interviews or user tests.

When we started the study, in March 2026, we expected it not to work. You can see it in the answers AI gives us day to day: it slips into a kind of slightly gratuitous empathy, a fake kindness. It smooths everything out, really. So we didn’t expect the results to be good at all, but we didn’t know how that would show up in the data. That’s why we wanted to compare several models and see their impact on the results, even though these are the March models and the study should probably be run again today.

And then the first protocol, built by the AI, wasn’t right. We’d explained to it very precisely that we wanted to simulate behaviours by drawing from the Nemotron database a sample representative of the one we’d surveyed ourselves. The task seemed simple enough. On its own initiative, without running it by us, it went for batch processing, to optimise (it’s always trying to save the tokens it uses). We hadn’t asked for that at all, and we noticed it a bit by chance, looking at the results. Rather than throw out that first run, we kept it in the comparison to measure the impact it had.

We had the answers of 1,000 real French people

In the summer of 2024, we ran a quantitative study for Cultura with a panel of 1,000 French people, on their cultural practices and their consumption habits. We tell the whole story in the case study. So we had each respondent’s socio-demographic profile and their answers to about ten questions: income, cultural budget, reading frequency, satisfaction, second-hand buying. By having an LLM answer the same questionnaire, from profiles drawn to approximate the same composition, we could measure the gap with the real answers, question by question.

Others have done this before us under the name “silicon samples” (Argyle et al., 2023, in Political Analysis), mostly in English and on opinion polls. We had a French questionnaire we’d designed ourselves, on a subject where social desirability weighs heavily.

Three ways to get the AI to answer

In March 2026, we generated 3,000 synthetic respondents, 1,000 in each of these three conditions:

  • batch processing, the one the AI had picked by itself: Claude Sonnet 4.6 receives the profiles in batches of 5 with a standard prompt, and works through them one batch after another;
  • embodiment with Sonnet 4.6: we ask it to embody each profile before answering, one at a time, 1,000 times, with the same model but a different method;
  • embodiment with Claude Opus 4.6, using the same method.

Generating the answers one by one takes time and costs more than in batches. The profiles come from the Nemotron database, where we drew 1,000 profiles aiming for the real panel’s composition by sex, age and occupational category. The LLM receives these profiles as they are, and it’s up to it to simulate the answers to the attitude and behaviour questions. The study covers 10 variables. Net monthly income is a separate case, because the prompt partly sets it by giving the occupational category and age bracket, and the LLM picks the exact income bracket. The other nine are entirely simulated: opinion on purchasing power, cultural leisure time, satisfaction, annual budget, budget change, reading frequency, number of books read per year, frequency of creative hobbies and second-hand buying.

The statistical tests

For each variable and each condition, we measured the gap between the real and synthetic distributions:

  • On categorical variables, we used a chi-square test of independence and a Cramér’s V (0 when the distributions are identical, 1 when they have nothing in common)
  • On numerical variables, a Mann-Whitney U and a rank-biserial r
  • To compare the conditions with each other on the same personas, a Wilcoxon signed-rank test
  • For all three conditions at once, a Friedman test
  • And a Bonferroni correction over the 30 tests (10 variables × 3 conditions), with a threshold of 0.00167

On average, batches stray furthest from the real answers

Variable by variable, here’s each condition’s gap to the real answers.

The lower the value, the closer the synthetic panel is to reality.

Overview: Cramér's V by variable Gap between the synthetic answers and the real ones, variable by variable (Cramér's V: the lower it is, the closer the synthetic answers are to the real ones).

For “Books read per year”, the value is a rank-biserial r (Mann-Whitney test), not a Cramér's V: it can't be compared directly with the other bars.

View the data
Overview: Cramér's V by variable
VariableIn batchesSonnet embodimentOpus embodimentClosest condition
Net monthly incomeV = 0.375V = 0.415V = 0.328Opus embodiment
Opinion on incomeV = 0.292V = 0.288V = 0.284Opus embodiment
Cultural leisure time/weekV = 0.391V = 0.234V = 0.257Sonnet embodiment
Cultural leisure satisfactionV = 0.407V = 0.004V = 0.140Sonnet embodiment
Annual cultural-goods budgetV = 0.363V = 0.375V = 0.297Opus embodiment
Budget changeV = 0.469V = 0.460V = 0.483Sonnet embodiment
Reading frequencyV = 0.336V = 0.242V = 0.311Sonnet embodiment
Books read per yearr = 0.088r = 0.101r = 0.026Opus embodiment
Creative-hobby frequencyV = 0.526V = 0.263V = 0.398Sonnet embodiment
Second-hand purchaseV = 0.350V = 0.008V = 0.120Sonnet embodiment

Across the nine simulated variables, batch processing is never the closest to the real answers. On average, Sonnet embodiment brings the gap down from 0.358 to 0.219, almost 40% less.

The lower the value, the closer the synthetic panel is to reality.

Average gap to the real answers over the 9 simulated variables (Cramér's V) Average gap to the real answers (Cramér's V) over the 9 simulated variables, by condition: the lower it is, the closer the synthetic answers are to the real ones.

Average over the 9 simulated variables; income, partly set by the prompt, is left out.

View the data
Average gap to the real answers over the 9 simulated variables (Cramér's V)
ConditionAverage gapMedian gapLargest gapVariables where it is closest
Sonnet embodiment0.2190.2420.4606/9
Opus embodiment0.2570.2840.4833/9
In batches0.3580.3630.5260/9

The budget from one year to the next

If you’re used to questionnaire results, gaps in points will probably speak to you more than effect sizes. We show them as gaps to the real answers, because the client’s original distribution remains confidential.

We asked respondents whether their cultural budget had changed compared with the previous year.

Budget change: gap between synthetic and real answers (the worst case) Gap in points between the synthetic answers and the real ones. All three conditions overestimate “the same” by about 39 points and underestimate “don't know” by a dozen points.

Zero = identical to the real answers; a bar pointing up over-represents the answer, pointing down under-represents it.

View the data
Budget change: gap between synthetic and real answers (the worst case)
AnswerIn batchesSonnet embodimentOpus embodiment
Less−17−17−18
The same+39+38+40
More−10−9−10
Don't know−12−13−12

All three conditions fail just as badly (V from 0.46 to 0.48, all significant, p < 10⁻⁸⁹). The LLM underestimates budget changes, in either direction: it pulls the answers towards “the same”, about 39 points above the real answers, and gives “don’t know” a dozen points less.

Second-hand buying

We asked respondents whether they buy second-hand cultural goods.

Second-hand buying: gap to the real answers (the best case) Gap in points to the real answers on second-hand buying: 32 points too many in batches, less than one point with Sonnet embodiment.

Zero = identical to the real answers; a bar pointing up over-represents the answer, pointing down under-represents it.

View the data
Second-hand buying: gap to the real answers (the best case)
AnswerIn batchesSonnet embodimentOpus embodiment
Yes+32+1+12
No−32−1−12

Among the synthetic respondents processed in batches, the share who say they buy second-hand is 32 points higher than among the real ones. Sonnet embodiment brings the gap down to 0.9 points, and it can no longer be statistically distinguished from the real answers (chi-square = 0.11, p = 0.74, V = 0.008).

Creative hobbies

Creative hobbies: gap to the real answers (the most contrasted case) Gap in points to the real answers by frequency. In batches, the share of people who never do any drops by almost 32 points; with Sonnet embodiment, the gap falls below 1 point.

Zero = identical to the real answers; a bar pointing up over-represents the answer, pointing down under-represents it.

View the data
Creative hobbies: gap to the real answers (the most contrasted case)
AnswerIn batchesSonnet embodimentOpus embodiment
Never−32+1−24
Rarely+18+9+18
2-3 times/week+21+4+9

This is where the conditions differ most from one another (V of 0.526 in batches, 0.263 with Sonnet embodiment). In batches, the share of people who never do any loses almost 32 points. With Sonnet embodiment, the gap is only 0.9 points.

3 results out of 30 that can’t be statistically distinguished from the real answers

Out of the 30 tests, 27 remain significant after the Bonferroni correction (p < 0.00167), so in 27 cases out of 30 the synthetic distribution differs statistically from the real one. The 3 exceptions:

The 3 tests that are not significant after the Bonferroni correction
VariableConditionChi-squarepCramér’s V
Leisure satisfactionSonnet embodiment0.030.860.004
Books read/yearOpus embodimentU = 470,6770.32r = 0.026
Second-hand buyingSonnet embodiment0.110.740.008

The method changes the answers

The Friedman test compares the three conditions at once, on the same profiles (the 964 that have a valid answer in all three). On 9 variables out of 10, they produce distributions that are significantly different from each other (p from 10⁻⁹⁶ to 10⁻¹⁵). The only exception is budget change (p = 0.061), where all three fail in the same way.

Effect size by variable and condition Green = close to the real answers, red = far from them
Effect size by variable and condition : Green = close to the real answers, red = far from them
Variable In batchesSonnet embodimentOpus embodiment
Net monthly income 0.375 0.415 0.328
Income opinion 0.292 0.288 0.284
Leisure time 0.391 0.234 0.257
Leisure satisfaction 0.407 0.004 n.s. 0.140
Cultural budget 0.363 0.375 0.297
Budget change 0.469 0.460 0.483
Reading freq. 0.336 0.242 0.311
Books/year 0.088 0.101 0.026 n.s.
Creative hobbies 0.526 0.263 0.398
Second-hand 0.350 0.008 n.s. 0.120

“Books/year” is measured with a rank-biserial r (Mann-Whitney U), not a Cramér's V. The “n.s.” cells can't be statistically distinguished from the real answers.

Everyone ends up in the middle

What struck us most was the smoothing. The synthetic answers have no extremes: almost no one admits their budget is going down, no one owns up to a behaviour that isn’t very socially acceptable. Whatever flavour you give the profiles, the model answers like an average profile. We measure it on five gaps:

  • on budget change, as we saw, no condition corrects the gap, the largest in the study;
  • the top brackets (income above €4,000, more than 20 hours of leisure a week) and the bottom ones (under €1,000, life “very difficult”) are consistently under-represented (V from 0.29 to 0.42);
  • socially valued behaviours, like second-hand buying or cultural engagement, are over-represented (V from 0.01 to 0.35 depending on the condition), and embodiment reduces this bias without removing it everywhere;
  • the synthetic profiles read more and do more cultural activities than the real respondents, and more often say they’re dissatisfied with not doing more (V from 0.24 to 0.53);
  • the synthetic answers clearly underestimate “don’t know” on the budget and how it changes.

That’s hugely problematic, because a quantitative study also shows how the answers are spread out, extremes included, beyond the average. So the synthetic answers don’t go looking for what really makes people different, those subtle differences that come down, in a way, kind of, to being human.

Where does the bias come from?

The synthetic answers overestimate cultural engagement. We still wanted to check that it didn’t come from the database, so we had Gemini 2.5 Pro score the 1,000 Nemotron profiles from 0 to 10, according to the density of actual cultural practices they describe, like “has played the electric piano since adolescence” or “a regular at the Musée des Beaux-Arts in Tours, where he enjoys the temporary exhibitions”. We checked this scoring with a second evaluator, Claude Sonnet 4.6, and the two agree (Cohen’s kappa = 0.77, Spearman correlation = 0.79).

Depending on the evaluator, 0.3% to 5% of the profiles describe someone with almost no cultural practice. That’s not much, but in our real panel, the share of respondents who never read and never do creative hobbies is in the same ballpark. The two measures aren’t identical. On one side, a model scores a written profile; on the other, people declare what they do. But the database doesn’t look much more “cultural” than the real respondents.

Cultural engagement of the Nemotron profiles Breakdown of the 1,000 Nemotron profiles by how many cultural activities they describe (scored 0 to 10 by Gemini 2.5 Pro).
View the data
Cultural engagement of the Nemotron profiles
Engagement levelScoreNemotron profiles
Low0-25.0%
Moderate3-424.1%
Significant5-637.2%
Intense7+33.7%

The model, for its part, follows what the profile says, with a moderate to strong correlation between a profile’s cultural score and its answers on all the variables.

The model follows what the profile says (Spearman correlations) Spearman correlation between a profile's cultural score and the model's answers, by variable.
View the data
The model follows what the profile says (Spearman correlations)
VariableSonnet embodiment (rho)Opus embodiment (rho)
Relationship to culture0.590.72
Reading frequency0.400.44
Creative hobbies0.460.37
Books read/year0.490.50
Second-hand purchase0.400.43

Among the 50 least cultural profiles, none says they’re passionate about culture, and 37 out of 49 answer that they never do creative hobbies with Sonnet (24 out of 50 with Opus). When a profile describes someone who spends time at the museum “exploring contemporary art exhibitions and attending printmaking workshops”, the model answers, logically, that culture matters to them. So the small number of profiles with little cultural practice doesn’t explain the gap. On creative hobbies, Opus embodiment stays 24 points below the real panel for the share of people who never do any, much further off than Sonnet.

For us, the problem was never really on the database side, which is what it is. Even if it hadn’t been generated by an AI, we know how to build databases of tens, even hundreds of thousands of people. The most basic CRM at any one of our clients holds a good part of Nemotron’s information, over a narrower scope, and it’s not exactly great quality either. The very existence of Nemotron hinted that our clients would want to hand the whole process over to an AI, from start to finish, and if that risk exists, it has to be tested. What interested us much more, though, was seeing how, despite everything, the model manages to embody people.

And on that point, we’re fairly convinced that even with a better database, the effect would remain, because of a kind of social desirability in AI agents that try to please. People might tell us that a suitable prompt can counterbalance it, but actually, even a suitable prompt would risk tipping into the opposite bias. Today, reproducing the normal curve of the diversity of behaviours doesn’t seem possible to us with current technology and the way it was designed. Our feeling is that the problem comes from the technology itself, not from the data. It doesn’t convey, somehow, people’s humanity. That’s probably a bit reassuring, and maybe that’s why we want to think it.

What could it be useful for?

You hear people say that using these synthetic panels to pre-test a questionnaire or explore segmentations before going into the field doesn’t cost much. We’re not so sure: a handful of agents asked to simulate somewhat heterogeneous behaviours would probably do just as well, and we don’t see a huge added value in using a database like Nemotron. Today, we’re genuinely questioning the point of this kind of approach. For the past six months, NVIDIA and the start-up they worked with have made a bit of a show of announcing it, which, by the way, didn’t get a huge echo.

The limits of this study

Going back over the study files in September 2026, we found two flaws, which we were able to fix after the fact:

  • the 1,000 profiles aren’t those of the real respondents but Nemotron profiles, and drawing them combined sex, age and occupational category as if these criteria were independent. So the synthetic panel is younger than the real one, with 75 fewer people aged 70 and over and 44 too many retirees aged 50 to 59. Weighting the synthetic answers to match the real composition barely moves the gaps (an average gap of 0.219 before and after for Sonnet embodiment, from 0.358 to 0.347 for batches), so this mismatch doesn’t explain what we’re measuring;
  • 162 of the 1,000 Sonnet embodiment answer files were misfiled, so the comparisons between conditions sometimes set two different profiles side by side. We redid the calculations, filing each answer with its profile, and the figures in this article come from that new calculation, without the conclusions changing.

We’re going to run the study again with current models, comparing Sonnet and Opus on 100 profiles to start with.

Want to discuss your project ?

Everyone talks about user experience, service design, ergonomics… It’s not very clear, but you’d like to explore all that or get better at it.

Contact us