When NVIDIA announced Nemotron, a database of “synthetic personas”, we figured our clients would fairly quickly come and ask us whether this technology, or another one, could replace real users. Clients, in very different sectors, are already asking us. They’d like to use the modelling they have, or hope to get produced, to simulate behaviours with an AI. Actually, from the very first versions of ChatGPT, around 2023, there were tests where it acted out behaviours. The first reaction of most of the clients we showed them to was: “Oh that’s great, we won’t need to go and interview users, consumers, customers, whatever you call them, anymore. We’ll just have to ask ChatGPT to play the part and put our questions to it directly.” Our competitors are looking into it too, we saw it again recently, and they’re checking whether it would be feasible to simulate one-on-one interviews or user tests.
When we started the study, in March 2026, we expected it not to work. You can see it in the answers AI gives us day to day: it slips into a kind of slightly gratuitous empathy, a fake kindness. It smooths everything out, really. So we didn’t expect the results to be good at all, but we didn’t know how that would show up in the data. That’s why we wanted to compare several models and see their impact on the results, even though these are the March models and the study should probably be run again today.
And then the first protocol, built by the AI, wasn’t right. We’d explained to it very precisely that we wanted to simulate behaviours by drawing from the Nemotron database a sample representative of the one we’d surveyed ourselves. The task seemed simple enough. On its own initiative, without running it by us, it went for batch processing, to optimise (it’s always trying to save the tokens it uses). We hadn’t asked for that at all, and we noticed it a bit by chance, looking at the results. Rather than throw out that first run, we kept it in the comparison to measure the impact it had.
We had the answers of 1,000 real French people
In the summer of 2024, we ran a quantitative study for Cultura with a panel of 1,000 French people, on their cultural practices and their consumption habits. We tell the whole story in the case study. So we had each respondent’s socio-demographic profile and their answers to about ten questions: income, cultural budget, reading frequency, satisfaction, second-hand buying. By having an LLM answer the same questionnaire, from profiles drawn to approximate the same composition, we could measure the gap with the real answers, question by question.
Others have done this before us under the name “silicon samples” (Argyle et al., 2023, in Political Analysis), mostly in English and on opinion polls. We had a French questionnaire we’d designed ourselves, on a subject where social desirability weighs heavily.
Three ways to get the AI to answer
In March 2026, we generated 3,000 synthetic respondents, 1,000 in each of these three conditions:
- batch processing, the one the AI had picked by itself: Claude Sonnet 4.6 receives the profiles in batches of 5 with a standard prompt, and works through them one batch after another;
- embodiment with Sonnet 4.6: we ask it to embody each profile before answering, one at a time, 1,000 times, with the same model but a different method;
- embodiment with Claude Opus 4.6, using the same method.
Generating the answers one by one takes time and costs more than in batches. The profiles come from the Nemotron database, where we drew 1,000 profiles aiming for the real panel’s composition by sex, age and occupational category. The LLM receives these profiles as they are, and it’s up to it to simulate the answers to the attitude and behaviour questions. The study covers 10 variables. Net monthly income is a separate case, because the prompt partly sets it by giving the occupational category and age bracket, and the LLM picks the exact income bracket. The other nine are entirely simulated: opinion on purchasing power, cultural leisure time, satisfaction, annual budget, budget change, reading frequency, number of books read per year, frequency of creative hobbies and second-hand buying.
The statistical tests
For each variable and each condition, we measured the gap between the real and synthetic distributions:
- On categorical variables, we used a chi-square test of independence and a Cramér’s V (0 when the distributions are identical, 1 when they have nothing in common)
- On numerical variables, a Mann-Whitney U and a rank-biserial r
- To compare the conditions with each other on the same personas, a Wilcoxon signed-rank test
- For all three conditions at once, a Friedman test
- And a Bonferroni correction over the 30 tests (10 variables × 3 conditions), with a threshold of 0.00167
On average, batches stray furthest from the real answers
Variable by variable, here’s each condition’s gap to the real answers.
The lower the value, the closer the synthetic panel is to reality.
For “Books read per year”, the value is a rank-biserial r (Mann-Whitney test), not a Cramér's V: it can't be compared directly with the other bars.
View the data
| Variable | In batches | Sonnet embodiment | Opus embodiment | Closest condition |
|---|---|---|---|---|
| Net monthly income | V = 0.375 | V = 0.415 | V = 0.328 | Opus embodiment |
| Opinion on income | V = 0.292 | V = 0.288 | V = 0.284 | Opus embodiment |
| Cultural leisure time/week | V = 0.391 | V = 0.234 | V = 0.257 | Sonnet embodiment |
| Cultural leisure satisfaction | V = 0.407 | V = 0.004 | V = 0.140 | Sonnet embodiment |
| Annual cultural-goods budget | V = 0.363 | V = 0.375 | V = 0.297 | Opus embodiment |
| Budget change | V = 0.469 | V = 0.460 | V = 0.483 | Sonnet embodiment |
| Reading frequency | V = 0.336 | V = 0.242 | V = 0.311 | Sonnet embodiment |
| Books read per year | r = 0.088 | r = 0.101 | r = 0.026 | Opus embodiment |
| Creative-hobby frequency | V = 0.526 | V = 0.263 | V = 0.398 | Sonnet embodiment |
| Second-hand purchase | V = 0.350 | V = 0.008 | V = 0.120 | Sonnet embodiment |
Across the nine simulated variables, batch processing is never the closest to the real answers. On average, Sonnet embodiment brings the gap down from 0.358 to 0.219, almost 40% less.
The lower the value, the closer the synthetic panel is to reality.
Average over the 9 simulated variables; income, partly set by the prompt, is left out.
View the data
| Condition | Average gap | Median gap | Largest gap | Variables where it is closest |
|---|---|---|---|---|
| Sonnet embodiment | 0.219 | 0.242 | 0.460 | 6/9 |
| Opus embodiment | 0.257 | 0.284 | 0.483 | 3/9 |
| In batches | 0.358 | 0.363 | 0.526 | 0/9 |
The budget from one year to the next
If you’re used to questionnaire results, gaps in points will probably speak to you more than effect sizes. We show them as gaps to the real answers, because the client’s original distribution remains confidential.
We asked respondents whether their cultural budget had changed compared with the previous year.
Zero = identical to the real answers; a bar pointing up over-represents the answer, pointing down under-represents it.
View the data
| Answer | In batches | Sonnet embodiment | Opus embodiment |
|---|---|---|---|
| Less | −17 | −17 | −18 |
| The same | +39 | +38 | +40 |
| More | −10 | −9 | −10 |
| Don't know | −12 | −13 | −12 |
All three conditions fail just as badly (V from 0.46 to 0.48, all significant, p < 10⁻⁸⁹). The LLM underestimates budget changes, in either direction: it pulls the answers towards “the same”, about 39 points above the real answers, and gives “don’t know” a dozen points less.
Second-hand buying
We asked respondents whether they buy second-hand cultural goods.
Zero = identical to the real answers; a bar pointing up over-represents the answer, pointing down under-represents it.
View the data
| Answer | In batches | Sonnet embodiment | Opus embodiment |
|---|---|---|---|
| Yes | +32 | +1 | +12 |
| No | −32 | −1 | −12 |
Among the synthetic respondents processed in batches, the share who say they buy second-hand is 32 points higher than among the real ones. Sonnet embodiment brings the gap down to 0.9 points, and it can no longer be statistically distinguished from the real answers (chi-square = 0.11, p = 0.74, V = 0.008).
Creative hobbies
Zero = identical to the real answers; a bar pointing up over-represents the answer, pointing down under-represents it.
View the data
| Answer | In batches | Sonnet embodiment | Opus embodiment |
|---|---|---|---|
| Never | −32 | +1 | −24 |
| Rarely | +18 | +9 | +18 |
| 2-3 times/week | +21 | +4 | +9 |
This is where the conditions differ most from one another (V of 0.526 in batches, 0.263 with Sonnet embodiment). In batches, the share of people who never do any loses almost 32 points. With Sonnet embodiment, the gap is only 0.9 points.
3 results out of 30 that can’t be statistically distinguished from the real answers
Out of the 30 tests, 27 remain significant after the Bonferroni correction (p < 0.00167), so in 27 cases out of 30 the synthetic distribution differs statistically from the real one. The 3 exceptions:
| Variable | Condition | Chi-square | p | Cramér’s V |
|---|---|---|---|---|
| Leisure satisfaction | Sonnet embodiment | 0.03 | 0.86 | 0.004 |
| Books read/year | Opus embodiment | U = 470,677 | 0.32 | r = 0.026 |
| Second-hand buying | Sonnet embodiment | 0.11 | 0.74 | 0.008 |
The method changes the answers
The Friedman test compares the three conditions at once, on the same profiles (the 964 that have a valid answer in all three). On 9 variables out of 10, they produce distributions that are significantly different from each other (p from 10⁻⁹⁶ to 10⁻¹⁵). The only exception is budget change (p = 0.061), where all three fail in the same way.
| Variable | In batches | Sonnet embodiment | Opus embodiment |
|---|---|---|---|
| Net monthly income | 0.375 | 0.415 | 0.328 |
| Income opinion | 0.292 | 0.288 | 0.284 |
| Leisure time | 0.391 | 0.234 | 0.257 |
| Leisure satisfaction | 0.407 | 0.004 n.s. | 0.140 |
| Cultural budget | 0.363 | 0.375 | 0.297 |
| Budget change | 0.469 | 0.460 | 0.483 |
| Reading freq. | 0.336 | 0.242 | 0.311 |
| Books/year | 0.088 | 0.101 | 0.026 n.s. |
| Creative hobbies | 0.526 | 0.263 | 0.398 |
| Second-hand | 0.350 | 0.008 n.s. | 0.120 |
“Books/year” is measured with a rank-biserial r (Mann-Whitney U), not a Cramér's V. The “n.s.” cells can't be statistically distinguished from the real answers.
Everyone ends up in the middle
What struck us most was the smoothing. The synthetic answers have no extremes: almost no one admits their budget is going down, no one owns up to a behaviour that isn’t very socially acceptable. Whatever flavour you give the profiles, the model answers like an average profile. We measure it on five gaps:
- on budget change, as we saw, no condition corrects the gap, the largest in the study;
- the top brackets (income above €4,000, more than 20 hours of leisure a week) and the bottom ones (under €1,000, life “very difficult”) are consistently under-represented (V from 0.29 to 0.42);
- socially valued behaviours, like second-hand buying or cultural engagement, are over-represented (V from 0.01 to 0.35 depending on the condition), and embodiment reduces this bias without removing it everywhere;
- the synthetic profiles read more and do more cultural activities than the real respondents, and more often say they’re dissatisfied with not doing more (V from 0.24 to 0.53);
- the synthetic answers clearly underestimate “don’t know” on the budget and how it changes.
That’s hugely problematic, because a quantitative study also shows how the answers are spread out, extremes included, beyond the average. So the synthetic answers don’t go looking for what really makes people different, those subtle differences that come down, in a way, kind of, to being human.
Where does the bias come from?
The synthetic answers overestimate cultural engagement. We still wanted to check that it didn’t come from the database, so we had Gemini 2.5 Pro score the 1,000 Nemotron profiles from 0 to 10, according to the density of actual cultural practices they describe, like “has played the electric piano since adolescence” or “a regular at the Musée des Beaux-Arts in Tours, where he enjoys the temporary exhibitions”. We checked this scoring with a second evaluator, Claude Sonnet 4.6, and the two agree (Cohen’s kappa = 0.77, Spearman correlation = 0.79).
Depending on the evaluator, 0.3% to 5% of the profiles describe someone with almost no cultural practice. That’s not much, but in our real panel, the share of respondents who never read and never do creative hobbies is in the same ballpark. The two measures aren’t identical. On one side, a model scores a written profile; on the other, people declare what they do. But the database doesn’t look much more “cultural” than the real respondents.
View the data
| Engagement level | Score | Nemotron profiles |
|---|---|---|
| Low | 0-2 | 5.0% |
| Moderate | 3-4 | 24.1% |
| Significant | 5-6 | 37.2% |
| Intense | 7+ | 33.7% |
The model, for its part, follows what the profile says, with a moderate to strong correlation between a profile’s cultural score and its answers on all the variables.
View the data
| Variable | Sonnet embodiment (rho) | Opus embodiment (rho) |
|---|---|---|
| Relationship to culture | 0.59 | 0.72 |
| Reading frequency | 0.40 | 0.44 |
| Creative hobbies | 0.46 | 0.37 |
| Books read/year | 0.49 | 0.50 |
| Second-hand purchase | 0.40 | 0.43 |
Among the 50 least cultural profiles, none says they’re passionate about culture, and 37 out of 49 answer that they never do creative hobbies with Sonnet (24 out of 50 with Opus). When a profile describes someone who spends time at the museum “exploring contemporary art exhibitions and attending printmaking workshops”, the model answers, logically, that culture matters to them. So the small number of profiles with little cultural practice doesn’t explain the gap. On creative hobbies, Opus embodiment stays 24 points below the real panel for the share of people who never do any, much further off than Sonnet.
For us, the problem was never really on the database side, which is what it is. Even if it hadn’t been generated by an AI, we know how to build databases of tens, even hundreds of thousands of people. The most basic CRM at any one of our clients holds a good part of Nemotron’s information, over a narrower scope, and it’s not exactly great quality either. The very existence of Nemotron hinted that our clients would want to hand the whole process over to an AI, from start to finish, and if that risk exists, it has to be tested. What interested us much more, though, was seeing how, despite everything, the model manages to embody people.
And on that point, we’re fairly convinced that even with a better database, the effect would remain, because of a kind of social desirability in AI agents that try to please. People might tell us that a suitable prompt can counterbalance it, but actually, even a suitable prompt would risk tipping into the opposite bias. Today, reproducing the normal curve of the diversity of behaviours doesn’t seem possible to us with current technology and the way it was designed. Our feeling is that the problem comes from the technology itself, not from the data. It doesn’t convey, somehow, people’s humanity. That’s probably a bit reassuring, and maybe that’s why we want to think it.
What could it be useful for?
You hear people say that using these synthetic panels to pre-test a questionnaire or explore segmentations before going into the field doesn’t cost much. We’re not so sure: a handful of agents asked to simulate somewhat heterogeneous behaviours would probably do just as well, and we don’t see a huge added value in using a database like Nemotron. Today, we’re genuinely questioning the point of this kind of approach. For the past six months, NVIDIA and the start-up they worked with have made a bit of a show of announcing it, which, by the way, didn’t get a huge echo.
The limits of this study
Going back over the study files in September 2026, we found two flaws, which we were able to fix after the fact:
- the 1,000 profiles aren’t those of the real respondents but Nemotron profiles, and drawing them combined sex, age and occupational category as if these criteria were independent. So the synthetic panel is younger than the real one, with 75 fewer people aged 70 and over and 44 too many retirees aged 50 to 59. Weighting the synthetic answers to match the real composition barely moves the gaps (an average gap of 0.219 before and after for Sonnet embodiment, from 0.358 to 0.347 for batches), so this mismatch doesn’t explain what we’re measuring;
- 162 of the 1,000 Sonnet embodiment answer files were misfiled, so the comparisons between conditions sometimes set two different profiles side by side. We redid the calculations, filing each answer with its profile, and the figures in this article come from that new calculation, without the conclusions changing.
We’re going to run the study again with current models, comparing Sonnet and Opus on 100 profiles to start with.