The State of Synthetic Audiences in Marketing Research
An evidence review, 2023 to 2026. What synthetic audiences can actually do today, what they cannot, and how to tell the difference before you bet a decision on one.
Synthetic audiences, meaning survey respondents simulated by large language models, entered marketing research in 2023 with results that did not justify commercial use. Willingness-to-pay estimates came back wrong-signed. Qualitative responses read as stereotypes. Subgroup analysis failed outright. The most-cited industry test of the period concluded the method was not good enough to supplement human sample.
Three years later the picture has changed materially. A study using 57 real personal care concept surveys and 9,300 human responses recovered 90% of human test-retest reliability while preserving realistic response distributions. A study of 70 nationally representative survey experiments predicted treatment effects at r = 0.85. A large-scale comparison across eight sampling sources found a synthetic audience landing within the range produced by seven human panels, in a study funded by the platform under test.
The improvement has two sources, and the second matters more than the first. Models got better. More importantly, the field learned that how you elicit a response and what you ground it in determine validity far more than raw model scale. The single largest accuracy gain in the literature came from changing the question format, holding the model constant.
Significant limitations persist and some appear structural. Synthetic respondents remain weak on individual-level variance, subgroup fidelity, causal estimation, and genuinely novel categories. A classifier can distinguish synthetic survey data from real census responses at near-perfect rates. Independent testing continues to find statistically significant divergence from human benchmarks on most questions.
The defensible position for a marketing organization in 2026 is bounded adoption.
Synthetic audiences are ready for early-stage screening, message triage, instrument piloting, and hard-to-reach audience exploration, provided the work is grounded in real data and validated against a human holdout. They are not ready for final go/no-go decisions, population estimates, claim substantiation, or segmentation deliverables. That boundary is likely to move over the next 24 months, and organizations that build validation capability now will be positioned to move with it.
Why this question is being asked now
Traditional consumer research is slow and expensive. Industry pricing benchmarks for 2026 put focus groups at $7,000 to $20,000 per group with most studies requiring four to six groups, a 1,000-respondent quantitative survey at $20,000 to $80,000, and a full concept testing programme for a large CPG company at $200,000 to $500,000 annually across 10 to 20 studies. Typical timelines run 4 to 12 weeks from brief to final report.
- CPG concept testing programme$200,000 to $500,000
Annual, across 10 to 20 studies.
- Quantitative survey, 1,000 respondents$20,000 to $80,000
Four to twelve weeks from brief to report.
- Focus groups$7,000 to $20,000
Per group. Most studies need four to six.
- Entry-level synthetic platform$480 to $1,188
Per year. Results in hours rather than weeks.
Bars are scaled to the upper bound of each range. The economic pressure to adopt is obvious, which is precisely why the validity question deserves careful treatment.
The market has moved quickly. Qualtrics added synthetic respondents to its survey platform. Toluna built over one million synthetic personas from its 79-million-member human panel. NIQ BASES, a standard vendor inside CPG innovation pipelines, now offers synthetic capability. Simile raised $100 million from Index Ventures in February 2026.
Practitioner sentiment has not kept pace with vendor activity. The Rival Group 2026 Market Research Trends Report found 42.75% of market researchers described themselves as not excited about synthetic respondents, even while welcoming other AI applications. A separate 2026 survey of UX researchers found 47% skeptical and wanting more evidence before trusting them, with 24% cautiously optimistic. The hesitation is concentrated among the people closest to the work, which is worth taking seriously.
The first generation failed, and the record is clear
Any honest account of this field has to start with how poorly it performed at launch. The 2023 evidence base is unambiguous.
- 2023
Foundational proof of concept, with caveats
Argyle and colleagues introduced the concept of algorithmic fidelity, showing that models conditioned on socio-demographic backstories could emulate response distributions for US subgroups. This established the idea was worth studying. It did not establish that it worked commercially. - 2023
Willingness to pay came back wrong-signed
Brand, Israeli and Ngwe tested GPT-3.5 on conjoint and pricing tasks. Their finding, in their own framing, was that estimates are sometimes comparable to human studies but are often inaccurate and in some cases wrong-signed. A pricing input that can invert direction is unusable. - 2023
Systematic distortion appeared immediately
Aher, Arriaga and Kalai replicated several classic findings, including the Ultimatum Game where three of four human results fell near the model trend line. The same work surfaced what they named hyper-accuracy distortion, a tendency for models to produce answers that were too correct and too consistent to be human. - 2023
Industry testing reached a blunt conclusion
Kantar ran a side-by-side test with 5,000 respondents. The synthetic sample showed strong positive bias and failed to capture subgroup trends. On repeated qualitative questions, responses veered towards the stereotypical and lacked variance or nuance. Kantar's stated conclusion was that off-the-shelf synthetic sample is not good enough to use as a supplement for human sample. - 2023
Opinion representation was skewed
OpinionQA found substantial misalignment between model opinions and 60 US demographic groups, comparable in magnitude to the Democrat-Republican divide on climate change. Misalignment persisted after steering and was worst for groups including those over 65 and those widowed.
That is the baseline. Anyone selling synthetic audiences in 2023 was selling something the evidence did not support.
What actually changed
Three things improved between 2023 and 2026. They are worth separating because they carry different implications for how a brand should evaluate a vendor.
Driver one: model capability
The obvious factor. GPT-3.5 gave way to GPT-4, GPT-4o, and successors. Reasoning, instruction following, and consistency all improved. This matters, and it is the driver vendors talk about most.
It is also the least interesting of the three, because it is the one a buyer has no control over and the one that explains the smallest share of the gain.
Driver two: elicitation method, which mattered more than model scale
This is the central finding of the last two years and it deserves emphasis.
Early work asked models to output a number directly. Give this concept a purchase intent score from 1 to 5. Models responded with distributions that were narrow, clustered toward the center, and lacking the variance real human panels produce. This was read as evidence of a deep structural flaw.
Researchers at PyMC Labs, working with Colgate-Palmolive, tested a different approach. Instead of requesting a number, they elicited a natural language response about purchase intent, then mapped that text to a Likert distribution using embedding similarity against reference statements. They called the method Semantic Similarity Rating.
The validation set was substantial and commercially real: 57 consumer surveys on personal care product concepts, 150 to 400 US participants each, 9,300 human responses total, with demographic attributes attached.
of human test-retest reliability, recovered by changing how the question was asked rather than which model answered it. Distributional similarity above 0.85.
The method recovered both panel-level response distributions and the relative ranking of concepts by mean purchase intent. It generalized to other question types, reaching ρ = 82% and K = 0.81 on concept relevance, and zero-shot elicitation outperformed a LightGBM classifier trained on demographic and product features. Accuracy required prompting the model to consider demographic attributes, and model response behavior with respect to age and income mirrored human response behavior reasonably well.
The model was GPT-4o at temperature 0.5. Nothing exotic. The variance collapse that the 2023 literature treated as a fundamental limitation turned out to be substantially an artifact of asking the question the wrong way.
Driver three: grounding data
The second large gain came from what the model is conditioned on.
Park and colleagues at Stanford, Northwestern, and Google DeepMind built generative agents from two-hour interviews with 1,052 real individuals. Those agents replicated their subjects' General Social Survey answers 85% as accurately as the individuals replicated their own answers two weeks later. Interview-grounded agents outperformed demographic-only agents and reduced accuracy bias across racial and ideological groups.
Toubia and colleagues at Columbia extended this with Twin-2K-500, building digital twins of 2,058 people from 500-plus questions each. Twins reached 72% accuracy on holdout data, an 88% relative accuracy against the test-retest baseline. Brand, Israeli and Ngwe found that fine-tuning on prior survey data improved alignment for existing and new features within a category.
The pattern across all three: richer grounding in real data about real people produces better results than demographic prompting. This is the second question to ask a vendor. What real data is this grounded in, and how much of it is mine?
The current evidence base
Organized by what is being measured. Concept testing and purchase intent is where the strongest results sit: the Colgate-Palmolive work above, alongside Li and colleagues finding agreement rates above 75% on brand similarity and product attribute ratings.
On predicting experimental outcomes, Hewitt and colleagues tested 70 pre-registered nationally representative US survey experiments, around 470 treatment effects across 105,000 to 119,000 participants. Correlation with actual treatment effects reached r = 0.85, rising to r = 0.90 for unpublished studies outside the training data. The models systematically overestimated effect magnitude. Toubia's digital twins replicated only about half of between- and within-subject treatment effects.
On population-level survey replication the picture is more mixed. Kieslich and colleagues tested six models against the World Values Survey and found 94.4% of 90 model-question comparisons were statistically different from the human benchmark at p ≤ 0.05. Rafikova and Voronin, simulating attitudes toward immigration, gender stereotypes, and family values, found ChatGPT produced responses aligned with liberal-prosocial norms that likely reflect alignment and safety training rather than stable ideological commitments.
The largest comparison came from Muthukrishna and Warner: 7,755 UK participants, one synthetic platform against seven human sampling sources, same instrument, pre-registered. The synthetic audience fell within the range of outcomes produced by human samples across most questions.
Separately, Wang, Zhang and Zhang built a transfer-learning estimator pooling LLM-generated and real choice data, recovering efficiency gains of 25 to 80 percent in conjoint estimation. This is a bounded, testable application that reduces human panel cost without replacing the panel.
What has not improved
The improvement narrative has real limits, and several of these appear structural rather than temporary.
Individual-level variance structure. Aggregate averages replicate well. The distribution of individual differences does not. Park, Schoenegger and Zhu ran replications of 14 Many Labs 2 studies with GPT-3.5 and replicated only 37.5% of the original results. Six studies could not be analysed at all because of what they named the "correct answer" effect, where different runs returned near-zero variation on questions probing political orientation, economic preference, and moral philosophy. In their Moral Foundations replication, the model identified as politically conservative in 99.6% of cases and as liberal in 99.3% of cases under reversed answer order, while both conditions showed right-leaning moral foundations.
STRAT7, testing two synthetic providers against real respondent data with Dunnhumby, found a bunching effect with fewer responses at the extremes of the scale than real data produces. Its 2026 follow-up found purely synthetic respondents breaking logical order in a price ordering exercise 68% of the time, falling to 32.8% when blended with real data. The Nuremberg Institute for Market Decisions, testing personalized digital twins built across ten demographic variables, found that models tend to overgeneralize and default to socially desirable responses, limiting their ability to reflect the nuance and diversity of real consumer opinion.
Subgroup and segment fidelity. This is the most persistent failure. Wang, Morgenstern and Dickerson tested four LLMs against 3,200 human participants across 16 demographic identities, and found models both misportray and flatten identity groups. Kantar found subgroup failure in 2023, and the STRAT7 work with Dunnhumby found divergent segmentation and key driver analysis, ruling synthetic out for both.
Detectability. Dominguez-Olmedo, Hardt and Mendler-Dünner evaluated 43 models against the US American Community Survey. Responses were governed by ordering and labeling biases, with a systematic pull toward the first option. Correcting for those biases pushed models toward uniformly random answers. A binary classifier could distinguish model-generated data from real census responses almost perfectly.
For an organization worried about data integrity, this is the single most uncomfortable finding in the literature.
Causal estimation. Gui and Toubia showed that varying a treatment such as price in blind simulations systematically affects variables that should remain constant, violating unconfoundedness and producing implausible demand curves that fail to slope downward. This is a logical problem with the simulation design, and model scale does not obviously fix it.
Novel categories. Synthetic respondents are backward-looking by construction. They are trained on how people responded to things that already exist. Performance on genuinely novel concepts with no analog in the training distribution is unreliable, and this limitation is structural.
Cultural and linguistic coverage. Durmus and colleagues found model responses skew toward the opinions of certain populations, particularly the USA and some European and South American countries.
Prompt sensitivity and drift. Bisbee and colleagues found ChatGPT persona averages tracked ANES averages while producing too little variance, different regression coefficients, sensitivity to prompt wording, and measurable drift over a three-month period. Drift means validation is a recurring obligation, not a one-time gate.
Two further critiques sit alongside these. Agnew and colleagues argue that substituting synthetic participants conflicts with the foundational research values of representation, inclusion, and understanding. Brucks and Toubia identify five systematic distortions that persist in digital twins even when built on validated data.
The benchmark problem: compared to what?
One finding from 2026 reframes the entire debate and deserves its own section.
The LSE / NYU study compared eight sources on an identical instrument: two opt-in panels, Prolific, two multi-source aggregators, two river samples recruited through Facebook and Instagram, and one synthetic benchmark. Low-quality responses including bots, duplicates, and failed attention checks were removed before analysis.
The synthetic audience landing within the human range was one finding. The other was that human panels disagreed with each other substantially.
difference in voting intention between human panels running the identical instrument. Statistically significant differences appeared on all seven substantive questions.
River samples showed nearly twice the rate of failed attention checks and duplicate accounts seen in Prolific and traditional panels. No single human platform performed best across demographic representativeness, response variability, and data quality considered together. After demographic weighting, cross-platform differences fell by more than 80%, indicating that sample composition rather than response behavior drove most of the divergence.
This matters for how the validity question gets framed. Human panel data is not a fixed ground truth. It is an estimate carrying its own substantial, and often undisclosed, source-dependent error. The honest comparison is between two imperfect instruments with different error profiles, and the choice of human recruitment source alone can materially change the strategic conclusion.
This point should be made carefully and not overplayed. Human panel variance does not license synthetic error. It does mean that a brand demanding perfect correspondence between synthetic and human results is holding synthetic to a standard human panels do not meet against each other. The relevant test is decision consistency, whether both methods lead to the same choice.
A related pressure runs in the same direction. Survey panels carry their own data quality problem. The eight-source comparison above found river samples showing nearly twice the rate of failed attention checks and duplicate accounts seen in Prolific and traditional panels. The Advertising Research Foundation has opened a dedicated initiative on survey fraud, low-quality respondents, and AI-assisted response behaviour, which indicates the industry treats this as unresolved. Some proportion of responses in any human panel is already machine-generated, undisclosed and unvalidated. Synthetic data that is labeled, inspectable, and reproducible has a governance argument that unverified human panel data does not.
Where it works and where it fails
The boundary below is the practical output of everything above. Both columns are supported by the same evidence base.
Early-stage concept screening
Narrowing a large concept pool before human validation. The Colgate-Palmolive work validated exactly this use case in exactly this category, recovering both distributions and concept rank order.
Message and copy triage
Directional reads on which messages resonate, with human validation on the finalists.
Survey instrument piloting
Testing question wording and structure before spending on fielding. Low risk, immediate cost saving.
Reducing human sample requirements
Efficiency gains of 25 to 80 percent in conjoint estimation, keeping humans in the loop while cutting cost.
Hypothesis generation
Producing candidate explanations for a research team to test properly.
Hard-to-reach audience exploration
B2B decision-makers, regulated buyers, multi-market executives, and low-incidence populations where human recruitment is slow or infeasible. Directional only, and weaker where training data is thin.
Subgroup augmentation
Statistical boosting of small segments within a real dataset, distinct from full synthesis. Independent validation across 28,630 respondents found effective sample size gains of roughly 1.85x to 3.5x for the smallest segments, with the caveat that weak inputs amplify noise. Sources disagree on how far this can be pushed. Fairgen reports gains boosting subsegments up to roughly 15% of responses. STRAT7 recommends capping synthetic contribution at no more than 5% of total sample. Treat 5% as the conservative bound until you have validated on your own data.
Early-stage concept screening
Narrowing a large concept pool before human validation. The Colgate-Palmolive work validated exactly this use case in exactly this category, recovering both distributions and concept rank order.
Message and copy triage
Directional reads on which messages resonate, with human validation on the finalists.
Survey instrument piloting
Testing question wording and structure before spending on fielding. Low risk, immediate cost saving.
Reducing human sample requirements
Efficiency gains of 25 to 80 percent in conjoint estimation, keeping humans in the loop while cutting cost.
Hypothesis generation
Producing candidate explanations for a research team to test properly.
Hard-to-reach audience exploration
B2B decision-makers, regulated buyers, multi-market executives, and low-incidence populations where human recruitment is slow or infeasible. Directional only, and weaker where training data is thin.
Subgroup augmentation
Statistical boosting of small segments within a real dataset, distinct from full synthesis. Independent validation across 28,630 respondents found effective sample size gains of roughly 1.85x to 3.5x for the smallest segments, with the caveat that weak inputs amplify noise. Sources disagree on how far this can be pushed. Fairgen reports gains boosting subsegments up to roughly 15% of responses. STRAT7 recommends capping synthetic contribution at no more than 5% of total sample. Treat 5% as the conservative bound until you have validated on your own data.
Final go/no-go decisions
Major capital allocation should not rest on simulated respondents.
Population estimates with defensible confidence intervals
Synthetic studies do not produce statistically defensible claims about what percentage of a market thinks something.
Advertising claim substantiation
Regulatory bodies including NAD in the US and ASA in the UK expect evidence from real consumers. State this to legal stakeholders before it is raised as an objection.
Segmentation deliverables
The most consistently documented failure across 2023, 2025, and 2026 testing.
Genuinely novel categories
Structural limitation with no clear path to resolution.
Deep qualitative insight
Nielsen Norman Group tested a synthetic platform against three of its own real-participant studies and concluded synthetic-user responses for many research activities are too shallow to be useful, with a tendency to please and to provide overly favorable feedback.
Regulated research
Clinical, financial, and safety-relevant contexts.
A validation protocol
Any organization adopting synthetic audiences should be able to show its parent company, its insights function, and its legal team a documented validation process. The following is a defensible standard.
Stage 1. Backtest against your own history. Select 8 to 12 completed studies where the human answer is already known. Span easy questions such as concept ranking and awareness and hard ones such as segmentation, pricing, and emotionally nuanced qualitative. Field the identical instruments to candidate synthetic platforms blind to the known results.
Stage 2. Score on metrics that expose the known failure modes. Correlation of averages alone will hide every documented weakness. Measure rank-order agreement on concept ordering, top-box and top-two-box agreement in percentage points, distributional similarity to catch variance collapse, variance and entropy against the human distribution, subgroup-level error reported separately for every segment that matters commercially, decision consistency, and a discriminator test of whether a simple classifier can separate synthetic from human responses.
Stage 3. Stress-test robustness. Randomize answer option order and re-run. Reverse question order and re-run. Quantify how much results move. Bisbee and colleagues documented meaningful drift over three months, so establish the baseline you will monitor against.
Stage 4. Set thresholds before seeing results. Committing in advance prevents rationalizing a disappointing result. Reasonable working bars for directional use: rank-order agreement above 0.8, top-box agreement within a few percentage points, decision consistency above 85% on backtested studies, and distributional similarity in the range the published literature achieves, with 0.85 as a useful reference point from the Colgate-Palmolive work. Anything below those bars is triage-only, and subgroup results should be assumed unreliable until specifically demonstrated otherwise.
Stage 5. Govern it as an ongoing obligation. Maintain a human holdout on live work. Re-validate quarterly, since models update and drift is documented. Keep an audit trail. Disclose synthetic methodology in any deliverable that travels beyond the immediate team.
Trajectory
The direction of travel is consistent. Wrong-signed pricing estimates in 2023. Recovery of 90% of human test-retest reliability on real CPG concept data in 2025. Publication in Nature in 2026. Every year of the last three has produced results better than the year before, across independent research groups using different methods.
The gains are coming from method as much as model. This is the more useful observation, because method improvements are faster and more reproducible than waiting for the next frontier model. The variance collapse problem that dominated the 2023 critique was substantially solved by changing the elicitation format. Similar unlock potential likely remains in grounding architecture, calibration against first-party data, and hybrid designs.
Some limitations may not resolve. Causal estimation, novel category performance, and individual-level variance structure have resisted improvement so far and have plausible structural explanations. It would be a mistake to assume all current weaknesses are temporary.
Two developments will materially update this assessment. The Advertising Research Foundation is running a controlled comparison using identical questionnaires across synthetic and traditional online samples, publishing monthly briefs with a synthesis report at the end of 2026. That will be the first substantial independent industry benchmark. Separately, independent replication of the Colgate-Palmolive elicitation results in other categories would establish whether that finding generalizes beyond personal care.
Organizations that build validation capability during this period will be able to expand usage as the evidence supports it. Organizations that wait for certainty will adopt later and without the internal muscle to evaluate what they are buying.
Conclusion
The evidence supports a specific and bounded position.
Synthetic audiences work well enough today for early-stage screening, message triage, instrument piloting, hypothesis generation, and exploration of hard-to-reach audiences, provided they are grounded in real data, use validated elicitation methods, and are checked against a human holdout. In those applications the speed and cost advantages are large and the risk is contained.
They do not work well enough for final decisions, population estimates, claim substantiation, segmentation, or novel categories. Those boundaries are supported by consistent evidence across multiple independent research groups and should be stated plainly to any stakeholder evaluating adoption.
The field has improved substantially and continues to improve. That improvement is real, documented, and driven by identifiable factors. It does not yet extend to every use case, and the organizations getting the most value are the ones that know precisely where the boundary sits and validate against it continuously.
Every claim, and how well it is evidenced
Each claim in this review, with its supporting figures and how strongly it is supported. Filter it. If you only trust strong evidence, read only that. If you want to know which findings were paid for by the companies selling the product, that is one click.
Showing 18 of 18 claims.
- Strong
Synthetic respondents can reproduce human purchase intent for concept testing
90% of human test-retest reliability; distributional similarity above 0.85; 57 surveys, 9,300 human responses
PyMC Labs and Colgate-Palmolive, 2025
- Strong
LLMs can predict social science experiment outcomes
r = 0.85 with actual treatment effects, r = 0.90 for unpublished studies; overestimates magnitude
Hewitt, Ashokkumar, Ghezae, Willer, Nature, 2026
- Strong
Interview-grounded agents approach human self-consistency
85% of two-week human test-retest reliability; 1,052 participants
Park et al., 2024
- Strong
Digital twins replicate averages but only half of experimental effects
72% holdout accuracy, 88% relative; around 50% of treatment effects
Toubia et al., Marketing Science, 2025
- Strong
LLMs match humans on brand perception and attribute analysis
Agreement above 75%, reported 75 to 85%
Li, Castelo, Katona, Sarvary, Marketing Science, 2024
- StrongInterested party
Off-the-shelf models produce positive bias and fail subgroups
5,000 respondents; failed subgroup pricing trends; stereotyped qualitative
Kantar, 2023
- Strong
Synthetic survey data is distinguishable from real survey data
Near-perfect classifier separation; option-order bias; 43 models tested
Dominguez-Olmedo, Hardt, Mendler-Dünner, NeurIPS, 2024
- Strong
LLMs misportray and flatten demographic identity groups
4 LLMs, 3,200 humans, 16 identities
Wang, Morgenstern, Dickerson, Nature Machine Intelligence, 2025
- Strong
Synthetic responses show too little variance and prompt sensitivity
Averages track ANES; variance too low; drift over 3 months
Bisbee et al., Political Analysis, 2024
- Strong
Early willingness-to-pay estimates were unreliable
Often inaccurate and in some cases wrong-signed, on GPT-3.5
Brand, Israeli, Ngwe, HBS WP 23-062, 2023
- Moderate
Most LLM responses differ statistically from human benchmarks
94.4% of 90 model-question comparisons at p ≤ 0.05
Kieslich et al., 2025
- ModerateVendor
Human sampling platforms disagree substantially with each other
Up to 20pp difference on voting intention across 8 sources; 7,755 participants; weighting cut differences 80%
Muthukrishna and Warner, LSE / NYU, 2026
- ModerateVendor
A synthetic audience can fall within the human sample range
Electric Twin within human range on most questions, pre-registered
Muthukrishna and Warner, LSE / NYU, 2026
- Moderate
LLM attitudinal responses carry systematic alignment bias
Responses skewed to liberal-prosocial norms across gender, migration, and family values domains
Rafikova and Voronin, J. Computational Social Science, 2026
- ModerateVendor
Synthetic augmentation boosts small segments
Effective sample size gains 1.85x to 3.5x; 28,630 respondents
Fairgen and Dig Insights
- Moderate
Synthetic qualitative output is shallow and over-agreeable
Three real-participant studies replicated
Nielsen Norman Group
- Moderate
Practitioner sentiment remains cautious
42.75% of researchers not excited; 47% of UX researchers skeptical
Rival Group, 2026; User Interviews, 2026
- Emerging
Synthetic improves conjoint estimation efficiency
Efficiency gains of 25 to 80 percent via transfer-learning estimator pooling LLM and real choice data
Wang, Zhang, Zhang, arXiv, 2024, revised 2026
How this was made
This is an evidence review, not original fieldwork. It compiles and compares published work from peer-reviewed journals, conference proceedings, preprints, and industry testing between 2023 and 2026. Every principal claim in the review is listed in the table above with its supporting figures, its source, and an assessment of how strongly it is evidenced. Nothing is asserted here that is not attributable to a listed source.
Where a study was funded by, or reported by, a party selling the product it evaluates, that is flagged in the table and at the point the finding is discussed. Some sources here are commercial parties reporting findings against their own commercial interest. Kantar sells both research services and synthetic sample, and its 2023 conclusion argued against the category. Findings of that shape are noted where they appear, and they are treated as stronger rather than weaker for it.
Cost benchmarks are quoted from a published 2026 buyer's guide rather than collected first-hand.
What we don't know
This review has no primary data of its own. It cannot tell you how any specific vendor performs on your category, your audience, or your questions. That is what the validation protocol in section eight is for, and running it is the only way to answer the question for your own organization.
The evidence base turns over quickly. Findings dated 2023 tested models several generations old and should be read as historical baseline rather than current performance. By the same logic, the 2026 findings here will age. Re-baselining is recommended after the Advertising Research Foundation synthesis report expected at the end of 2026, which will be the first substantial independent industry benchmark and may well contradict parts of this.
The strongest single result in this review, the Colgate-Palmolive elicitation work, has not been independently replicated in another category. It is personal care, one product class, one research team. Until it is reproduced elsewhere, treating it as a general property of synthetic audiences rather than a promising result in one setting would be overreading it. The method has drawn independent interest, with the Nuremberg Institute for Market Decisions identifying semantic similarity rating as a promising direction for addressing known limitations in synthetic survey data. That is corroboration of the approach rather than replication of the result, and the caveat stands.
The claim that human panels disagree with each other, which does real work in section six, rests substantially on a vendor-funded study. The design was pre-registered and the human fieldwork independent, which matters, but a single study should not carry an argument this load-bearing on its own.
Several sources cited here are industry reports rather than peer-reviewed work. They are marked as such in the source list.
- [1]
Argyle, Busby, Fulda, Gubler, Rytting, Wingate. Out of One, Many: Using Language Models to Simulate Human Samples, Political Analysis 31(3), 2023.
- [2]
Aher, Arriaga, Kalai. Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies, ICML, 2023.
- [3]
Brand, Israeli, Ngwe. Using LLMs for Market Research, Harvard Business School WP 23-062, 2023.
- [4]
Santurkar, Durmus, Ladhak, Lee, Liang, Hashimoto. Whose Opinions Do Language Models Reflect?, ICML, 2023.
- [5]
Kantar. What is synthetic sample and is it all it's cracked up to be?, 2023.
Industry side-by-side test, 5,000 respondents.
- [6]
Durmus et al.. Towards Measuring the Representation of Subjective Global Opinions in Language Models, arXiv:2306.16388, 2023.
- [7]
Gui, Toubia. The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective, arXiv:2312.15524, 2023.
- [8]
Li, Castelo, Katona, Sarvary. Determining the Validity of Large Language Models for Automated Perceptual Analysis, Marketing Science 43(2), 2024.
- [9]
Park et al.. Generative Agent Simulations of 1,000 People, arXiv:2411.10109, 2024.
- [10]
Park, Schoenegger, Zhu. Diminished diversity-of-thought in a standard large language model, Behavior Research Methods 56(6), 2024.
Preprint: https://arxiv.org/pdf/2302.07267
- [11]
Bisbee, Clinton, Dorff, Kenkel, Larson. Synthetic Replacements for Human Survey Data? The Perils of Large Language Models, Political Analysis 32(4), 2024.
- [12]
Dominguez-Olmedo, Hardt, Mendler-Dünner. Questioning the Survey Responses of Large Language Models, NeurIPS, 2024.
- [13]
Agnew et al.. The Illusion of Artificial Inclusion, CHI, 2024.
- [14]
PyMC Labs and Colgate-Palmolive. LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings, arXiv:2510.08338, 2025.
- [15]
Toubia, Gui, Peng, Merlau, Li, Chen. Twin-2K-500, Marketing Science 44(6), 2025.
- [16]
Wang, Morgenstern, Dickerson. Large language models that replace human participants can harmfully misportray and flatten identity groups, Nature Machine Intelligence 7(3), 2025.
- [17]
Brucks, Toubia. Digital Twins are Funhouse Mirrors, 2025.
- [18]
Kieslich et al.. Synthetic social data: trials and tribulations, arXiv:2510.19952, 2025.
- [19]
Hewitt, Ashokkumar, Ghezae, Willer. Large language models can predict the results of social science experiments, Nature, 2026.
- [20]
Rafikova, Voronin. ChatGPT as a research proxy: simulating human attitudes in social science research, Journal of Computational Social Science 9(17), 2026.
- [21]
Muthukrishna, Warner (LSE / NYU). Comparative study of synthetic and human sampling sources, 7,755 UK participants, 2026.
Funded by Electric Twin, the platform tested. Both authors hold equity.
- [22]
Wang, Mengxin (UT Dallas), Zhang, Dennis J. (Washington University in St. Louis), Zhang, Heng (Arizona State University). Large Language Models for Market Research: A Data-augmentation Approach, arXiv:2412.19363, December 2024, revised 2026.
Also at https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5057769
- [23]
STRAT7 with Dunnhumby. Putting Synthetic Data to the Test: A Real-World Evaluation, 2025.
Independent comparison of two synthetic providers against real respondent data from Dunnhumby's Shopper Thoughts community. Recommends capping synthetic at no more than 5% of total sample, for boosting underrepresented demographics only, and rules it out for segmentation, key driver analysis, and behavioural prediction.
- [24]
STRAT7. Synthetic data: Is this as good as it gets?, 2026.
Synthetic respondents landed 2 to 3 percentage points from headline results generated by real respondents. Willingness-to-pay prices ran 16% above those given by real people. In a price ordering exercise, purely synthetic respondents broke logical order 68% of the time, falling to 32.8% when blended with real data.
- [25]
Kaiser, C., Kaiser, J., Schallner, R., Manewitsch, V., Rau, L. (NIM, Nuremberg Institute for Market Decisions). Leaving Insight to Digital Twins? Promise, Progress and Limits of Synthetic Respondents, NIM Marketing Intelligence Review 18(1), 48-53, 2026.
DOI: https://doi.org/10.2478/nimmir-2026-0008
- [26]
Nielsen Norman Group. Synthetic Users, 2025.
- [27]
Fairgen and Dig Insights. Synthetic data validation, independent study, 2025.
Vendor-reported, with one independent validation across 28,630 respondents.
- [28]
Rival Group. Market Research Trends 2026, 2026.
- [29]
User Interviews. State of Synthetic Users, 2026.
- [30]
Advertising Research Foundation. Original Research initiative on synthetic data, 2026.
- [31]
Advertising Research Foundation. Respondent Verification and Fraud Detection Initiative, 2026.
A separate ARF programme from the synthetic data work, addressing survey fraud, low-quality respondents, and AI-assisted response behaviour. Initial synthesis and research design proposals in 2026, with empirical work to follow.
- [32]
Qualtrics. Synthetic research validation methodology and customer pilots, 2026.
- [33]
FishDog. AI Consumer Panels: The 2026 Buyer's Guide, 2026.
Source of the 2026 cost benchmarks and vendor landscape quoted here.
- [34]
Background that informed this review without supporting a specific claim above. Listed for completeness, and unnumbered because nothing in the text rests on it.
Horton. Large Language Models as Simulated Economic Agents, NBER Working Paper 31122, 2023.
ICC/ESOMAR. International Code, 2024 revision, and ESOMAR Congress 2024 papers, 2024.
Market Research Society, MRS Delphi Group. Using synthetic participants for market research, second instalment in a three-part series on generative AI in research, 2025.
Coverage: https://www.research-live.com/article/news/mrs-delphi-group-releases-synthetic-data-report/id/5126013
Greenbook. The Research Stack Has a New Layer: Synthetic, 2026.
MeasuringU. A Review of Experiments with Synthetic Users, 2025.
Viewpoints AI. Replication benchmark across 133 published marketing findings, 2024.
Kaiser, C., Manewitsch, V., Schallner, R., Steck, L. (NIM). Generative AI in Market Research, 2024.
Companion to the 2026 digital twins piece, and an earlier study. Listed here for completeness; no claim above rests on it.