What is synthetic data used for?
Synthetic data is used in two different fields. In software and data science the goal is usually privacy: models are trained and systems are tested on artificial records that follow the same distributions as real customer records. In market and consumer research the goal is speed and access: getting a first answer on how an idea, a price or a message will be received without running fieldwork.
Synthetic Consumer Lab works in the second field. The synthetic data we produce is not a copy of a table. It is the distribution of answers that a defined audience gives to a specific question.
How is synthetic data generated for consumer research?
The process has three steps.
- A sample is built. More than 50 socioeconomic and demographic attributes, such as age, income, city, education and shopping channel, are turned into profiles using statistical distributions and past consumer research.
- Each profile is represented by its own AI agent. The same question goes to everyone in the sample, and each agent weighs it against their own budget and needs.
- The answers are collected and analysed. The result is an output you can act on: a price-demand curve, a preference share, an importance ranking or a summary of themes.
When can synthetic data be trusted?
Synthetic data is only as trustworthy as its comparison with real consumer data. Academic studies show that large language models can reproduce the direction and ranking of human answers in many surveys and experiments, but that this varies by topic, subgroup and question type. Results are best read as indicators of direction and magnitude, not as a single exact number.
Synthetic data is most useful early on: eliminating options, narrowing a price range, deciding which hypothesis to take into fieldwork. For high-stakes decisions, validation with real consumers is still needed.
What are its limits?
Synthetic data is a model output, and it is limited to the world the model knows.
- Error grows in very new product categories and in niche audiences the model has seen little data about.
- Models tend to give answers that are close to the average and a little too reasonable, so extreme behaviour can be under-represented.
- The gap between stated intent and real purchase behaviour exists in synthetic data too.
| Synthetic data | Fieldwork | |
|---|---|---|
| Time | Minutes to hours | Weeks |
| Cost | Low per question | High per participant |
| Repeatability | Unlimited scenarios with the same sample | Each wave needs new fieldwork |
| Level of evidence | Indicator of direction and magnitude | Statements from real consumers |
| Best used for | Early screening, price and message tests | Final validation, new categories |
Frequently asked questions
Is synthetic data real data?
No. Synthetic data is generated by a model. Its value depends on how well it reflects the statistical structure of real data, which is why it has to be validated against real consumer data.
Does synthetic data contain personal data?
Synthetic consumer profiles are not records of real people. They are generated from statistical distributions, so an answer cannot be traced back to a specific real person.
Does synthetic data replace fieldwork?
Not entirely. Synthetic data eliminates options quickly and sharpens hypotheses. For critical decisions, research with real consumers remains the strongest evidence.
Which kinds of research can use synthetic data?
Willingness-to-pay (WTP) and price tests, conjoint analysis, MaxDiff, multiple-choice surveys, focus groups, ad and message tests, and A/B comparisons.
Sources
The assessments on this page draw on published academic work on simulating survey respondents with large language models. You can read our summaries on the articles page. Read the article summaries.
Related pages
Last updated: September 2026