Categories
Learn how to evaluate synthetic research tools with 5 key questions on data generation, model validation, and human insight.
The synthetic research market has exploded over the past 12 months. For many research teams, figuring out which synthetic partner best serves their needs can feel like a daunting task.
The confusion starts with the word itself. Ask researchers what "synthetic" means and the answers splinter: conversational AI, digital twins, personas, generated responses, AI insights. Each term points to a different method, and the method behind the output determines whether the data can support the decision you're trying to make.
It’s one reason why evaluating a synthetic partner must not start with brand recognition or funding rounds, but in your provider’s ability to answer these core questions clearly.
Ask this first. At its core, every synthetic solution uses the same approach: it takes text as input, processes it through a model's learned patterns, and predicts a human-like output. A persona generator takes a general-purpose AI model and gives it detailed instructions to answer like a certain type of person. A digital twin goes a step further, feeding the model a library of real documents or past answers so it can imitate a specific individual. Both add a layer of specificity on top of a general-purpose model that everyone else can access too. That added specificity has a ceiling, though, the model underneath still wasn't built to predict things like how someone answers a survey.
At Qualtrics we take a different path. Rather than wrapping a public LLM in prompts, we fine-tuned a foundational model with the language of research itself. Each survey is a conversation between the survey writer and the respondent, human or synthetic, that captures idiosyncrasies and contradictions that real survey takers make along the way.
Every vendor has a "secret sauce." They should be able to articulate what differentiates them, even if they don't reveal all of it.
Our synthetic model is fine-tuned on anonymized, aggregated survey responses from millions of real survey takers, collected through the Qualtrics platform. As we scaled training, we found that breadth of questions mattered more than depth of respondents. By exposing the model to many different question types, we taught it how surveys and respondents relate, which helps it avoid the repetitive, low-variety answers that plague general-use LLMs.
That tracks with a basic data science principle: a model is most accurate when the population you want to represent looks like the data it was trained on. A trillion-parameter model trained on the open internet can do a lot of things, but reproducing how real people answer a specific survey isn't one of its core strengths. Scale and specificity, together, are what differentiate this kind of model.
Different synthetic applications — personas, conversational AI, digital twins, simulated individual-level data — require different forms of validation. Ask a synthetic vendor to see a direct comparison between the AI-generated output and the human perspective it's trying to emulate. Look qualitatively for similar breadth and depth of response and quantitatively for synthetic and human distributions placed side by side.
In our own validation, Edge Audiences achieved an average deviation of 0.07 standard deviations from human response means (Cohen's d) — a 12x improvement over general-use LLMs on the same questions. But an average can hide a bad distribution, so we also checked KL divergence, mean differences, and top-box percentages, and confirmed that relationships between answers hold up well enough to support correlation matrices and segmentation.

Picking a vendor is only half the job. The other half is knowing which questions to trust synthetic data with in the first place. Don’t pick the vendor that is selling you synthetic for projects where it’s not appropriate.
Synthetic performs best on enduring, attitudinal questions. It's better at "How likely are you to try Brand X?" than "How was your last visit to Brand X?" In other words, likelihood, importance, future intent, and stable attitudes track well. Very specific, recent lived experiences don't.
Synthetic can stand on its own for early exploration, comparing options, message and concept iteration, and fast directional reads, where speed and flexibility matter most. Pair it with human data for high-stakes decisions, recent lived experiences, continuity with an existing tracker or trendline (comparability over time is where human data still adds real value), and specialized audiences underrepresented in the model's training data. The rule is simple: match the method to the decision you're making. That's the discipline a good vendor should be reinforcing,
As synthetic tools take on bigger questions, governance isn't optional. A vendor should be able to explain how data enters the model, who consents to that, and what happens to individual identity along the way, not just point to a list of certifications. At minimum, that means training only on anonymized, aggregated data; keeping customers in control of their own data; and building in privacy safeguards like access controls and the ability to remove data on request.
Your vendor should speak to these governance and security questions with the same clarity they bring to model architecture.
As you talk with vendors, or think about incorporating AI into your own workflows, remember that these solutions should not be about replacing the good work you are already doing. Faster and cheaper aren't the only points a vendor should be highlighting. The real value is using these tools in new ways to reach insights you couldn't get to before. Come with questions, ask for the data behind any accuracy claim, and match the method to the decision at hand. That's what earns synthetic data its place in the researcher's repertoire.
Comments
Comments are moderated to ensure respect towards the author and to prevent spam or self-promotion. Your comment may be edited, rejected, or approved based on these criteria. By commenting, you accept these terms and take responsibility for your contributions.
Disclaimer
The views, opinions, data, and methodologies expressed above are those of the contributor(s) and do not necessarily reflect or represent the official policies, positions, or beliefs of Greenbook.
More from Derrick McLean, PhD
Partner Content
4 min read
Learn how to evaluate synthetic research tools and build confidence in AI-generated data for better business decisions.
Partner Content
7 min read
Qualtrics examines how synthetic data performs against academic benchmarks, addressing trust and validation gaps in AI-driven research.
ARTICLES
Explore why AI is an opportunity for insights teams to reinvent their role, increase business impact, and improve decision-making.
7 min read
Synthetic respondents use AI to simulate survey participants. Learn how they work, when they're accurate, and when real respondents are still essentia...
Explore how cognitive offloading and augmented intelligence are changing market research, creativity, empathy, and decision-making.
Discover what the C-Suite wants from insights teams as Mehra discusses Chime's growth, career lessons, and AI's impact on research.
Sign Up for
Updates
Get content that matters, written by top insights industry experts, delivered right to your inbox.