Simile builds AI simulations of synthetic consumers that predict behavior with 85% accuracy – replacing six-month focus group studies with two-minute runs.
ENTRY ANGLES
Divergence layer tracking Simile prediction accuracy against real-world outcomes · Regulatory behavior simulation for FDA drug approval (NDA filing risk analysis) · Professional expert behavioral prediction in financial services compliance
VERTICALS
CAPABILITIES
Access to expert behavioral data for fine-tuning, Enterprise B2B sales into regulated industries, Simulation validation methodology
One customer stopped a half-billion-dollar strategic move on the basis of a Simile simulation. Another uses the platform to model how analysts will respond to specific phrasing before an earnings call is written. Joon Sung Park built the underlying technology for his Stanford dissertation: a 2023 project called Smallville that populated a game village with 25 GPT-3.5-powered AI villagers who woke up, held parties, formed relationships, and coordinated tasks without explicit scripting. The paper was widely cited as a research curiosity. What it was, in hindsight, was a proof of concept for an enterprise product worth $2 billion in fourteen months.
Simile trains AI agents on real demographic, behavioral, and psychological data to create synthetic populations that companies can query before committing to real decisions. The output looks like market research but runs on a different timeline: a study that would require six months of focus group recruitment and analysis takes two minutes. Park's accuracy claim is specific – 85% of the rate at which people predict their own behavior – and precise in its implications. People are poor self-predictors; a model slightly worse at predicting them for others still outperforms human research on latency alone, which is the variable that actually determines whether insight arrives before a decision is committed to.
CVS Health uses the platform for product concept testing. Gallup – the company whose name is synonymous with human opinion polling for eighty years – uses it as a customer. Deloitte uses it. The $200 million Series B that closed July 2026, raising total funding past $300 million at a $2 billion valuation, was led by Greenoaks and joined by Index, Bain Capital Ventures, and CVS Health Ventures – the corporate arm of a customer investing in its own supplier. Revenue has grown fivefold in five months.
The global market research industry processes roughly $76 billion annually, a figure that understates total spend because it excludes internal insight teams at large corporations. Traditional research has a timing problem more than a cost problem: the insight that matters most is the one that arrives before a decision is committed to, and survey pipelines that require weeks of design and months of fielding rarely do. Simile's proposition is that the timing problem is the primary one, and that most research decisions could be improved with a faster, slightly less precise answer.
The pricing trajectory Park has sketched publicly is worth noting. He predicts that a single simulation session will cost $10–20 million to run and will sell for $100 million within two to three years. Current pricing is evidently lower, since revenue has grown fivefold in five months on a customer base of recognizable brands. The unit economics at that scale are unusual: each simulation requires compute but not human labor, marginal cost per run trends toward zero as infrastructure scales, and the buyer population – large enterprises making expensive strategic decisions – values the output in terms of mistake avoidance rather than research spend.
The most telling customer relationship is Gallup's. Gallup has built its brand on the proposition that you learn what people think by asking them. Its decision to use a platform whose founder believes synthetic panels will overtake human panels within three years suggests either hedging or genuine conviction about the trajectory. Deloitte's participation matters for a different reason: consulting firms that adopt tools internally either abandon them quietly or build service offerings on top of them. A Simile-powered consulting practice would reach enterprise clients that the platform's direct sales team would not.
Simile's credibility scales with its track record, and its track record accumulates privately. The half-billion-dollar mistake averted case is compelling but unverifiable publicly. The cases where the simulation was wrong – where synthetic populations predicted one thing and the market did another – are also private, which is where the platform's long-term commercial risk sits.
Simile has validated that enterprises will pay significant money to replace slow human research with fast synthetic research. What it hasn't built is the divergence layer: a product that cross-references Simile outputs against observed post-decision market data, identifies where synthetic populations predicted incorrectly, and feeds that divergence back into model calibration. That loop is how clinical AI systems earn institutional trust over time, and it's the infrastructure that converts Simile from a research replacement into a decision intelligence system that gets more accurate with every deployment. Simile itself has no obvious incentive to build this – the divergence data would reveal failures as well as successes – which is why it's an opening for a third-party product rather than a roadmap item.
The vertical entry that doesn't compete with Simile's general-purpose engine: regulatory prediction for pharmaceutical drug approval. The FDA reviewer population is finite, their decision logic is extensively documented in public guidance, inspection databases, and Complete Response Letters, and the cost of predicting incorrectly is measured in years of development time and hundreds of millions of dollars. A model trained specifically on FDA reviewer behavior – not general consumer behavior – would answer a different question than Simile does and sell to a different buyer at a different price point. The first customer is a pharmaceutical company in a late-phase NDA preparation where the anticipated bottleneck is regulatory objections to a specific labeling claim or clinical data presentation. The value is priced against the cost of a Complete Response Letter, which sets the revenue ceiling at roughly $50–100 million per customer per filing cycle.