Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Latent Space: The AI Engineer Podcast
6 DAYS AGO
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Simulation: the new Scaling Law — Joon Sung Park, Simile AI

Latent Space: The AI Engineer Podcast
6 DAYS AGO
Unprocessed episode, you can be the first!
Shownote
Shownote
When we first dicsussed the Summer of Simulative AI
[https://www.latent.space/p/sim-ai] in 2024 we knew it would be a brief summer,
but it has recently come back with a vengeance with SimGym in April
[https://www.latent.space/p/shopify] and now Simile AI’s $2B Series B
[https://techcrunch.com/2026/07/30/synthetic-user-startup-simile-raises-200m-at-2b-valuation-5-months-after-100m-series-a/],
backed by GreenOaks and Index Ventures with prominent backers like Fei-Fei Li
and Andrej Karpathy, running tens of millions of simulations for Fortune 100
clients like CVS [https://www.simile.com/#customers] and 85–99% accuracy
[https://www.simile.com/#research] vs human focus groups.
Time to catch up on why this Second Summer of simulation is working!
From creating Smallville [https://arxiv.org/pdf/2304.03442], the landmark 2023
paper on Generative Agents that showed AI characters could remember, plan,
socialize, and develop emergent behaviors, to now building foundation models of
human behavior, Joon Sung Park is trying to answer a much bigger question: what
if we could simulate the world before making decisions in it? In this episode,
the Simile co-founder and CEO joins us to unpack the path from generative agents
to digital twins, why today’s frontier models still fail to capture how humans
actually behave, and what it would take to eventually simulate all 8 billion
people on Earth.
We go deep on Simile’s approach to modeling human behavior: long-form
interviews, observational and transaction data, randomized controlled trials,
population-level and individual-level models, and post-training on the causal
mechanisms behind why people make decisions. Joon explains how his research
created digital twins that reproduced human behavior and attitudes 85% as
accurately as people reproduced their own responses, why models optimized to be
rational can be bad simulations of irrational humans, and why understanding
“social physics” may require changing model weights rather than simply prompting
frontier LLMs.
We also explore the much larger ambition behind simulation: testing products and
policies before deploying them, finding counterintuitive paths toward desired
outcomes, modeling emergent behavior across entire societies, and potentially
tackling problems like climate change, democratic instability, and UBI. Joon
reflects on scaling laws for simulation, the economics of data-center-scale
simulated worlds, the connection to Thomas Schelling and psychohistory, why
simulation is surprisingly similar to painting, and whether we might already be
living in one.
We discuss:
* How Smallville and Generative Agents led to Simile
* Why Joon’s team asked: “What if we can just recreate the world that we live
in?”
* Why useful personal agents require deep models of their users
* Memory architectures, Markdown files, and the limits of prompting
* “Social physics” and behavioral foundation models
* Why web data captures what people say more than what they actually do
* Interviews, transactions, observational data, and randomized controlled trials
* Why predicting the future matters less than understanding how to shape it
* How Simile creates representative simulated populations
* Simulation versus prediction and the connection to Foundation’s psychohistory
* How to evaluate simulations instead of simply stacking LLM hallucinations
* Creating digital twins of 1,000 real people and reaching 85% behavioral
accuracy
* Why frontier models can struggle to reproduce real human behavior
* Why good simulations need to reproduce human biases and mistakes
* Post-training models on randomized controlled trials
* Population-level versus individual-level simulation
* Scaling laws for human simulation
* The long-term ambition to simulate all 8 billion people on Earth
* Whether simulations could help solve climate change or detect collapsing
democracy
* Thomas Schelling and the history of agent-based modeling
* Why future simulations could require an entire data center
* Multi-agent simulations and what happens when simulated people interact
* Replacing expensive human panels with synthetic populations
* Why market research is only the starting point for simulation
* Why Joon sees simulation as surprisingly similar to painting
* Using simulation to study questions like UBI
* Whether we are already living in a simulation
* Why AGI and simulation may be the twin technologies of advanced civilizations
Joon Sung Park
* LinkedIn: https://www.linkedin.com/in/joonspark
[https://www.linkedin.com/in/joonspark]
* X: https://x.com/joon_s_pk [https://x.com/joon_s_pk]
* Website: https://www.joonsungpark.com [https://www.joonsungpark.com/]
* Simile: https://www.simile.com [https://www.simile.com/]
Timestamps
00:00:00 Introduction and Joon’s Path from Art to AI
00:01:46 Smallville, Generative Agents, and the Origins of Simulation
00:05:03 “Let’s Just Create a World” and the Future of Personal Agents
00:09:53 Social Physics and Behavioral Foundation Models
00:14:08 Prediction vs. Simulation: How Do You Shape the Future?
00:16:59 How Simile Models Real People and Populations
00:25:35 Evaluating Simulations, Digital Twins, and 85% Accuracy
00:30:23 Post-Training Models to Reproduce Human Behavior
00:40:04 Scaling Laws and Simulating 8 Billion People
00:43:10 From Schelling to Society-Scale Agent Simulations
00:46:13 The Cost and Economics of Simulating the World
00:52:05 Real-World Use Cases, Synthetic Populations, and the Market
00:57:27 The Future of Simulation, Painting, and UBI
01:04:23 Are We Already Living in a Simulation?
01:06:08 Building Simile and Hiring
Transcript
Introduction: Joon Sung Park, Simile, and the Story So Far
Vibhu [00:00:00]: Today, we have Joon in the podcast. Excited to kick this one
off. Very exciting company. I wanna kick off and ask you the question, talk us
through the story of your life. How have you gotten here?
Joon [00:00:13]: Yeah, for sure. I’m really excited to be here. A story of my
life. So I was born in Korea, and I lived there for a good 11 years or so of my
life, and then my family moved to Boston. So we moved when I was 11, and my
parents were doctors, so they were going through their postdoctoral studies. My
dad was a surgeon, so he was doing his sabbatical years at the Boston Children’s
Hospital. So I grew up there, not too close to tech. I was very much a music and
artsy, painting kind of guy.
Vibhu [00:00:49]: Painting.
Joon [00:00:49]: Exactly. I got into painting a little bit later, in high
school, but that’s what I used to do. And then I grew up mostly in the East
Coast after Korea. So I lived a good number of years in New Hampshire, and then
I went to college in Pennsylvania. And I got into more of this tech scene, in
college. So I was originally trained to be an artist. I thought that would be my
professional career. So it wasn’t a hobby. It was like, “Hey, let’s make a
living out of this.” And then gradually, I got really interested in this idea
of, hey, the greatest artist often creates their own medium, and the best medium
that we had available today was in computation. So I decided to go deeper into
that, and one thing led to another, and we can go deeper into this, but I
decided that research was something that I gradually got interested in, and here
I am.
Smallville, Generative Agents, and the 2023 Breakout Paper
Swyx [00:01:46]: So there’s a lot that you packed into the research components.
You had one of the best papers of 2023, which was the generative agents paper,
commonly known as the Smallville paper.
Swyx [00:01:58]: Feel free to call back to anything else that you mentioned, but
most people would have heard of you from this. Do you have any statistics on how
many people have, like, read it? arXiv gives you something, right? Some stats.
Joon [00:02:10]: Yeah, it’s a good question. How many people have read it, I’m
not sure.
Joon [00:02:14]: I know we do keep track of citations, and they are going up
quite fast.
Swyx [00:02:23]: Yeah, Google Scholar has 7,200 citations.
Vibhu [00:02:25]: I feel like it made a bigger hit than that, and it was a
pretty instrumental paper. It got cited so many times.
Swyx [00:02:34]: It is frequently the answer when people ask, “What is the best
paper you’ve read recently?” It’s this one.
Vibhu [00:02:39]: I thought the memory component was pretty underrated. It was a
very good early memory system, and one of the biggest papers.
Foundation Models and the Search for Killer Applications
Joon [00:02:47]: Yeah, so maybe I can talk a little bit about how this
particular paper came together. So when I got into research, it was back in 2020
when I started my PhD program at Stanford, and that was the year, when we were
about to get GPT-3 to be available. So we already had GPT-2, and you could sense
that there was this new class of models that was just becoming available in the
market, and the team got very intrigued. And the general consensus was, “Well,
is this model going to be useful for anything?” “It’s really strange that these
models are not trained to do any particular task.” But we decided to take a bet.
So a large group of scholars at Stanford, and it was led by one of my
co-founders, Percy Liang, and we came together
Swyx [00:03:35]: Who coined foundation models.
Joon [00:03:36]: Who coined the term foundation models. We wrote this paper,
where that term came from called Opportunities and Risks of Foundation Models.
And during that process, really the thing that I started to think deeply about
was, here is a model that is fundamentally new in our ecosystem. The reason why
this was new was it wasn’t, again, trained to do anything in particular, but its
premise was it could do anything and everything. It was like a stem cell, if you
were to take a biology analogy. And I got really interested in this idea that,
well, if we were to really think about what are the killer applications that
this particular technology would enable, what would that be? Many of my
colleagues were using this for simple classification, simple generations.
Interesting that these models can do that, but from an interaction perspective,
not that interesting. We’ve known how to do that for many decades. And what we
came down to was these models are trained on this very broad data from the web,
right? So these are human behavioral data. It’s social media, Wikipedia, all
these data. So if you poke at the right angle, then you could see human behavior
that would just pop out that’s quite realistic, and we’ve never seen that
before.
The Time Machine Game and Recreating the World
Joon [00:04:45]: So that got us really interested. The exercise that we decided
to do, with this particular group of colleagues, Michael Bernstein, Percy Liang,
and myself, who ended up becoming my co-founder at Simile, we sat down and we
played this game that we call the time machine game.
Joon [00:05:03]: Imagine we were to get on a time machine and fast-forward 10
years and look back. What would have been the single application that will have
mattered that would be the most interesting and inspiring? And when we thought,
“Well, what if we can just recreate the world that we live in?” it’s really hard
to get more ambitious than that. Like, let’s just create a world.
Joon [00:05:24]: And that’s where we started. And initially, we had this paper
that was a precursor to the generative agents paper called Social Simulacra.
Swyx [00:05:32]: Before you go further, were there other candidates for the most
ambitious thing in the time machine exercise? What was number two or number
three?
Personal Agents, User Models, and Why Simulation Came First
Joon [00:05:44]: There is a close second that we were considering, which ended
up becoming more of these automation tools, especially the vision around really
personalized agents that would do things for you.
Swyx [00:05:59]: That’s also happening.
Joon [00:06:00]: It’s also happening. But it was interesting for us, right, in
that the reason why, we decided to go with the idea of simulation, one, I was a
huge science fiction nerd, and this idea of creating simulation, I was
personally really just fascinated. I loved the idea. It’s really cool to see,
like, a game town like this and just see these agents live in it. But at the
same time, my bet was if you were to create a really amazing personal assistant
out of this technology, what you need first is an amazing model of your users.
So I told a model, “Hey, can you go buy late dinner for me?” And it orders
Hawaiian pizza, and I do not like pineapples on my pizza. Then it totally
failed. The way for it to not make that mistake is only by having a deep
understanding of who I am. And I gave a very simple and dumb example here, but
you can imagine how this core understanding of people is instrumental. This is
how, if we have our family and closest friends, they have a good mental model of
who we are. That’s the basis of our social connection. So our bet also was this
technology around simulation, creating accurate representation of people ought
to precede the more complex agents that would automate the world that we live
in. So that was the bet. But that was a very close second, and I’m still very
much fascinated by it. I think there’s a lot of interesting work that’s going
around. My hot take here, though, is I don’t think we’ve seen a true personal
assistant that’s useful, in ways that meet the ambition of that particular line
of work. I think there are early applications that are interesting, and if you
talk to even ChatGPT nowadays or Claude, they know a lot about us. So a lot of
the generation it’s doing, I do think it’s much more tailored, but I think the
ambition is quite large in that field, and I don’t think we quite have all the
right ingredients just yet.
Swyx [00:08:01]: So OpenClaw and these personal agents, what do you want to see
from them that they don’t currently have?
Memory, Markdown, and the Limits of Prompting
Joon [00:08:09]: I do think it’s slowly getting there, but I do generally want
them to have much deeper understanding of the person. Right now, you look at the
models. OpenClaw, what it’s leveraging is a Markdown file, and I think it’s
quite clever, right? So if you look at the generative agents paper, this was the
same intuition that we had, where initially when we were creating the memory
architecture for the generative agents, and, like, this is, like, back in 2022,
so we didn’t really quite have the idea of even agentive architecture or the
term agent. But the intuition that we shared with some of the work that’s coming
out today was we initially thought, “Well, do we want to make the memory into,
let’s say, knowledge graph? Do we want to train a bespoke model?” All of these
things. And what we decided to do was, “No. Just forget about all this.” These
language models are quite good at modeling text and understanding and reasoning
about text. So just put everything in a Markdown file or a text file. You’re
done. I thought that was quite interesting that we could do that, and there’s a
lot of strength in doing that. But also, there are limitations. It’s the way you
retrieve and make sense of data that’s extremely large, it takes a lot of work.
So I think that technology is getting better. I also do, however, think, there
are certain things you just cannot shape just by prompting the model. So to some
degree, you do need to touch the parameters of the model itself. So there is
this work that I do think does need to happen, and it is happening. The question
is, how far can we take it? How do we source data, and how do you also create an
ecosystem where people are continuously feeding data to this model so it’s
learning about you?
Vibhu [00:09:50]: What’s the intuition between why you need to do it in the
model?
Social Physics and Behavior Foundation Models
Joon [00:09:53]: My intuition behind the actual when do you train or even
post-train a model versus just prompt a model is if the model has to learn the
underlying physics of the world that it’s operating in. So it has to learn new
social physics. The places where it doesn’t have to train are the places where
it already has the physics. We trust the physics. It already has the base
statistics, but it’s just trying to react to an environment. Then I think you
can just prompt your way into getting the actions out of it. I don’t think the
models that are out in the open have yet learned the complete mapping of social
physics of humanity. This is one of the core theses of Simile, right? And one of
the core reasons why that is the case is if you look at the data that the model
was trained on, these models were trained on the web data, like, whatever was
available on the web. And these are really interesting data sets, but they are
fundamentally the self-exposed attitudinal data with some behavior data that’s
sprinkled around here and there. And it has yet to learn the really deep
behavioral nature of people, not just what people say they do online, but what
they do in real life. And this is one of what I would consider to be the dark
knowledge of humanity that we haven’t quite captured. And it’s these data that
would also need to get factored into the model creation.
Vibhu [00:11:21]: You call it behavior foundation model.
Vibhu [00:11:23]: There’s a good one-liner here, but outside of that, what type
of data do you need? What are you changing on the model level? How do you go
about modeling, doing a behavior foundation model?
The Three Data Buckets: Interviews, Behavior, and Causality
Joon [00:11:35]: We think about data in three buckets. So one bucket is
interview data. It’s quite interesting. Rich qualitative data is interesting.
It’s not behavioral, but we would literally ask people, “Hey, tell me the story
of your life.”
Vibhu [00:11:53]: It’s just what we’re doing here exactly.
Joon [00:11:54]: The question that you all asked at the beginning of this
interview literally is the question we also ask. And we ask our participants to
go a little bit deeper, than how far I went. Maybe I can give more of my life
story in lieu of this. But the reason why that data is interesting is by
learning about this very long-tail information about people, you get a lot of
texture around this model, like, this person as a model. So even understanding
their childhood memory or even their trauma, their first love, these things,
quite informative in ways that’s really hard to predict. So that’s one. Then
there are two tranches of what I would consider to be the behavioral data. One
kind of behavioral data is observational. So these might be like transaction
data, or these might be data that you can get by scraping the web, right? So you
can imagine why these data sets would be interesting, right, because they give
you the base statistics of people’s behavior.
Joon [00:12:55]: But then there is the last category of data, that I personally
think is perhaps the most important, which is the data that describes the causal
mechanism, the whys of people. Some of this is covered by the interview data,
the qualitative, because people talk about why they made certain decisions. But
really, where you get to see the most behavioral aspect of this is in randomized
controlled trials, like RCTs. Imagine you have the same setup, but you have a
few different variables that you are trying to tweak. Can you get realistic
human behavior out of it in ways where, imagine you had this particular option.
Imagine you’re even trying to choose whether you’re going to drink coffee or
not. The day you drink coffee versus the day you didn’t drink coffee, does your
behavior change? That’s a data set that describes a causal mechanism. This is
quite important in modeling people. The reason why this is important is
oftentimes when people come to us, or not just to us, but the reason why people
are interested in simulation isn’t because they want to predict the future. If
you’re trying to win against the stock market, predicting the future is
interesting.
Prediction vs. Simulation: Shaping the Future
Joon [00:14:08]: But most people, most decision-makers, what they want to know
is, how can we shape the future? It doesn’t really help you to hear that your
sales are going to tank in two quarters. They’re just gonna say, “Wow, that
sucks.” What they want to know is, well, what do we need to do now to avoid that
future? That’s the causal mechanism. And this is also very hard data to come by,
right, because the world is our ground truth, but it happens once. So in a very
controlled setup where everything is equal except for one variable, this kind of
data set rarely happens. So this is a reason why this data set is both hard to
come by and quite important if you’re trying to model human behavior.
Swyx [00:14:50]: So behavior, I think, is the hardest data set to acquire. What
is out there? What is even possible? You’re not going to know a lot of details
about my life. I don’t even have data for myself on my own health or habits, and
I just don’t log everything. So how can you have that data?
Joon [00:15:14]: So we run a lot of randomized controlled trials.
Swyx [00:15:17]: But you put people in the lab, they watch them sleep, or what?
Joon [00:15:20]: We do care a lot about the consent process. People know that we
invite them to be a member of this community to both share data and have
themselves represented in different forms. But we bring a lot of people to the
lab, or virtual lab, where we design experiments that would pose them real
behavioral decisions. And often in these experimental setups, what makes the
difference between what is attitudinal versus behavioral is whether the stake in
your decision is real. That’s ultimately what makes it behavioral. So in these
setups, we are inspired by our colleagues in social sciences, psychology, and so
forth. So when they run studies, the techniques they utilize is imagine there’s
an online store that you’re inviting people to come by. Then whatever they
purchase in this experiment, they actually get that item delivered. Like, these
are the things that make the stakes real. So we run a lot of these experiments,
and we also do partner with firms. Right now, we also have customers who are
quite excited to at least give us a glimpse of the behaviors that their users
exhibit so that we can get a little bit deeper understanding of how people
behave in these different platforms.
How Customers Use Simile: Populations, Queries, and Experiments
Vibhu [00:16:39]: I think on the customer side, they have a lot of data about
their users, who has bought. They have the action data.
Vibhu [00:16:47]: Can you walk us through an example of what someone comes to
you for? What questions would they want solved? Do you customize a model for
them? Do you have something off the shelf? What does that look like?
Joon [00:16:59]: Today, when people leverage our models, it’s often to better
understand the population of their interest. So usually, the start of the
relationship, we come together and hear about what population they want us to
model, right? So it might be that if you’re a CPG company that’s selling to all
of the US, then maybe it’s fairly straightforward. You want to model the gen pop
of the US. But at the same time, if there is a vertical or if there’s a market
that they’re trying to go into, imagine, they want to better understand, let’s
say, people in their 20s and 30s living in California. That’s a much more
specific population. So we hear about this population, and we go recruit these
people, with consent, and with incentives, and we collect some of their data and
create a model of these people. Then what our product allows you to do is query
them. So it can take as input a filter that is a description of the population
that you want to talk to, just like the one I just mentioned, and an
environment. The environment can literally be survey questions, behavioral
experiments, It can be A/B testing. Oftentimes, the core use cases are things
like concept testing, to start with. But also, people sometimes want to do focus
groups or one of the fun use cases that we also serve is even modeling things
like earnings calls for public companies.
Joon [00:18:21]: So these are the use cases that we often start with.
Swyx [00:18:23]: Concept testing, is that an established term? I’ve never heard
of concept testing.
Concept Testing, Gallup, and Politics
Joon [00:18:27]: Yeah. So it has to do with they have, let’s say, different
messaging, different products, different ideas.
Swyx [00:18:32]: It’s like a marketing exercise.
Swyx [00:18:33]: Okay, got it. Got it. Politics?
Joon [00:18:36]: We do, have a strategic partnership with Gallup, and of course,
Gallup is deep into policy space and so forth. Right now, we have not worked
deeply with politics, like that area just yet, however.
Swyx [00:18:49]: I’m curious if there is demand or if they really would have
different needs that somehow fundamentally don’t mix with your existing, users
or people.
Joon [00:19:00]: I think there’s certainly demand.
Joon [00:19:02]: But we are very much mindful of how this technology gets
adopted and the societal impact that we’ll end up having with this technology.
And I do see politics as an area where a company has to be particularly
thoughtful about the way they operate and make impact. So this is where we also
want to make sure that we form enough of guardrail and perspective on how to
leverage this technology before we go on to serve markets like the politics.
Swyx [00:19:29]: I’ll give people an example. one of my favorite shows is The
West Wing. I don’t know if people have watched.
Swyx [00:19:34]: One of the key storylines is, like, the president has, multiple
sclerosis, but they haven’t. they need to figure out how to disclose it. So they
run a poll with a fake governor and ask people to respond on the poll,
Counterfactuals, Polling, and When Simulation Is Useful
Swyx [00:19:47]: They try to make decisions based on the results of that poll
on, like, how well they’ll be received, like where, how should we play this?
Swyx [00:19:54]: And I’m like, well, I think those counterfactual things, I
would use a simulation for this if I could trust it.
Joon [00:20:01]: For sure.
Joon [00:20:02]: In that show, how’d it go?
Swyx [00:20:04]: In that show, it was, like a foregone conclusion. They were
like, “We know it’s bad. We just don’t know how bad.” And then the poll came
back. It was like, “It’s really bad.” And then they just did it anyway.
Joon [00:20:14]: Part of it is to show, right? So you’re, you’re looking at the
idea
Swyx [00:20:17]: Maximizing drama.
Joon [00:20:18]: How bad could it be? Oh, it’s horrible.
Swyx [00:20:20]: And to some extent, I think that is part of the trick of the,
or the challenge or with being a customer of yours, which is that if I know
it’s. if I roughly know and can intuit
Swyx [00:20:35]: What the effect is going to be, do I need you? What sensitivity
of it, of effect do I need in order to make a decision, right? So for example,
if I, my approval rating is 50%
Swyx [00:20:48]: And I, they have this negative piece, news item comes out, and
it drops to 30.
Swyx [00:20:52]: If it drops to 20, if it drops to 40, do I care? No. It, I know
it drops. It’s negative. So when do I care about simulations?
Joon [00:21:01]: You do something that’s clearly bad, that’s not popular, and
people don’t like you, like, yeah, it’s like
Swyx [00:21:05]: You don’t need a simulation.
Joon [00:21:07]: Yeah. Well, so there are a couple of things. one is, there are
use cases where, like every day, developers, designers, policymakers, marketers,
every single day, they create assets. They create new products. And turns out,
it’s many of the decisions in hindsight is obvious. Yes, of course this is bad,
but we still run those studies because understanding the magnitude and
understanding how acute something is quite difficult, even if, we feel like, of
course, like this makes sense. this is the reason why we make so many mistakes.
Like, every time somebody goes online and say something that has huge backlash,
you look at that and like, “What an idiot.” However, it’s tough. That’s one.
There’s also another aspect here, which is, again, this is the reason why
simulation is different from prediction. In simulation, in the ideal case
scenario. So what simulation is trying to show is it’s trying to show each step
of the way or each step that we need to take to get to a certain outcome, right?
So in the most advanced simulations, sometimes the next step that we’re
suggesting might be quite counterintuitive. The analogy that I sometimes give,
and I ground it in a more realistic example, but, I, as I mentioned, I’m a huge
fan of science fiction, and I don’t know how, many of the audience members have
read, like, things like the Foundation series by Asimov.
Simulation as a Path, Not Just a Prediction
Swyx [00:22:37]: Oh, yeah. We’ve mentioned psychohistory a number of times.
Joon [00:22:39]: Okay, fantastic. So I might be, talking to the right crew. If
you read Foundation series, literally the first act is there’s a group of
scientists who have found out that, “Oh, our galactic empire is going to
collapse, and we’re going to have 30,000 years of unrest.” And they run
psychohistory, the simulator that tries to teach them, “Okay, how can we keep
this unrest to a 1,000 years?” And they plan this out, and the first step of
that plan is to get the scientists who say, “Okay, this is coming,” exiled into
this random place in this, galax- galaxy.
Swyx [00:23:18]: Terminus.
Joon [00:23:19]: Exactly. And that’s so counterintuitive. Like, what a strange
move that you literally sent the group of scientists who was raising voice
around this potential collapse of galactic empire into nowhere. How is that the
right first move? Well, it turns out in this particular simulation, that was the
move.
Joon [00:23:40]: It’s these things, right? And the reason why these reasoning is
possible is because you’re showing the step function or each step that results
in a particular outcome. So really what simulation allows you to do in its
highest form is you give it not a problem or question, like what would people
answer to the survey? That’s not what we do. What we tell it is, “Here is a goal
that we have. In the context of foundation, we want to keep the unrest to a
1,000 years. What is the path that we need to take now to get to that particular
future?” And that’s what simulation allows you to do. Now, translating that into
real market, imagine you’re a automobile company and you’re about to release a,
EV, and you’re trying to understand, well, how do we market EV, to make sure
that our stock price goes up? But what if the answer comes down that, well, you
can market your EV in XYZ way, but that might change people’s perception around
the cars that’s not EV and make your overall sales to go down. Not very
intuitive, especially all you’re trying to optimize is EV salesss, and that’s
the only thing that you’re tracking, then that might result in a completely
wrong solution, or at least different solution than what you would have
expected, whether it’s right or wrong.
Joon [00:24:57]: That’s the power of simulation.
Swyx [00:24:58]: For listeners, we covered a similar topic with Mikhail Parakhin
from Shopify, where they are working on SimGym. I don’t know if he ever talked
to you about it. it’s very similar.
Joon [00:25:07]: I
Swyx [00:25:07]: The goal is increased conversion, but then the journey is very
unusual.
Joon [00:25:12]: Journey is unusual.
Swyx [00:25:12]: Yeah. The-- He’s trying to look for interventions on a shopping
trajectory, which is similar to what you’re saying. Like, it’s not about the
attitudinal, is your word for it.
Swyx [00:25:24]: It’s about behavior.
Joon [00:25:25]: It’s about behavior.
Swyx [00:25:25]: And that’s exactly the difference, right? It’s, like, not about
the near-term direction about-- but it’s more about, like, how do you affect
multiple turns of interactions.
Vibhu [00:25:35]: You had a good quote at the start about this as well. It’s not
about people wanting to know the outcome. It’s about how they can change it,
change the way to get there, something like that. But I wanna take it back to
how do we know this is grounded? Like
Grounding and Evaluating Digital Twins
Vibhu [00:25:47]: How do you run evals? How do you test that simulations come
through? if I was to do the same thing that you described with, say, your
favorite LLM, Opus, GPT-5.6, have some agent to map out these things
Vibhu [00:26:02]: How different are the answers we would get if I give it the
same goal, the same objective, make a decent system? You’re saying that you need
to change the model weight. You have your own solution to this. But how far off
are we, and how do you check if it’s grounded? you have some interesting stuff
on your site that points to how you run real evals, but if you could take us
through that side. I think that’s one of the big concerns that people have.
They’re like, “LLMs hallucinate.”
Vibhu [00:26:27]: “You’re just hallucinating layer after layer,” right?
Joon [00:26:30]: The way we do this, and this is the paper that we worked on
after the generative agents paper that really became the, at least for Simile
and also the field of simulation and synthetic panels, really became the
foundation. Yeah, this is the paper. the paper is called Generative Agent
Simulations of 1000 People. Here’s what we’ve done. For this paper, we brought
1,000 people that’s representatively sampled from the US to a virtual lab. And
what we have done was we spent two hours collecting fairly wide-ranging data. In
this particular study, we focused a lot on this interview data, that was, whose
script was taken from this project called American Voices Project. And then we
would also pair that with a lot of behavior data and so forth, whatever we can
collect within two hours. And then we would send these people away for a couple
of weeks. And during that time, I would use this data to create their digital
twins. And I would bring the humans, participants back after 2 weeks and have
them complete a battery of surveys, experiments, behavior studies. So we have
the list here, which included things like behavioral economics games. We would
run literally, like, Big Five personality test, General Social Survey. We would
also go ahead and run the randomized controlled trials that were published on
PNAS. And we would have their digital twins predict how the source individuals
would have acted in these studies and surveys. And this is where we could
replicate people’s behaviors and attitudes 85 percent as accurately as people
would replicate their own. So that was the first really paper that gave this
validated results that we can model individuals in an accurate way. And what we
ended up finding now, of course, in AI space, so this paper came out at the end
of 2024. AI space, a year and a half, 2 years, that’s a lifetime.
85% Accuracy and Why Frontier Models Miss Human Behavior
Swyx [00:28:24]: Yeah. Just, for listeners who are not seeing the YouTube, I
just wanna say, like, the headline figure is 85 percent accuracy, like, which is
a big improvement over all the other
Swyx [00:28:34]: Methods that you showed.
Joon [00:28:36]: But the part that was particularly striking to us, especially
as we improved this technology even further, was the generative AI models like
ChatGPT, Claude that’s coming out, it does give you the right foundation.
However, what they do not consider is the true attitudinal and behavioral aspect
of people, especially in the population that you care about. So what these
models are really good at today is they’re trying to become the super rational,
objective machines, right? So you go get their data from places like Mercor,
Scale. You talk to professional programmers, scientists to create model that’s
amazing at reasoning. That’s what they do. Simile doesn’t care about any of
this. The models that we’re talking about here, what we’re trying to create are
models that are as dumb as I am, right? So if I make some mistakes, the model
has to make the same mistake.
Swyx [00:29:34]: Oh, that’s very hard.
Joon [00:29:35]: That’s very hard.
Swyx [00:29:36]: You’re solving Murphy’s paradox.
Joon [00:29:37]: That’s exactly. And this is a completely different data and
training objective. This is also where we see quite a bit of discrepancy in the
performance in human behavior prediction between the frontier models, Simile’s
model, and the models being created in this space, where in some cases, the
model performance of frontier models go all the way down to 20, 30 percent,
especially if you go into that more niche population on topics that our
customers would care about. On more gen pop, it might be around 50 to 60
percent. So it’s not very robust. Like, you wouldn’t want to make your decision
off of these and these findings. If you can bring that up to 85 percent, that is
ultimately what people end up getting very excited about.
Swyx [00:30:20]: Yeah. Do we wanna keep going on the paper, routes?
Joon [00:30:23]: Yeah, for sure. So the last one, was an interesting one. So
this, paper was the follow-up paper that we had, to the 1000 agents paper, where
the idea was now can we augment the models even further and post-train a model
based on a lot of randomized controlled trials? So this was an interesting one.
The data is always the most interesting part of modeling in many ways. The data
that we got here was there’s this, there’s this platform called Open Science
Framework. So some, the audience might be familiar with this. And there has
been, especially in the social sciences over the past 5 years or so, there has
been this concern around replicability of studies. And so it was a bit of a
crisis, the scientists acknowledged, where we rerun the study and we don’t see
the same finding.
Post-Training on RCTs and Replication Studies
Vibhu [00:31:12]: Oof.
Joon [00:31:12]: It’s tough. And the reason why it’s there-- that was often the
case was there’s this survival bias where the papers that get published often
need to maintain what we call the value of less than 0.05 in the experiments
that we ran. That suggests that only-- there’s only 5% chance that the results
that we saw is false positive. But the tricky part was all the papers that were
not published, and there’s still a 5% chance that whatever we publish is totally
just randomly generated. Like, there’s a 5% chance that, hey, this effect is not
real, but it just happened to be real because of the sampling bias. So because
of that, what scientists started to do was they started to register their
studies. So before running an experiment, they would go to this platform and
say, “Here is the data. Here is the population that we’re collecting, and here’s
the hypotheses.” And they would just say, “Here is our hypothesis.” Like, “This
is what we believe.” And you cannot retroactively change those hypotheses. This
is what gives us more scientific statistical confidence that whatever effect
that you ended up seeing is true. So that ended up creating this really
interesting platform where there’s one platform that has now contains tens of
thousands of real-world experiments and hypotheses. And a lot of these are
really high-quality, like, professionally designed behavior studies and random-
randomized controlled trials. So we got the data and the studies from this
platform and used that to make a point. And this particular, model is not,
something that we’re serving commercially because this was a part of the open
science. But this particular data set, helped us make a point that by collecting
a lot of these randomized controlled trials, that are really well-designed, we
can make significant improvement in model’s capability to predict human
behaviors. So that’s what this paper was about.
Vibhu [00:33:10]: Is this stuff done on a individual level? Like, do I need to
tune the model per individual, per company? Is there foundation model changes
and then some slight post-training? Anything you can share there?
Population-Level vs. Individual-Level Models
Joon [00:33:21]: So this particular model was trained. the data we had at the
level of individuals, but this particular model was trained. We experimented
with both. And this is what we end up doing at Simile too. We always train 2,
distinct model. One is what we call the population-level model. The other is
what we call the individual-level model. And both take very similar input, which
is the description of a subpopulation or individual and a stimuli. In this
particular work, we’ve done the same. Here, the results that we are reporting
are much more geared towards individuals because we do think that is a harder
task in many ways, but that’s what we have done.
Vibhu [00:34:02]: You seen anything on the questions that humans can solve that
models can’t solve? So like
Human Biases, Mundane Choices, and What Models Miss
Vibhu [00:34:09]: Currently, it’s, I live 5 minutes walk away from a car wash.
It’s a 10-minute drive. Should I walk or drive?
Joon [00:34:16]: Huh.
Vibhu [00:34:16]: The model will say, “Oh, walk to the car wash.” And, you don’t
have your car.
Vibhu [00:34:20]: Is anything like this a problem in simulation? You would
assume, like, very simple for human to think about, but if the model is saying
you should walk to the car wash, anything here?
Joon [00:34:32]: It’s less, what can we solve, but I think it’s more about what
biases or mistakes do people make that models miss. Like, imagine that you are,
like the. When I was still at Stanford, I lived in Palo Alto. So it’s about, I
would say, 40-minute walk from the campus. You ask the model, “Okay, let’s go
home. What can I, what can I do?” It would likely call an Uber or, give me, the
bus time. But for the longest time, I really liked walking back. And the reason
why I wanted to do that was not for efficiency. It really helped me think. And I
like to walk for, half an hour or 40 minutes or so a day, where I just get to,
just think about ideas, research, just get lost in my thoughts. That’s very
human activity. Unless the model has seen that and understands the importance of
that activity, it would miss these kinds of features. So that I think, is
fundamentally what we’re trying to model. Like, what is fundamentally human
might not be the most efficient thing to do, might not be the right thing to do,
but things that make us who we are.
Swyx [00:35:43]: I’m curious if, there are some data sets that you really want
that would materially help you. One version of this may be interesting, which is
more valuable to you to acquire as a data set, all of LinkedIn, all of Twitter,
all of Facebook?
What Data Matters: Social Media, Transactions, and Facebook
Joon [00:35:57]: It’s a little bit hard to rank, in part because, there’s,
there’s this product saying where no feedback is wrong because it teaches you
something about your users. Doesn’t matter what feedback.
Joon [00:36:11]: I think it’s a little bit like that.
Swyx [00:36:12]: So just whatever is bigger.
Vibhu [00:36:13]: What about a different domain? Say it was. What about all of
Amazon data?
Joon [00:36:17]: Oh, yeah.
Vibhu [00:36:18]: Shopping data, right?
Joon [00:36:18]: Shopping data. So Amazon data is interesting in that it’s very
much behavioral, although, like, what people do on social media, you could
squint and say that is also behavioral. But the transaction data is always
interesting. It is also most commonly available, however.
Joon [00:36:33]: If we were to look at purely social media, like if you really,
if I were, if I had to really pick, Facebook likely is interesting because I do
think it is most a default version of people. Because you go to LinkedIn, it’s
very much professional environment. So people put up their, they have their
guards up, right? And that still is interesting because that is true human
attitude and behavior, but it is not your base state. you go to Twitter-
Twitter, people have their own crazy personas, or depending on who you are.
Like, my Twitter profile and, persona is very much, initially was I was very
much an academic. “Hey, I’m here to share my studies.” Now, I share, things
that’s related to Simile. But Facebook is one of those more private space where
people just connect with their friends. In that way, I do think it shows you a
little bit more about who that person is. So if I had to pick, I’d likely pick,
Facebook.
Swyx [00:37:30]: Yeah. And you’re interested in, like, the whole person and
their background and philosophy. I, is it too clinical or too machine
learning-oriented to just say this is just ways to inject variance and biases?
The broad question, is, like, is this any better than a randomized, like,
combinatorial explosion version? So we have a link to the Tencent
Billion Personas, Synthetic Demographics, and Bespoke Data
Swyx [00:37:54]: Billion persona paper, where they did not do any of the
groundwork that you are doing.
Swyx [00:37:59]: They just did like a cross matrix of here’s all the professions
in the world, here’s all the people, possible backgrounds in the world, do a dot
product across all of them, and that’s it. That’s your prompt for a billion
people.
Swyx [00:38:12]: This will do something. I don’t know if it’ll do what you do,
but it gets you some way, some percent of the way there.
Joon [00:38:18]: So this was an interesting paper. Like, what I admired about
this paper when it came out was the scale. And you do gradually want to be able
to simulate really large societies and interactions. So the scale is definitely
admirable. it is relying heavily on the known statistics that went into training
the model. So to the extent that you believe that statistics is correct, this is
not a bad way to go about this. But the thesis here, and this is something that
we also have seen in the market, like if this works, then we have solved
simulation.
Joon [00:38:54]: It,
Swyx [00:38:55]: Because I survey, like, okay, 5% of the US population is in
construction.
Swyx [00:39:01]: The other 5% is in medicine, whatever, right? And then you just
keep going down the list, and then you do the other side. 5% has, like, the big
5 personality
Swyx [00:39:08]: Of, like, neurotic or whatever. That’s it.
Joon [00:39:11]: That’s it. So if you believe that the underlying data set and
the platform that we’re leveraging has all the right statistics, then this will
have solved it. you’re at that point merely retrieving the knowledge that is
already embedded in the model, in the model parameters. That’s not,
unfortunately, what we see, where there is such detailed and also niche
knowledge about people that if you just take one example, it might feel very
mundane, but it’s quite rich when you put together, that you do need to do a lot
of bespoke data collection to better understand people. And this is also, I
think what makes this particular, job fun, which you want to deeply understand
people, and the process of deeply understanding them requires a lot of attention
to the details. And you do need to pay attention to and pay respect to the daily
lives that people lead.
Scaling Simulation: From Thousands to Societies
Vibhu [00:40:04]: I wanna talk about scaling simulation.
Vibhu [00:40:07]: So what can’t we simulate, what can we simulate, and how does
scaling affect this? So how big are the models? What if we go from, 8B, like,
couple 100 billion
Vibhu [00:40:18]: Like billion000 parameters, billion000? Do we get scaling? Any
interesting emergence? Like, at a certain scale, at a certain amount of
training, you uncover anything unusual and any learnings from that?
Joon [00:40:31]: What we are seeing is at Simile, so we do post-train our own
model. The thing that we’re seeing is the early glimpse of scaling law in
simulations. The more data about humans and more compute you ingest, you start
to get predictive and predictable gains of the model performance in simulating
it, simulating people.
Vibhu [00:40:51]: Ooh. We need a scaling law curve.
Joon [00:40:52]: It’s scaling law. Whenever you find it’s a beautiful thing. And
we’re starting to see the glimpse of it, which is quite exciting. But if you
talk about the ambition of simulation as a whole, it’s not merely about building
a model. It’s about building a model, then creating the agents that become the
individuals in a much larger ecosystem. So they’re creating this multi-agent
simulation. Down the line, you want these multi-agent simulation to also live in
a very rich environment, right? What we are really trying to get to at that
point is, hey, can we create. All right, let’s do a time machine game again, and
5 years, 10 years into the future, can we create a simulation of 8 billion
people living on Earth? I think that’s quite interesting. And that really is the
vision. And once you get to that state, the questions that you can help answer
for the society also start to change from my perspective. The answers are
fundamentally about emergence of the emergent behavior of society and large
groups of people.
Joon [00:41:53]: So the questions that I get excited by, and maybe this is a
stodgy- a bit. I have my, academic side of me.
Joon [00:42:01]: And for me, it’s questions like, can we help solve climate
change? If you look at climate change as a problem space, this is what we, like
social scientists would often call it the wicked problems, problem where you
have many actors with competing incentives for trying to make a very complex
decision and coordinating that coordination decision. Very difficult to really
solve in real life, which is also the reason why we couldn’t solve it. Can
simulation help us solve that? Another one is, can we understand the signals for
collapsing democracy, or can we understand or can we uncover the origin story of
the monetary system? These are societal questions that we never really had a
good way of answering. If we can create simulations of our society, you have to
believe that these are the problems that we can solve. So that’s really the
ambition of this field. And, I also think, yes, I think there’s a Nobel Prize to
be won there, which wouldn’t be surprising. And I think there’s some amazing
societal impact that we can have to help people make better decisions.
Climate Change, Democracy, and Societal Simulation
Swyx [00:43:04]: Nobel Prize in economics?
Joon [00:43:06]: In economics.
Swyx [00:43:06]: Oh, I see. I see. Rooting for you to write that paper.
Joon [00:43:10]: One of these days. But, one of the scholars that I was deeply
inspired by, When I was coming into the space of simulation, is this scholar,
named Thomas Schelling.
Schelling, Agent-Based Models, and the Nobel Prize
Swyx [00:43:23]: Schelling point?
Joon [00:43:24]: So the canonical example of the work that he’s done was he was
one of the creators of agent-based modeling. So this was, like, in the 1970s and
80s. It’s very early days, but this was truly one of the first exemplars of
simulations. And one of the canonical model from that time, and of course many
of these simulations are trying to tackle the societal problems that’s most
relevant for their era, it was called the model of segregation. So racial
segregation was a big topic, that, we cared about. And what they’ve done was
they created this grid world where they had red dots and blue dots. And these
dots were, back in the day, like, they were the agents, and they had a simple
rule that governed their behavior. If certain percentage of your neighbors are
of different color and if that goes above certain threshold, then you move to a
new location at random.
Joon [00:44:21]: One of the striking finding of this paper or this agent-based
model was for the longest time, people thought the segregation within society
was caused by explicit and overt racism.
Joon [00:44:34]: But if you look at this model, people’s preference towards
living with people of the same color, that preference can be very minute.
Joon [00:44:42]: But the very small difference causes the society to segregate
completely over time. This was very counterintuitive for a lot of people. And
this particular work ended up informing housing policies. Mixed income housing,
got really inspired by this work. And Thomas Schelling ends up winning the Nobel
Prize for having laid the groundwork for very early versions of simulations. The
opportunity that I do see here in the more scientific terms, is agent-based
models for the longest, had impact in the 1980s, 90s, to some extent, early
2000s, but it has now gotten forgotten by the community a little bit. Because as
you can imagine, red dots and blue dots is not really a rich description of
people.
Joon [00:45:31]: But with the emergence of things like generative AI and, in
particular, generative agents, we do have an opportunity to create these
agent-based models that are high fidelity enough to help us make really complex
decisions. And that’s the opportunity that I see. If that truly works, then yes,
that is the work that will result in a Nobel Prize.
Swyx [00:45:53]: Yeah. For what it’s worth, and I grew up in Singapore. 80% of
Singapore is in public housing, and public housing has, enforced racial quotas
for exactly that reason, which is very interesting. okay, so we talk about
scaling, we talk about all these, the agent possible applications.
Cost, Reuse, and the Economics of Simulation
Swyx [00:46:13]: I’m scared about the cost. if you even-- let’s just keep it to
the US, about 8 billion people.
Swyx [00:46:21]: But, how much does it cost to model so many hundreds of
millions of people?
Joon [00:46:26]: Oftentimes today, we don’t start at that scale, this stage of
the, of industry and simulation as technology. But we can get our users
extremely rich and meaningful insights even by modeling thousands, tens of
thousands of people. And today what we do is every week we are collecting data
on the scale of tens of thousands people’s data, and we have panel partnerships
that gets us to tens of millions of people globally. So that’s what we do today.
Swyx [00:46:55]: And just as a side note once you’ve collected one person for
one study
Swyx [00:46:59]: Can you reuse that same person for all the subsequent studies?
Joon [00:47:03]: That’s exactly right.
Swyx [00:47:03]: Okay.
Joon [00:47:04]: The beauty of this model and these agents is the fact that they
are domain-agnostic.
Joon [00:47:08]: That what you’re really trying to understand is what is the
fundamental nature of these people? What’s their social physics? And there are a
lot of, a lot of, people that does change over time. Like, even, like, even
things like, how many times have you gone have you been to, like, CVS the past
week? that will change. But there’s so many traits about people that are also
known to never change. Like, your risk tolerance doesn’t really change over
time. It’s very consistent. So it’s these things that we’re trying to learn. But
the scale we are operating is right now hundreds or, tens of thousands to
hundreds of thousands. And in many of the core use cases that we are deployed
in, and this is more than enough population, to cover those. Really, at that
point, what you care about is less the number of people, but more do you have
the right subpopulation of interest covered? And this is also the reason why
people want a larger sample. It’s not because they want, stronger statistical
guarantees. It’s more that can they filter down to any population of their
interest. However, you can also imagine in 10 years, if we truly believe that
the compute is going to scale, that we’ll have much more availability for
compute, and our ambition for simulation is also going to scale accordingly,
there’s definitely a reason for us to create an entire data center worth of
simulations.
Joon [00:48:35]: Or in my hunch here is I do think in the next some number of
years, we will start creating simulations that will cost as much as training a
foundation model. But perhaps it’s going to be so valuable to the society that
it would be a no-brainer. Right now, even today, like, we are training bunch of
new foundation model just so we can say we trained one and we spent tens of
millions. But if we can create a simulation at the level of society that would
solve climate change, I would run that today. I would raise the money right now
just to run that.
Multi-Agent Simulation and Social Influence
Swyx [00:49:10]: Amazing. the follow-up question is, does it also compound if
you let the simulations talk to each other?
Swyx [00:49:18]: Or do they already do that today? They don’t, right, as far as
I understand?
Joon [00:49:22]: It depends on what simulation you’re trying to run.
Joon [00:49:24]: In the multi-agent simulation setup, the agents do talk to each
other.
Swyx [00:49:28]: Right, which is exactly Smallville, right?
Joon [00:49:29]: That’s right.
Swyx [00:49:30]: But a lot of times, for example, in commerce, you’re just by
yourself, so there’s no point talking. which is way cheaper.
Vibhu [00:49:37]: But they use all these levels, right? Like, you decide what
you will buy based on what other people around you buy and talk about, right?
Swyx [00:49:43]: It depends.
Vibhu [00:49:44]: It depends.
Swyx [00:49:45]: Again, I’m, I’m coming at this from a cost point of view. I’m
like, “Oh my God.” Like
Vibhu [00:49:48]: I think
Swyx [00:49:49]: If there is, like, some combinatorial thing of, like, thousands
of people talking to thousands of people, then that one million X’s might cost.
Vibhu [00:49:56]: I have a very different view as the cost point aside. Like,
running these studies in reality is a lot more expensive, right? Running any
study like this is you gotta have people do it, you gotta sign people up. It’s
very expensive and sometimes, like, not feasible to run the study.
Vibhu [00:50:14]: But the outcome or the decisions you make are very expensive
on them, right? So spend X million on something that, the overall process costs
100 million might as well, right? There’s, there’s a lot of value to be had
there. It’s a small cost, but I’m excited on the cost side.
Joon [00:50:33]: To some extent, and when you deploy technology, you often want
to deploy in a way where you can replace existing budget or you can make things
more efficient, and that is the best way to deploy. However, the way you capture
the long-term value of the technology is making the argument that, no, it’s the
upside, that by making this better decision using simulation, you have saved
yourself or made yourself hundreds of millions or even billions of dollars, and
that’s a case to be made.
Vibhu [00:51:06]: Random tangent question. So if you’re doing a lot of
inference, a lot of model multi-agent stuff, are you at the point where it makes
sense to, train a model that’ very sparse? You’re expecting to do multi-million
dollar runs. Are you thinking about this in model architecture standpoint or
inference efficiency, or, you’re still at the research phase of it works, we’re
not super there yet?
Joon [00:51:34]: Efficiency, we do think quite a bit about. this is technology
that is deployed now in some of the largest enterprise companies in the world,
and we do process significant number of queries, that are trying to, simulate
the populations in the world. So efficiency is a consistent thing. we don’t want
to over-optimize too early, so I wouldn’t say, like, this is the higher bid
Right now, but this is definitely something that we think pretty carefully
about.
Swyx [00:52:05]: Yeah. Are there other case studies? So we, you talked about
CVS, talked about Gallup, Deloitte, Wealthfront.
Efficiency, Enterprise Use, and Real-World Case Studies
Joon [00:52:12]: Wealthfront is an interesting one, because one of the things
they were trying to do, they were one of the first customers that wanted to do
product testing that goes beyond just asking people what they think about, let’s
say, behavior experiments and so forth. So there, really what we had to do was
reason about multimodal input, so images, but also you can also imagine, like,
these agents traversing through Figma mockups or websites. So some of the things
that our agents can also do is it can be given a domain, like, or, like, a
website URL and go use it for a while. It’s these things. And Wealthfront was
one of the first, customers, that was very excited about this possibility.
Vibhu [00:52:53]: What have people been asking? Like, is there any demand that
we have not covered? Like, UI testing, right?
Vibhu [00:52:59]: I wanna try a new. I wanna ship a new feature, test the UI,
simulate how people will do it. Any interesting things that you’re seeing demand
for?
Product Testing, Websites, and Synthetic Panels
Joon [00:53:08]: Today, a lot of the demand does come from like, the places
where people have historically used human panels, we can now replace with
agents, and these synthetic populations. And this is not replacing human panel.
in many ways, the simulation that Simile is building is grounded. So the way
that I think about this is we are trying to represent humanity at scale. And in
that way, the use cases are what we would expect, but it’s the scale of
deployment that surprises me.
Joon [00:53:44]: Turns out there are so many decisions that people make every
day in these organizations, groups, and we want to be able to say, “We listen to
people. We have consulted our users.” But in reality, that is rarely the case
because getting to people and asking them many questions, it’s difficult. It’s
both costly, time-consuming, but most importantly, people are just not
available. If I had to answer 1000 survey questions for this one particular,
vendor, even if I wanted to do that, like, I would never do it. And that’s very
much the case. What simulation can do is ensure that the voices of people are
always represented in rooms where the decisions for them is made, right? So all
the stakeholders of this particular product launch, ideally they’re consulted.
That’s what this technology really is trying to enable.
Market Size, TAM, and Human Decision-Making
Swyx [00:54:39]: In my mind, that means it skews towards more consumer focus,
right? Like, anything with a wide enough customer base where you do benefit from
the diversity that you represent. What are some rough statistics, just for
people who are not familiar with this market in general, what’s the market size
that. I’m sure you have some, like, rough numbers. market size is, like, a vague
question
Swyx [00:55:01]: But, like, how much do people spend?
Joon [00:55:03]: So market research is a $100 billion industry.
Joon [00:55:06]: But the thing about simulation is not a tool for market
research. Simulation is a tool for human decision-making. So the question around
what is a TAM here is quite tricky, right? Because it’s easy to say, “Well,
market research TAM is roughly 100 million or 100 billion.” so is it a TAM? And
not really, right? Because in many ways, you’re trying to inform all human
decision-making. You’re trying to inform every decision that are made about
humans for humans. What is a TAM for that? It’s really unclear. And I’ll be
honest. Like, I have a scientific background, I have a research background, so I
didn’t come into the field calculating, oh, what is the TAM for human
decision-making? But I just had to assume, well, if we can inform every decision
that is made about human for human, that has to be big.
Swyx [00:55:58]: Some- something valuable.
Joon [00:55:59]: Exactly.
Swyx [00:55:59]: To some extent, you are a unicorn founder now, and you have to
care as a CEO. But, like, I do think, like, yeah, when you go into these
boardrooms with people that you’re quoting millions of dollars of contracts for,
like, you have to say, “Well, here’s what you spend on humans-”
Swyx [00:56:15]: “. And here’s what we save you, and it’s 85% similar.”
Joon [00:56:19]: And certainly, the value case, is something that we care deeply
about. Like, what is the value that we provide to the users and the
decision-makers? But this is also where, like, as a founder, I think valuation
only tells one very superficial aspect of the story, and I try not to think too
much about valuation, in general, because that’s not what also motivates a team
or certainly doesn’t. I’m, I-- Again, the interesting thing about researchers is
we are happy living in academia, getting paid next to. we get paid okay. we
don’t get paid that much, as a researcher here in academia, but it’s the impact
and it’s the, it’s the value that we can provide to the individuals and the
society that really drives us. And in that way, ultimately what drives us is the
impact. Does the simulation we provide have a real impact in people’s
decision-making in ways that progresses our society forward? If the answer is
yes, then yes. that has to be great business, and we see that in numbers, and we
do care deeply about that upside story, but that’s the heart of it.
Where Simulation Goes Next
Vibhu [00:57:27]: Do you have any timeline predictions? So we talked about
scaling laws of simulations.
Vibhu [00:57:33]: You brought up, okay, maybe one day we can simulate how to
solve climate change.
Vibhu [00:57:38]: Where are we now?
Vibhu [00:57:40]: If that’s not the end state, what is an end state, and what
does progress look like?
Joon [00:57:45]: So what I sometimes tell people is simulation as industry, it
feels a lot like where GPT-3.5, GPT-4 was, for the AGI saga, which is we have
now technology that is powerful enough to do real damage on the verticals that
we are tackling. At the same time, there’s a lot of progress that is yet to
come. And that’s, I think, where this is. So the way I see it, I do think there
will continue to be breakthroughs both in data, in algorithms, and there will be
much more aggressive scaling that will also happen over the next few years. But
I think that’s roughly where we are.
Swyx [00:58:27]: I think that was about the rough set of topics. Anything else
that we should have asked you or you wish people asked you more about Simile?
Simulation as Painting and Understanding Human Essence
Joon [00:58:38]: I think the, what’s, for me, what’s quite fascinating about
simulation, it is very impactful technology, but it is also very interesting
technology, both in terms of, like, what it means for human society, our
philosophy. And the way I sometimes interpret simulation is. So going back to my
background, I as I mentioned earlier, I started my career as a painter. it was a
professional pursuit, and I did oil painting, for figures. So I got my training
originally in the realism studios, and that’s what I spent a lot of my, years,
doing. Simulation is a lot like painting, right? The best paintings teach you
something deep about the subject that you’re trying to represent. And it is
always not a perfect representation. It-- No painting is perfect. There’s always
some small differences and discrepancy, but what it does is it tries to
highlight the thing that matters the most about the subject.
Swyx [00:59:47]: The essential
Joon [00:59:49]: The essential essence.
Swyx [00:59:49]: Yes. He, you, he’s brought up some of your work.
Vibhu [00:59:53]: Just nice to put it up.
Joon [00:59:54]: Yeah. So these are some of the works. So this is from, my,
personal website that I maintain when, I was still a researcher.
Swyx [01:00:00]: I think a lot of people will say, like a Picasso, like anything
postmodern is, like, very much focused on the essence.
Swyx [01:00:09]: Right. yeah, but I don’t know if any one of these evokes
something that you like to tell the story of.
Joon [01:00:15]: No, it’s one of those things where, each of these paintings,
drawings, whatever it may be, it is trying to surface something about the
subject that you feel deeply about onto the surface. when I was a painter, and
artist, the topic that I cared really deeply about was, the more mundane aspect
of human lives. This shows up in some of the, some of the work that I’ve done,
where, like, I did this entire study of a rural town where I went around and
took photos of people for not really doing anything special, but just living
their everyday lives. I thought that was the most interesting thing. I’m
somebody who has this perspective where, the world is oriented around this
fractal shape, and you have two choices to understand the fractal shape. You
either go outward and try to explore as much as you can to understand the
broader shape of the fractal, or you go inward because, the outward resembles
the inward, shapes. And understanding the mundane aspect of it was very much
that. Simulation has a lot of this, right? You’re trying to understand even the
most mundane aspect of people. When put together- teaches you something really
deep about that individual and the society. So I think that’s what’s interesting
about simulation, the way, the same way that AGI helped us better understand or
really think critically about humanity and human intelligence, simulation is
really an exercise of understanding more about human society and our collective
lives. So that I find to be, yeah, particularly interesting.
Swyx [01:01:56]: Yeah. Now you’re reminding me that some of the best
biographers, documentarians, and even photographers, they’re taking a photo of
you.
Swyx [01:02:05]: But before I take a photo of you, I must spend-- I must, like,
follow you for a week just to understand you?
Swyx [01:02:11]: Which some artists, some do. Part of your work, there’s a very
famous book called Working. I don’t know if you’ve, been referred to it before.
Swyx [01:02:18]: It’s very famous, like, to the point of having a Wikipedia page
Swyx [01:02:23]: About this like, really depth understanding and interview of
people as they, about their lives, which seems mundane, but is told in a very,
compelling way. Yeah, 1970s as well.
Joon [01:02:34]: Okay. It was an amazing decade.
Vibhu [01:02:39]: Before closing question
UBI, Future Questions, and the Value of Simulation
Swyx [01:02:41]: Okay, here we go
Vibhu [01:02:41]: You said that you started Simile with your 10-year question,
right? If we do that now, 10 years down, what can we simulate? What would you
simulate if, like, if you’ve made significant progress, are there any questions
outside of the ones that we brought up? Any- anything that you think is most
impactful? Anything that you would go vision 10 years out?
Joon [01:03:03]: In many ways, as I mentioned, I am somebody who is very much
impact-driven. So the what would inspire me is I would want to ask, 10 years
later, what would be the most important societal question that we as a society
have to ask? I would love to tackle that. Like, do we need UBI? That could be an
interesting one.
Swyx [01:03:24]: Ooh, has anyone done that?
Joon [01:03:25]: Well, we were thinking about it.
Vibhu [01:03:27]: Can we get access? Can we just
Swyx [01:03:28]: So OpenAI, this is, like, just trivia now. Like, OpenAI, or I
think Sam Altman funded a study on this
Swyx [01:03:35]: In Africa, and the answer was no.
Joon [01:03:37]: The answer was no. But, what, was it something about the
implementation?
Swyx [01:03:41]: Yeah, I know. It was a skill issue.
Joon [01:03:43]: Or was it something about, But this is the thing. See, when Sam
Vibhu [01:03:46]: Funny news article
Joon [01:03:46]: Altman funded this particular,
Swyx [01:03:50]: He spent 14 million dollars? Oh my God.
Vibhu [01:03:52]: It’s a little more.
Joon [01:03:52]: Quite a bit. But this is the thing. This is the reason why you
want to run a simulation. You spend 5 years, 40 million dollars on this one
study and have one finding, but if you can run simulation many times instantly,
then that’s the value.
Swyx [01:04:07]: I feel like that one could-- you could have done in a
simulation. Like, if you can do the housing study, you can do the UBI one. Like,
I, come on.
Vibhu [01:04:13]: I think sometimes people will spend the money because they
wanna verify what you think, right? Like, sometimes you just wanna. Is it right?
Like, you gotta test it.
Swyx [01:04:23]: Okay, closing question. What are the chances we are in a
simulation right now?
Are We Already in a Simulation?
Joon [01:04:28]: So it’s a fun question, and I assert at some point I just
answer, yeah, we’re definitely in a simulation. But what I do, feel, however,
is, whether we are in a simulation or not, that, I don’t think that makes our
experience any less real. And I think that’s fundamentally, like, what I believe
in. Maybe we live in a simulation, maybe not, but for
Swyx [01:04:48]: It’s real to us. Yeah.
Joon [01:04:49]: Yeah. For me, I don’t really care.
Swyx [01:04:50]: Yeah. Unless you die and you wake up in, like, the level higher
or below.
Joon [01:04:55]: That would be interesting.
Vibhu [01:04:55]: I feel like you wouldn’t care. Once you die, then you find out
you’re in a higher level.
Joon [01:05:01]: I worry about it when I die.
Swyx [01:05:04]: I think the other thing that. Okay, so I like the mathematical
answer to this, which is, like, the, sheer number of possibilities that you are
in a simulation far outweigh the sheer number of possibilities that you’re not.
Swyx [01:05:16]: Except for the simplest answer, which is, it is computationally
very expensive to have you be a simulation. okay, great. You’ve been very
generous with your time. Congrats on all your success. I met you just after your
Smallville paper and had no idea that you could build, like, such an enormous
company. And then now you’re like, “Well, it’s a $100 billion market, but that’s
just where we’re starting.” So this is, very exciting.
Vibhu [01:05:42]: I think $100 billion market was not the term. That was only
part of it.
Swyx [01:05:45]: Yeah, exactly. It’s, if you’re thinking too small.
Joon [01:05:48]: Well, I do believe that, maybe my final note here might be,
again, I love science fiction. You look at any advanced civilization in science
fictions, there’s 2 twin pillar, technology. One’s AGI in some form, and the
other is simulation. So I think the market’s pretty big here.
Simile as Research Lab and Product Company
Vibhu [01:06:08]: Tell us about the company. You guys just raised a lot. You’re
half a research lab, half a company. you’re hiring. Where are you based?
Joon [01:06:15]: Yeah. So we’re based in Mission Rock, so not too far away from,
where we are right now. So we’re in SF, but we are also bicoastal. So we have
our, team. I would say our headquarter is in SF, and we have a lot of our
technical talent in SF, and we do have a smaller office that just opened up in
New York. We are, as a company, an interesting one in that today, there are AI
neo labs and then there are AI product companies. Simile truly is both. So this
is a company that was founded by 4 founders, myself, Michael Bernstein, Percy
Liang, Lainie Yallen. Michael, Percy, and I are all researchers. So of course,
Michael was one of the authors of the ImageNet, kickstarted the AI revolution
back in 2013, has been instrumental in human-centered AI. Percy coined the term
foundation model, and is a, one of the greats of the AI researchers today. And
Lanie is my business counterpart, where she led some of the fastest-growing AI
native companies from their seed to A and B. But we have this DNA at the company
where the vision of the technology that we’re creating is continuously
developing, that we are getting people who were my lab mates. We are about 60
people right now.
Joon [01:07:28]: 15%, almost 20% of the company population are just my lab mates
from Microsoft Research lab.
Joon [01:07:36]: And we It’s quite fun because many of them then had gone on to
OpenAI, Google Gemini, and these places. And so it’s been a few years since we
really got together and had a chance to work together. But now they’re coming
back and really building out this vision that I find to be quite exciting, and
that excitement is shared. So there’s that motion at Simile where we are a group
of researchers trying to do something that no one is working on that we find to
be the most impactful potentially. But at the same time, this is, again,
technology that can make impact today. So we have an amazing group of engineers,
product people, and designers, who are sitting here with us trying to imagine
what does it look like to help people understand what simulation can do and make
real-world decisions with this. Having both and then deploying it to some of the
largest customers in the world today, it feels quite unique.
Swyx [01:08:30]: Yeah, it’s very compelling. One part of it was this is the call
to action. Like, who are you hiring? You’ve done part of it, which is you have--
you’ve got a very talented group. Who are you hiring? Like, what roles?
Hiring and Closing
Joon [01:08:41]: So honestly, at this point, we’re hiring across
Swyx [01:08:43]: Everything
Joon [01:08:43]: All, section. we are always excited to bring on, amazing
research talent.
Joon [01:08:49]: So if you’re interested in working with, our lab mates, we are
always welcoming of amazing, researchers. But also we, hire, amazing engineers,
that some of whom I, like, I respect the most. Many of them come from places
where we have personal connections with, so many of the members are from Figma,
Notion, Rive, and so forth, but also more broadly from the companies that we as
a team have really admired. So engineers both in the product side, infra side,
we’re all looking for those hires.
Swyx [01:09:24]: Well, lots of people. I think you made a really good case. So
thanks, and, we’ll see you in the simulation.
Joon [01:09:30]: Amazing.
Joon [01:09:31]: See you all there.
This is a public episode. If you'd like to discuss this with other subscribers
or get access to bonus episodes, visit www.latent.space/subscribe
[https://www.latent.space/subscribe?utm_medium=podcast&utm_campaign=CTA_2]
Highlights
Highlights
Chapters
Chapters
Transcript
Transcript
