Fei-Fei Li is Solving the Hardest Problem in Robotics | World Labs with a16z
By a16z
Summary
Topics Covered
- Spatial Intelligence Is the Next AI Frontier
- Real-to-Sim-to-Real Unlocks Scalable Robot Learning
- One-Third of Robot Tasks Are Cleaning
- Counterfactual Reasoning: The Hidden Power of Simulation
- Specialized Bodies Beat General-Purpose Humanoids
Full Transcript
We are building the next frontier of AI which is what we call spatial intelligence.
As Cynics, we are developing what we call a real to see to real pipeline. We
can replace all the data, all the evaluation we need in the real environment by using the data that can generate at a scalable way in our digital world.
Think about human intelligence. We do a lot of simulation in our head. Why?
There's a very important role simulation plays that real world data doesn't play which is counterfactual. Reason
what we are building is a consistent world consistent both over space over time over different viewpoints and over different type of interactions. My north
star is I want the robot to work.
The world we live in can be multiverse that we create technology to allow people builders developers to act within different spaces. Do you believe we'll
different spaces. Do you believe we'll ever be able to build robots that have the power efficiency of a human being?
How far away are we from this? Is this
like 5 years or this is like never?
The TLDDR is All right. Well, it's great to have you
All right. Well, it's great to have you both here. So, Feay, for the listeners
both here. So, Feay, for the listeners that may not have the background, maybe you can give an overview of what World Labs does.
Yeah. Well, World Lab has is a two-year-old startup. I I I think we
two-year-old startup. I I I think we should just recognize it's a frontier model lab. It's uh we we are building
model lab. It's uh we we are building the next frontier of AI which is what we call spatial intelligence. And uh
spatial intelligence is about um creating AI that has the ability to generate uh understand
reason with and interact with spaces whether it's physical or virtual. And of
course uh a means to an end towards spatial intelligence is building large world models. And that's what uh World
world models. And that's what uh World Labs is mostly focused on.
Yeah. So you've been saying this since the very beginning which is um you know the machine's ability to perceive and reason about spaces and act on spaces
but I always had the assumption that the acting on spaces was some like longist future thing but now you're acquiring a robotics company and so maybe talk a little bit about the timeliness of this
and the intentions.
Yeah. So first of all, it doesn't just take robotics to act within spaces or to interact, right? I mean, look at the
interact, right? I mean, look at the creative field, whether it's VFX or gaming and or design. Many use cases,
you can create and act within virtual spaces. World Warlab's thesis has always
spaces. World Warlab's thesis has always been that um the the the the world we live in can be multiverse that we create
technology to allow people builders developers to act within different spaces.
Having said that the ability to act within the physical space is one of the most exciting and most profoundly
important capability of the future AI world. So robotics is very much that. So
world. So robotics is very much that. So
worldlap has always believed that robotics is a important application as well as use case of uh
spatial intelligence and world modeling.
So by joining force with uh inviting scen team to world labs is part of our long-term vision and uh mission. We
we've always committed to that. Amazing.
So, Yunu, you're the co-founder of Scenix. So, maybe provide everyone with
Scenix. So, maybe provide everyone with a quick um overview of your background and what Scenix does.
Yeah. So, I'm Yunu. So, I'm currently co-founder of Cynics and also assistant professor at Columbia University. Wow.
So, my research started from my PhD at MIT and then postto with Fay really at the Stanford University.
That's great.
The world is small.
The world is small.
It is.
Throughout my career, my goal has been very simple. trying to help the robots
very simple. trying to help the robots better perceive and interact with the physical world. So I'm a very practical
physical world. So I'm a very practical person. I want my robot to work in the
person. I want my robot to work in the real physical environments.
So for cynics the unique opportunity we see is that there has been a lot of like a bottlenecks. Right now we see faced by
a bottlenecks. Right now we see faced by the developments of general purpose robots especially around training and also around evaluations. So as we are
developing what we call a real to sim to real pipeline. Okay. We want to map the
real pipeline. Okay. We want to map the real environments into the digital world that has the best alignments with the real environments. By alignments we mean
real environments. By alignments we mean that whatever happens in the digital world is also going to happen in the real environments such that we can replace all the data all the evaluation we need in the real
environment by using the data that can generate at a scalable way in our digital world. So that is how everything
digital world. So that is how everything started in Synex. We put together a very very strong and best teams around robotics, robot learning and also simulation and rendering trying to build
this realtom real stack to solve some of the key bottlenecks.
It's amazing that you two work together.
Yeah. And uh there is a funny story here because you would think because we worked together he was my amazing postto we've been talking about this and world
lab um um integration for a long time.
It's actually not true. They came into World Wars as a customer.
Really?
When we when we released the first version of our generative model called Marble last winter around uh November, December, Cynics just signed up.
No kidding. AS A CUSTOMER.
YES. And I didn't even know what it was.
And then I realized this is Vindrew's company. I called Vindro. I'm like,
company. I called Vindro. I'm like,
"Wow, this is your company." and you guys and and then we realized there's so much synergy.
Maybe Faith could just quickly describe what Marble is.
Yeah, Marble is the code name for the base model that World War has been training and iterating on that the fundamental capability right now of Marble that is publicly released is to
take a a prompt. It can be an image, it can be a few images and uh or a text and turn that into a geometrically
consistent world that can be represented in 3D geometry whether it's gausian splat or mesh. Really what Synix team is
doing is trying to solve this extremely difficult problem in robotics which is the lack of data. M
the lack of data in training, the lack of data in uh evaluation. This is very very different from language models where data is abundant on the internet.
And we know that um in order for robotics to work, we have to somehow unlock the power of scaling law. But
where does that come from? This is
something that that is a profound problem that everybody's battling with in in robotics. It'd actually be great to talk about this energy like you have put together a very very talented team.
You have put together a very talented team and so like to what extent is there overlap to what extent is this an extension? Maybe talk a little bit about
extension? Maybe talk a little bit about that. Yeah, that's actually like how
that. Yeah, that's actually like how complimentary it is.
It's it's actually the the TLDDR is is very complimentary and with the shared mission. So is one of the three uh
mission. So is one of the three uh technical co-founders. The other two are
technical co-founders. The other two are Changi Jan, another Colombia professor who has been a world-class technologist in simulation. Wow. And Changi has his
in simulation. Wow. And Changi has his background in also um VFX. He worked at Weta, he worked at Tencent, he's being
an entrepreneur. Uh then there's Sunonni
an entrepreneur. Uh then there's Sunonni who is a phenomenal engineering leader who was also in a uh startup uh that was
acquired by Amazon many years ago. So he
worked in many different tech stacks in the computer vision field in uh in Amazon. So when when we started talking
Amazon. So when when we started talking more seriously I recognized that uh a couple of things that Phoenix has from a talent point of view is extremely
complimentary to to uh worldaps. one is
obviously incredible um uh thought leadership and and just technical prowess in robotics
right so uh from really from hardware full stack robotics and even when he was my posttock at Stanford at that time you already had your faculty offer so you
were there only for one year I wanted you for more than one year but he had to go become a have the real job so uh he was a full stack researcher in in
robotics from modeling to to hardware.
Uh and of course uh VR and his student uh students at Synenix was that pool of talent world hasn't had yet. Then on the
Changi side is just incredible simulation um um capability, right? he's
such a senior um researcher and technologist in simulation and and uh what world labs is doing is very much um
interfacing the world of simulation. So,
so I think what they don't have um obviously is on the generative model side as well as the computer vision 3D reconstruction side we're also very
strong at world labs. So that's a technology that scenics needs. So
together these uh these two sides come together and make it much more complete.
Fe's motivation in this as like this is an extension and a compliment to get into robotics you know having been in your situation which is deciding when to sell a company it would be great to hear from you on like how you think about
joining world labs and kind of the fit there and like why you made the decision to do it. Yeah. So at the very beginning we were deciding okay we want to just keep going but after chatting with F
after seeing all the synergies that happening in the middle it just makes perfect sense for the forces to join each other. So in any sense as cynics
each other. So in any sense as cynics what we have been doing is real to seem to real is to do this reconstruction of the environment. So we capture the
the environment. So we capture the appearance of the environment geometry of the environments and also the dynamics of the environment meaning how the environment is going to change when you apply actions. So this tensor
reconstruction right now is still a little bit on the heavier side and what world labs right now has been doing involves a lot of profound capabilities around sparse reconstruction and
generations. So we see a lot of
generations. So we see a lot of opportunities of leveraging like a marble and other like capabilities as world labs in order to do very efficient reconstructions and modeling of the environments.
So can we expect a foundation model for robotics from world labs?
World Lab is building a foundation model. As you know, Martin, we're
model. As you know, Martin, we're building a base model and uh as the technology has been evolving, some of the most exciting base models
are omniodels, right? They take they take multimodal input, they have multimodal outputs. And uh what is a
multimodal outputs. And uh what is a foundation model for robotics? Uh it's
very likely going to involve actions. M
it's very likely going to involve the output of actions in addition to the state of the world and we're definitely not ruling this out.
Yeah. Great.
So for example for the foundation models it's essentially needs to be a multimodal model. So it has to take into
multimodal model. So it has to take into account frame text image deps and different kind of modalities and action is a very very important parts of that modalities. So if you think about frame
modalities. So if you think about frame actions as inputs that is essentially a forward simulator that is going to predict how the environment is going to change when you apply a specific action.
When the action is output this is essentially a policy model that is trying to predict given a specific goal like what should be the action you take in the real environment to get you closer to that goal. So this kind of
omni models actually can benefit a lot and actually provide huge amount of values for the robotics communities in trying to understand how to model the environments and at the same time how to act in the environments and this can
also acts as a backbone for you to fine-tune into specific robotic applications to making sure it's really leave up to the reliability and efficiency that expected by the clients.
you know if uh Yun if you don't uh if you don't mind a kind of a lay investor question I see a lot of robotics companies and a very popular approach right now for the robotics companies
that come in is like we'll use a video model you know and like you know that's the the predominant method where this is you know 3D and simulation it's a very different approach and so maybe you
could contrast you know this popular approach of just using video only versus kind of what the ambition here is yeah so in order to create words where the robot can learn the words as I
mentioned is to capture the essential structure of the problem and one of the very important necessary like requirements for those words will be consistency so that is where I actually
see there's very very strong synergies with marble because what we are building is a consistent world consistence both over space over time over different viewpoints and over different type of
interactions and marble the generated words from marble is also provides an infrastructure a component of that entire words that we believe is necessary for the robot to learn.
Imagine if a robot push an object forwards. The object just magically
forwards. The object just magically disappear which has been a problem of many of the existing like video prediction models. This won't provides
prediction models. This won't provides good enough signal for the robot to know like what is the right thing to do. But
obviously right now there has been a lot of investigation on building better and better and stronger and stronger like a video models. So we actually see a way
video models. So we actually see a way where some of the infrastructure we build can provide as a initial momentums and to going through this data flywheel
of going from this like a more simulationdriven models into like a robot policy models which going to do the execution in the real environment collecting new data the data will come back in where the model doesn't
necessarily have to be physics only or learning only but somewhere in the middle which be able to capture the essential structure of the problem but at the same times be able to scale and become better and better as you
accumulate more data.
You know, um I've worked now fay very closely for a while and and and you've always had this north star which has driven this and you know you you've articulated variously as kind of 3D and
and in a number of other ways. And I'm
just wondering for you is there also a similar philosophical northstar or you're more the pragmatic like I am like build the system like do the thing.
My northstar is to make robots work.
Amazing. Yeah. in the real environment.
I'm a very practical person. I want the robot to work. One interesting thing that's actually coming from my collaborations with FIP during my postto, we are building this kind of benchmark. We actually send out surveys
benchmark. We actually send out surveys asking the general public what they want their robots to do for them.
Yeah.
Among the thousand tasks we collected, onethird of the tasks are about clean.
People just don't like to do those like a d and dirty tasks. And those are the scenarios we really want to making sure we have robotic solutions to deal with.
One thing I really like about Synenix uh Martin especially uh uh continuing your question there's a lot of robotics companies building models and all that.
One thing I truly like about Synenix is is Renu and his co-founders have such an incredibly pragmatic approach to
robotics. They especially they come from
robotics. They especially they come from academia right Sunny doesn't but Vindra and Chani come from academia but their first instinct is work with design
partners and customers in real industry whether it's is uh labs at um industry labs or or warehouses or um um electronics
electronics you know uh assembly that is such a refreshing actually a refreshing way of approaching robotics and that that really tr made me
very excited to work with them.
Maybe this is for you and but I'll just be this is this is personal curiosity which is it seems to me that for robotics you have to be pretty exact. I
mean not perfect but pretty close. But
for the creative use cases which worldaps has done a lot of you kind of don't need to because you know I mean you know even sometimes like being wrong is stylistic or intentional or whatever.
And so from a technical perspective what is the challenge here for reconciling these two things or do they never get reconciled like there always be two points in the design space? So they will
be like reconciled in the in the long terms of course and um modeling of the environments um doesn't have to be perfect. The model doesn't have to be
perfect. The model doesn't have to be perfect in robotics.
And by the way is there again this is pure curiosity but is there like um a bit more formal way to say that like what does that mean not to be perfect?
It has to be pretty close.
So so let me putting this way for example models over the developments of all different kind of robotic applications has been a very important cornerstones. Yeah. If you look at all
cornerstones. Yeah. If you look at all the existing robotic applications like plane, drones, Roomba or even for quadripad robots, bipedal robots, model has been the way for them to actually work and be able to
transfer from simulation to the real environment.
I see.
But if you look at those locomotion robots like quadriped robots, bipedal robots, they can walking on snows, they can walking on bushes. But you don't need to have a simulator. They can
simulate all the bushes and snows like very precisely. you needs to have a
very precisely. you needs to have a simulation that capture the essential structure of the problem and do a whole different kind of randomizations inside the digital environment. So that is what we're aiming for. So basically with
Synenix and together with word labaps we're trying to investigate like what is the level of fidelity we needs to model the mass massive worlds besides the robots I said we'll be able to transfer
the robotic systems training the simulated environment digital worlds back into the real scenarios. As an
investor, I've heard other researchers say like Sergey Lavine say simulation will always eventually deviate from the physical world and real world data
collection is absolutely critical. And
so maybe talk a little bit about like the viability of this approach where simulation is a cornerstone as opposed to some other approach.
So they don't contradict with each other. So if you think about the
other. So if you think about the simulation, simulation is essentially trying to predict how the environment is going going to change when you apply the actions and this is essentially a model
of the worlds. It doesn't necessarily have to be pure physics. It can be a combination between both physics and also learning. We are collecting real
also learning. We are collecting real world data. We will be using those real
world data. We will be using those real world data. It just at different stages
world data. It just at different stages of this like a data fly. Maybe at the very beginning we have stronger emphasize on we have more physics to making sure we have the right consistency and right structure for us
to learn the uh the the worlds for us to train the robot policies. But as we accumulate more and more data both through data collection and also through the collaboration with our clients we'll
have the data that will be moving towards more towards more learning based like modeling of the environments. So
this kind of transition and also this kind of data flow is really an enabling factors of both getting the best of both physics and the and and geometry and consistency as well as all the power and
magics from the data and compute.
I want to add to this and be slightly philosophical here is there isn't a a binary choice between simulation or no
simulation. All this come um in together
simulation. All this come um in together um to to make robotics work. Think about
human intelligence. We do a lot of simulation in our head. You know why?
There's a very important role simulation plays that real world data doesn't play which is counterfactual reasoning is that you play out events that c hasn't
happened or cannot happen or you don't have enough data to make it happen in real world. And while you play it out,
real world. And while you play it out, you learn how to act in it. Humans do
this all the time. We probably don't, you know, we just I know you were at World Cups.
I was at the World.
Congratulations to Spain winning. I'm
sure in the planning of every game there is simulation, whether it's digital or on the on the whiteboard or whatever, that simulation, the role simulation plays is counterfactual reasoning. And
that's really important in robotics because we just do not have cannot possibly have enough real world data for that. Here's a real life uh example the
that. Here's a real life uh example the the industry of self-driving cars. Whimo
has officially said they use billions of hours of simulation and and actually Whimo is more simulationheavy than just real world
data heavy. So these are real examples
data heavy. So these are real examples and and as you know Martin Andrew too cars are the simplest kind of robots.
Yeah.
Yeah. So so clearly simulation plays a huge role in robotic learning.
I also want to add to that. So like
there are if you put things more specific simulation can provide two levels of benefits. The first one is reliability and the second one is efficiency. So for reliability, if
efficiency. So for reliability, if you're thinking about a robotic system working reliable in the real environments, you need data to provide systematic coverage of all the state space and the variations that robots
might encounter. That's how you can
might encounter. That's how you can learn of how that is robust. So with
simulation, you can do systematic randomizations and control and the variations of lighting, frictions, geometries, object types and also all different kind of physical parameters to
making sure you have sufficient coverage of the state space. So this is what can give the robotic systems reliability and second is about efficiency. So right now
many people are doing tele operation and if you look at many of the television device imagining all the actual skeletons you are using you're actually collecting the data at a speed that is actually slower than human actually
doing the task.
But for many of our clients human speed to them is not good enough. They want
faster than human speeds. Yeah.
So for the robot to like move faster, it's not as simple as just drive the robot faster because the gravity doesn't change.
But in simulation, you can do systematic speed up of the robots behaviors to train the robots such that it considers all the dynamics changes of the environments. So this is what's can give
environments. So this is what's can give like our clients for them efficiency. So
both for the reliability and efficiency though there are some kind of like a very unique like values where simulation can provide. You've talked about the
can provide. You've talked about the technology and the platform, what it does. Maybe talk about the specific use
does. Maybe talk about the specific use cases people use it for.
There are essential like two specific use cases especially around both training and also around evaluations.
Okay.
Um starting from the evaluations.
So evaluation is something like people often overlooked in the robotics but if you are tuning like robotic models you have to know how well it works and that is the only source of
information for you to iterate.
Yeah. By the way, a lot of a lot every every AI person really understands what eval are and uses it all the time. NonAI
people, it often means something a little different. So maybe it's even
little different. So maybe it's even worth just describing specifically what you mean by evaluation.
Okay. So what I mean by evaluation is you'll be able to understand for this specific checkpoints how well does it perform? Does it perform for example 95%
perform? Does it perform for example 95% of the time or 99.9% of the time? And the key criteria people use in industry is how long does it
take? How long in walk work walk time
take? How long in walk work walk time does it take for you to distinguish between a checkpoints that is 90% from a checkpoint that is 90 92 points. And if
you only do that in the real environment that's just takes so long for you to do the distinguishment. And if you really
the distinguishment. And if you really think about also the robotic evaluations right now people are doing in the real environments the iteration speeds is multiple orders of magnitude slower than
iterations of those language models.
Yeah.
So not only is like the robotic tasks very varied very diverse.
Oh yeah because like you actually have to do the thing right atoms have to move through space.
Exactly. But only there are the laws of physics have to do.
Have you watched those robotics videos?
Every video has like 10 x 8x because it moves so slowly.
Exactly. So not only it's slow, it's dangerous, it's costly, but at the same time the speeds is also like multiple orders of magnitudes like slower. So
some of our clients actually needs this digital environment that's can be used to evaluate their robotic like systems and because our digital environment has proven alignments with the real world.
So meaning whatever happens in the sim is also likely to happen in the real environment. If a checkpoint is working
environment. If a checkpoint is working better in the simulation is also highly likely to also work better in the real environments as we have also been discussed in the blog post. So that
actually give our clients very strong confidence in actually using the data using the signal from the digital environment to do scalable safe and much faster evaluations of their robotic
systems. Great. So that is on the
systems. Great. So that is on the evaluation then on the training. So on
the training sides, so basically like I also mentioned it's about controllability. So you want to control
controllability. So you want to control all the different possible variations of states parameters lighting frictions physical parameters like even object geometry, object types. So you want to
making sure you have sufficient coverage of all different kind of scenarios such that you'll be able to generate like a informative data for your robots to be robust. And this is just going to be so
robust. And this is just going to be so hard to do just in the real environments like we discussed. If you do tele operation, the speed at which you're collecting data is slow. You're also
limited by like how many robots you have, how many tele operation device you have. There's like a whole different
have. There's like a whole different kind of like challenges around all the data operations around it. But in
simulation everything can be controllable, everything can be systematic and everything can be understand at a level where you know exactly and making claims about exactly
what distribution you have covered to develop confidence about within that distribution. we know the robot will
distribution. we know the robot will work. So those kind of confidence and
work. So those kind of confidence and efficiency and scalability is something that our clients also value to use our like a digital words for the training of robotic systems.
Here's the crazy thing even before uh Synix and we are talking our inbound customers for Marble were already seeing this kind of demands. we just cannot
serve these customers. But we are already getting a lot of phone calls from from robotics early stage robotics companies who are developing their
models all the way to downstream very pragmatic use cases and uh we're seeing these uh these needs. When people hear you're going into robotics, what they're
going to envision is you're pulling out a 3D printer and you're going to be making hardware and then you're going to be programming the brain of a robot and sticking it in the robot and then you've
got a robot. And I don't think that's what you guys are talking about here. So
maybe talk about what you know where this fits in the life cycle of creating a robot and like where you will end and where the rest of the ecosystem will
will begin. So what we have been
will begin. So what we have been building you can imagine is a infrastructure like with the softwares around this infrastructures for people to for them build words such that robot
can learn and evaluate and this infrastructures is naturally model agnostic and embodiment agnostic. So I
just want to be very clear just because this is actually a very subtle for you it's obvious but it's a very subtle point which is um from what you said that's not building a robot it's building an environment which
another company can place their robot brain exactly to navigate and to learn.
Yeah. So for our customers right now they have all different kind of robots.
Some are using for them single robot arm some are using bio some are using a fixed arm. Some are using like a mobile
fixed arm. Some are using like a mobile manipulators. Some using grippers, some
manipulators. Some using grippers, some are using some more elaborate versions of the end factors. So our platform right now is just naturally embodiment agnostic. We can very easily integrate
agnostic. We can very easily integrate different kind of robotic embodiment be able to put them into the works we generated with digitalized such as we will be able to give those individual
robots capabilities of doing the right tasks and at the right levels of reliability and efficiency in the real environments.
Yeah.
And we are also for example model agnostic. So we can just using the data
agnostic. So we can just using the data generated by our words to train different models either from scratch or doing post training of existing foundation models like vision language
action models or word action models. So
to us it doesn't matter we just want to making sure we have the infrastructure we have all the words such as the robot can work reliably in the real environment. you know, you you you have
environment. you know, you you you have uh told me um that you think uh a lot of the predictions around humanoids were a little bit aggressive and we're likely to see more constrained rollouts like
warehouses or whatever. Can you talk a little bit about that and like how that impacts what you're going to be tackling here at uh the like world labs?
So that's a very good question. So if
you look at for example all the uh progressions of robotic applications in the real environments it has always followed the trend from going from fully structured environments into semiructured environments and then into
unstructured environments.
For fully structured environments what do we mean that you have knowledge and control over all the configurations within the environments like factories like factories or for car manufacturing mice those has been automated for
decades.
Yeah. Yeah. Yeah. And then you have for example semiructured environments which you have certain controls over the environments for example like the Amazon warehouses or for example like
restaurants hotels where you have certain control over the environment to just make the task easier for your robots but there are obviously many other like objects or for example clothes those are the object you don't have control and then for the
unstructured environments it's like your home and my host those is I would say the the grand challenge Especially Especially my house, trust me. Three dogs,
five-year-old.
Yes, dogs.
Exactly.
If you're thinking about where does the robustness coming from, robustness coming from a sufficient coverage of the scenarios that robots might encounter.
So, it's so much easier and more approachable at least like right now to focus more on the semiructured environments before we move on to fully unstructured environments. So, we will
unstructured environments. So, we will move into that direction. It's just we want to uh take a more sustainable and more realistic approach towards it.
I think your point here is that humanoids mimics human body and evolution has optimized human body for
unstructured environment.
And so our fingers, our legs are not the best apparatus to do one thing. For
example, if if if our only goal as a species is to climb trees, we will not have this body necessarily, right? So,
we'll have different kind of fingers.
But what humans end up having are involve evolved into is this this body shape that can be very general but not necessarily best at everything. And that
is for the survival of unstructured environment. But but but from a business
environment. But but but from a business point of view, from a pragmatic um technology point of view that this
unstructured environment and a generalized body is actually the hardest problem to solve. It's not necessarily even the right way to solve the problem.
It's we specialize. So we take more more specialized body to solve a narrower problem. But the challenge for scenics
problem. But the challenge for scenics is is that to be more body agnostic so that their infrastructure can serve different bodies and different
semistructured environments.
You know a common lens to look at exactly this question is an economic lens right which is um well you compare it to like generative LLMs where they can create uh pros or
code 10,000 times faster than a human being a bunch cheaper than a human being. So the economic case makes sense
being. So the economic case makes sense because our brains aren't very efficient at that. However, our brains and our
at that. However, our brains and our bodies are very efficient at 3D navigation, right? You know, like moving
navigation, right? You know, like moving through the world or picking things up.
And so this is just a prediction question, but do you believe we'll ever be able to build robots, at least in the foreseeable future, that have the power
efficiency of a human being when it comes to to menial tasks? So let's say just basically you know minimum wage or something like that like how far away are we from this? Is this like 5 years
or this is like never I think it's going to take a very long time. So if you really think about like a robots in the real environment in the end it it will always be a system. So every working
robot in the real environment is a system work. It's you need to be very
system work. It's you need to be very mindful and thoughtful about how the system are coming together. the
hardware, the software, the brain, even like to the details of for example what's the friction coefficients of your of your fingers. So there's a lot of things you have to consider to make
these things uh a reality and it will take iterations. But what I am excited
take iterations. But what I am excited about is that I have always been at the state of the arts of robot learning and also trying to push the state of the art forward. Yeah.
forward. Yeah.
But the state-ofthe-arts always moving faster than I expected. So what's I'm focusing on and trying to investigate right now is very different from for
them when I started my PhD. So this is a speak to how fast the whole ecosystem has been evolving and all the moving pieces started coming together or building this robotic systems but we
also have to be like calibrated about our predictions. So we will see a lot of
our predictions. So we will see a lot of like a progress but to achieve for example human level efficiency and capabilities it will take uh longer.
Martin, the hardest thing in today's AI is to have the right measured optimism, right?
That's right. It's totally true. Yeah.
I mean, even LLM does not have human brain efficiency. Human brain operates
brain efficiency. Human brain operates on 30 watts.
Yeah, that's true.
That's so we are far from that. So,
but that but I mean performance to power it may be close, right? in
narrow task like software like generating an image or or software engineering than it is right. Yeah, I
think so. I don't think we're anywhere close when it comes to robotics. Does
this change how you think about your um like strategically the level of ambition that your team can go after? I mean does it is it changed that or is it still very much in line with what you expected to do when you started?
It definitely changed the trajectories in a very profound manners. So we see a lot of unlock in be able to do this whole process do the modeling of the environments yeah much more efficient
and much more scalable manners especially in partner together with word labaps and I also want to add to f if you think about for example the current states of the language models
so those are models that's with incredible capabilities but still you don't just blind trust it to book your flight tickets or make your hotel reservations you still hopefully there's still a person who reading the
output from those language models.
Yeah.
But that is very different from how people and we'll be using for them robotic models because for robotic models out of the box the robots has to work reliably in the real environment.
So and we don't even have the data we don't even have all the necessary infrastructures around those for the robots to just out of box work reliable in the real environments. So for that reasons be able to create this digital
worlds this scalable digital worlds where the robot can learn in and evaluates within that is going to unlock so much more potentials for be able to replace all the like a costly and unsafe
data in the real environments with the data generated from the the the the words for the robots to be able to do scalable learning and evaluations.
You know I've I've seen uh I've seen many of kind of these kind of integrations. they actually work very
integrations. they actually work very well at this stage when they have this much alignment which is great. Um but
there's always like this question of do you integrate now into what's happening now or do you keep things quite separate and provide kind of like a long-term trajectory that will you know be
realized you know in in in the year time frame. How are you thinking about this
frame. How are you thinking about this fee? Is this something that integrates
fee? Is this something that integrates right away or is this kind of a separate longer term?
This is a great question. I think at this point you know Changi Sunonny Justin Ben and have been talking about
this at this point we are going to take it thoughtfully we're not rushing to integrate everything from codebase to teams
because I think Synex does have a very uh um wellthought and I wouldn't call it standalone completely but fairly uh
contained um tech stack as well as their customers as well as the the the kind of products they're building. We're going to take
they're building. We're going to take time. We definitely will we already on
time. We definitely will we already on the simulation side as well as the the potential base model um action condition model side. We already are starting to
model side. We already are starting to talk and also they are using Marvel as a internal customer. So we we will be
internal customer. So we we will be integrating but uh we're not rushing to blend the team as like a full salad bowl.
Yeah.
How are you thinking about geographies with this with scenics move? Is it going to stay in the same place?
We're going to Vundra is going to move.
Oh well, welcome here. Moving to San Francisco. Yeah, Florence and the
Francisco. Yeah, Florence and the Renaissance. Perfect.
Renaissance. Perfect.
I think we our labs is officially becoming a by coastal company where the headquarter is in San Francisco. I've,
you know, I live in Palo Alto. I feel
like I'm in a different state. We we but uh we but um I'm actually excited that we're going to have an office in New
York that can help us to attract talent on East Coast and also we have been talking about making sure that in both offices we set up the robots so that we
get to basically test out and and and uh mature our engineering stack so that we can work with robots remotely because we have to do that for for our customers.
Anyway, so may maybe just to be very concrete, Fay, maybe let's just pencil out like what is the the perfect success case in two years? Like what product do you
two years? Like what product do you have? Who's engaging with it? How do
have? Who's engaging with it? How do
they use it? Just just the crisp like what? We're very happy that uh Synix
what? We're very happy that uh Synix team and World Apps team will have validated
um customers in in a small number of important vertical use cases where our
system, our infrastructure has proven to be truly beneficial to their automation
needs And these customers became our lighthouse examples to scale our business.
And how how early let's say someone listening to this uh is running a robotics company? How at
what stage do they engage with world labs? Is it really early on? Is it
labs? Is it really early on? Is it
somewhere in the middle?
So right now um for our customers because we are building this kind of real to sim to real pipelines like where the simulation is essentially the words we're going to provide the training evaluation grounds.
Some customers they needs only the real to sim part they want to digitalize the task they care about and be able to do the evaluations of their robotic systems. Some customers they need this real to
sim to this entire pipeline such as they will be able to like have policies like running on their hardwares. So our
platform is also designed in a way that is flexible depending on what our clients needs and at the same times uh the clients we are working with are actually pretty close to the deployments
like a stage. So basically they are working on very very practical tasks.
those tasks when replaced when we have a robotic solutions that are there can just create value immediately and they have like at least like uh tons or hundreds of like this kind of situations
they are thinking about to do the automations for. So as like together
automations for. So as like together with word labaps we'll be able to develop reliable solutions for those scenarios as we have already showed we have a number of scenarios already
instantiated in our blog post and we'll be able to further our investigation to see how they can actually solve the key like like requirements and and and also constraints faced by the real world deployments.
Great. I want to be very specific about this. Is it ever too late or too early
this. Is it ever too late or too early to call World Labs if you're a robotics company?
No.
We we want everybody to call us. We want
to learn about your use case.
Wonderful. If you listen, if you're listening to this and you're anywhere close to a robotics project or robotics company, uh please track World Labs.
Yes. Thank you.
Definitely open for business.
We are open for business. All right.
Not too early.
All right. If you're doing robotics, call World Labs. Thank you both very much for coming.
Thank you.
Loading video analysis...