Biomni: A General-Purpose Biomedical AI Agent | Kexin Huang
By BioTuring Team
Summary
Topics Covered
- Highlights from 00:05-12:03
- Highlights from 11:51-24:07
- Highlights from 23:57-37:09
- Highlights from 37:00-48:35
- Highlights from 48:25-58:48
Full Transcript
so much Kind. So this is this is my great pleasure you know today introducing Kessin um he's a computer scientist of course
at the computer science department at Stanford University and uh in the jur let's go back group
right uh so uh I don't know if if people works in the in graph theory and high performance computing you know this is
this is the This is the this map was the library that Julia wrote like 15 years ago that solve a lot of things for the community.
But let me go back and uh introduce Kin right because he's uh he's the boss today.
So uh Kin research focus on leveraging AI to drive novel deployable
interpretable biomedical discoveries while also attacking fundamental AI challenges such as multimodal
uh modeling, uncertainty quantifications, agentic reasoning, right?
And Kessin's work has been published of course in nature medicines, nature biotechnology and many others you know uh very
um high impact journals but at the same time as computer scientist he also publish in the conferences like nurips
IC ICML ICO ISBC and other and other conferences Okay. So his research has been featured
Okay. So his research has been featured in uh Forbes in MIT technology reviews and he also contribute to machine
learning research at leading companies such as Genitech, GSK, FISA.
So today he present his works about uh Biomni. This is a John work with Aviv
Biomni. This is a John work with Aviv Rev with Jouer Lecovc and others. So
this reflects both practical and fundamental work in machine learning. And this is you know I read the paper and enjoyed so
much.
Um because Bi Omni is a generalpurpose biomedical AI agent that was designed to autonomously execute a wide spectrum of
research task across diverse biomedical subfield. And in the future what we are
subfield. And in the future what we are seeing now I I think we can think of the future of when we as scientists working
there would be a virtual assistant of scientists knowing almost everything there you know that people know and suggest us what we need to do in order
to fill in the gaps. So this is this is the dream of science right and uh with that said
uh Kessen the stage is yours. Uh
thank you so much. Cool. Um thanks for having me. Uh it's a really great
having me. Uh it's a really great pleasure to be here and share Bomni and also share our uh vision around AI agents for medicine. As son said, we we
do feel like agent is more of the I guess the it's very it has very different capabilities compared to previous generations of AI models for bio medicine and I think we're all like really excited about it. So look really
looking forward to the talk today and looking for the discussion um as well.
Um so I guess uh I'll spend maybe the first 40 minutes to give the talk and then maybe left uh around 15 minutes uh for question and answer. I guess sounds
good.
Um let me share my screen.
Right. So you pre prefer to leave the questions to the uh to the end right Ken? Oh honestly both works because is
Ken? Oh honestly both works because is that is there any preference on say no no because usually if with this kind of
interesting talk like this usually you know if we allow the talks to be intermediate then you know it would lead to a three hours talk easily. I see.
Okay sounds good. Yeah then we can uh leave it in the end. Um yeah. All right.
Cool. Sounds good. Yeah. So uh very excited to be here. Um so I'll talk about our I guess uh our work uh in our group around like AI uh agent for for
bio medicine uh in general with a one one representative work with pi omni um but also I want to share a bit a little bit of like our understanding of AI agents for bi for biome medicine why is
it important maybe in the first five minutes as well. Um so as you probably all know the biomedical research so fundamental like drug discovery, clinical diagnosis, agriculture and even
like consumer health like the rain here and the manufacturing, cosmetics like it uh is basically the the the cornerstone
of many different um industry um here and and for the past uh two years we actually we realized there's a kind of a um kind of a observation that is like if
we look at the biomedical research progress for us as in in as a whole uh partially is limited by like a lack of ideas but another big reason is just
limited by human bandwidth um like when I for for example I talk to um uh PIs in the lab they usually have like 50% of data just analyze uh because they don't have enough budget to hire
bioformmatician there's just not enough bmatitionian to analyze them and I and if I talk to some some scientist friend they always have crazy ideas like they have a list of ideas just on on the
backlog but no but they can just they just cannot do it because they don't have the capacity to do it. Uh so this basically motivates this like human biometical progress is actually limited
by human bandwidth. Um so what we mean by this is like as time goes by the supply or you know the current amount of people that have this biomeical expertise is kind of growing slowly but
the demand is just growing like exponentially because of these exper new experimental technologies. um also
experimental technologies. um also because of human creativity. Um so
there's a huge deficit of of demand and supply. So this also motivate us to you
supply. So this also motivate us to you know need a new completely new way of scaling this kind of expertise and one way to scale expertise is train more scientists but uh this is just like
impossible right now especially with this funding cut and everything. So
training a scale scientist is like expensive slow um and it's it's just not scalable. So we we must fundamentally
scalable. So we we must fundamentally rethink how to scale uh biomedical expertise and and recently we have seen AI agents can scale expertise in other domains. So
just do a very very brief overview of what is AI agent. Basically as human uh gives a instruction uh the RM large language model uh will basically understand the instruction and perform
some action to an environment and this environment will based on the action provide you know perform action and then provide feedback to the RM and and uh and then the OM can decide if it needs
to do further action or it can stop if it the instruction is kind of a you know achieved. So, so this kind of
achieved. So, so this kind of autonomously it's it's almost like a like a almost like a real human like human scientist uh that can you know perform action to the to the to the
entire world and then and then get feedback from them and and so on and and this framework has shown to scale expertise in cursor uh like for coding agent I know as a as a computer science
PhD student I actually don't need to code anymore because you know the coding agent has scaled expertise scaled the software engineering expertise uh for Similarly for hobby there's like uh
similar for legal industry there's like hobby and and so on. So we're
interesting to see if we can extend agent to do bio medicine uh as well.
So so this this is kind of a motivation uh for past two years of research uh which is like can we create a virtual AI biologist that automates biometical
research. So if we really create this AI
research. So if we really create this AI agent that can perform biomatic research then for example we can perform all of these tasks much much faster and potentially more accurate and also the
amazing thing about automation is like we can tackle many of the task at the same time. So this really enables some
same time. So this really enables some new capabilities, right? Because
previously we are limited by we because we know we have band we have limited bandwidth. So we can only do like two
bandwidth. So we can only do like two genes, three genes but now they can do automation. We we we never need to
automation. We we we never need to select which genes to perform on we just do 20,000 genes all at once. So we can tackle many of the task at the same time. Um so this the the power of
time. Um so this the the power of automation just enables a lot of new capabilities uh um in in biio medicine and it can also enable like suggestion
and explore like novel hypothesis and and and make like no targets and things like that. So there's there's so much
like that. So there's there's so much applications uh if we can build really like a virtual AI biologist and uh and a and a similar effect having
the cursor it also happened to biologist um where a biologist organ this AI scientist or like research assistant it uh it can basically have the impact of entire specialist team. So a biologist
can actually perform bio binformatics job can actually uh if you are if you're a microbiologist you can also have uh have the ability to do structural
biology um this kind of cross domain um kind of a kind of specialization as well and uh and also you can make a biomeical research very universally accessible
right because this agent if it's trained by knowledge from you know the top labs then it can also deliver to to to knowledge to some like I don't know like a primary school in the in the in in the
developing countries. Um so it's also a
developing countries. Um so it's also a very powerful um accessibility tool as well. So we really believe that it can
well. So we really believe that it can amplify and extend the humans kind of a biologist potential. Okay. So so so
biologist potential. Okay. So so so that's kind of a very quick motivation of why it's so exciting to build a AI biologist. Uh next I'll talk about you
biologist. Uh next I'll talk about you know how do we really build one like what is what is our approach to to build a AI biological agent.
So how is balance research being done today? Um so so so basically if we want
today? Um so so so basically if we want to build build a biology we first need to understand how is this ball research being done today so that we can you know learn from it and simulate it. And in
our view if we make it a bit abstract um then it's like high level then it's like two step it's first one is generating hypothesis here a biologist say this gene causes some phenotype which is like
tail loss. It can perform reasonings. It
tail loss. It can perform reasonings. It
can perform some experiment. It can per do some like data analysis like exploratory data analysis. So these are many many individual task enabled to generate a new hypothesis
and then uh after a hypothesis is generated it also needs to validate a hypothesis. Um so so so some people like
hypothesis. Um so so so some people like bio bio statistician will reason why a variant could cause a phenotype. um then
it will perform some p value calculation from these massive databases like gtags, pdb, bio bank and so on or it can also be a you know a real well biologist you can perform some targeted well
experiments. So overall there are just
experiments. So overall there are just lots of individual tasks involved when you validate a hypothesis.
So this basically motivates us to do this kind of a um kind of summary or like overview of biomatic research in a nutshell. So we we think it's in the
nutshell. So we we think it's in the outer loop. It's basically a generation
outer loop. It's basically a generation and a validation loop where a hypothesis is generated and then it's validated and the validated knowledge will be informed to update the new hypothesis and and
this kind of going through this iterations and within this outer loop in each each step there's an inner loop where there's gazillions of individual uh workflows like running a database
query running a binatic analysis reasoning experiments and so on that that is happening that are necessary to perform to generate hypothesis or to
validate hypothesis.
So if we if we view this as like this formulation then it's uh there's also a recipe for building a air biologist just emerge. So basically we need to build a
emerge. So basically we need to build a a genetic system that for hypology generation also hypology validation and also we need a generalist biomeic agent that can perform a wide range of
biomeical tasks.
So so so in this talk we will focus on on on this part um which is biom omni u uh which is our new framework that can enables to tackle the generalist
biometical agent that can just perform a wide range of different tasks. So we
also have different works trying to build a specialist agentic systems for hypos generation and hypos validation but uh this is for another day.
Cool. So let's del into biom. Uh so I want to first acknowledge my co-authors.
Uh so it's an amazing team effort. Um
special shout out to Serena Hansen, Jerry and Ma and also our um collaborators from across many different labs and and organizations and also our
adviserss uh here.
So as I as I mentioned like within the inner inner loop um after you have an idea that you want to validate or you want to generate a hypothesis this often involve many many different steps many
many small individual tasks and uh if we summarize a little bit it's actually a loop of reasoning and actions and basically you perform some reasoning okay you you you you want do this
analysis you perform it and see the analysis result and you do another step reasoning and you want to okay you to probably do another stop and it just continues until the answer is is is
arrived and we know that there's a lot of progress in our reasoning but actually the biometical action space is just undefined. So, so this basically
just undefined. So, so this basically motivates our research because when we first want to build this generous agent, it just there's no way for the LM to actually perform these actions.
Basically, uh it cannot do it cannot call machine learning predictors like biom models, bio machineing predictors.
It cannot do specialized databases queries. It can also not there's no
queries. It can also not there's no environment to run like binformatics uh softwares like scampy. um it of course it does not have access to the right
experiments and also the statistical analysis and so on. Basically in summary is like we don't have an environment uh without environment is basically like a scientist without a hand. It just keep
thinking but it cannot perform actions um so but we know like real scientists actually interact with the environment deeply and then in the environment
inform them to perform the next step.
So this basically motivates our uh first step which is to create an environment for LM agent. Um then the first natural question is like how do we create a environment because biometical action
space is actually very large complex and scattered. So how do we create the
scattered. So how do we create the action space in a in a in a unified uh fashion. So that that's a that's the
fashion. So that that's a that's the first problem that we want to um tackle.
So we we we utilize a new method that we propose which is like we are inspired by this bio archive where they have like 25 subjects and in each subject there's
like a lot of papers um so we can actually extract a 100 recent paper papers and these are basically a proxy of the snapshot of the research uh that
I belonging to this sub field and then for each paper we actually ask an action discovery agent that basically reads through each chunk of the paper and uh and after reading we asked the action
discovery agent what are some of the tools software and databases that are necessary like to generate or reproduce result in this chunk of papers. So after
we screening uh each for each subject we basically get a very comprehensive list of database tool and softwares. So
theoretically speaking, if we um implement everything into into environment and uh if everything is implemented really nicely almost to the 100% kind of fidelity, then
theoretically speaking with enough reasoning ability, we can an agent can reproduce everything like every to every basically 100 research paper in the past
you know in the past past few years for across all the subjects uh of inbound archive of course did this kind of theoretically but I think This G gives
the idea like this is a very systematic approach uh to generate an action space uh that span entire fields of of bio medicine.
So after we generate like a list of uh things this basically inform us where to look at then we need to implement these actions like we really need to make the
RM compatible uh with these actions. Um
so we think a lot about it and then we actually divide it into three categories of of of these genres of actions software database and specialist tools.
Um and and for the first six months of the project a team of five students are basically working uh day and night and try to implement like a lot of engineering work try to implement this
first version of the of the environment.
So days and night sorry both days and night. Yeah, but
mostly days. But yeah, uh and then and then so for the for the softwares for example, just give you a taste, it's like you know plink, scampy, gcta,
biopython, sk image, viana, l homer all of these packages spanning both like python r and and and command line tools.
And then we also have the uh the databases um which are like nomad bind open genetics uh PDB self hygiene regular DB clean and so on and so on and
uh and so the databases we have like two kind of kind of databases the first one this kind of web API and there's also a a suite of this this uh this like
there's no web API but like each raw raw raw data sets so these are like um like um uh for example put this chemical uh
libraries um and and and there's also a lot of small small files like cosmeic and stuff like that. So um so that's the databases and then we also have a lot of
specialist tools. So these are like well
specialist tools. So these are like well protocols like very specialized advanced AI models and also knows because we find there's a lot of things that just does not know because scientists have a lot
of implicit knowledge um and these are not captured in the on the on the internet. So there's no way for the RM
internet. So there's no way for the RM to figure it out. So for this kind of information we just kind of extracted codified from expert um and we implement ourself. Um so these include like I
ourself. Um so these include like I don't know eight automatic property predictions like some wab cloning like golden gate cloning protocols um divoc
also like universal cell embedding models parameter design and so on. So
there's a lot of these specialist tool uh as well.
Cool. Like overall we just get a single environment with diverse tools, software and and databases and it also bridge across sub fields of uh bio medicine.
And then after we have the environment, then the question is how do we build a general agent that can utilize this environment really well and can perform all kinds of different tasks. And so so
so the so for example given a query from the from the biologist um the first step is the retrieval step because our environment is like massive. So it
cannot put everything into the context.
It will just be a waste of tokens because most of the tools sorry most of the tool are just not relevant for the question uh at hand. So that that's why we first need a retrieval step where we
can retrieve the the the the software the databases and the fetchless tool that are needed to answer this question and we experiment across different kind
of a retrieval method and we find that the embedding based um rag based uh approach actually fails often uh so because biological knowledge a biomedical tool use is like very
requires some kind of reasoning ability as well so that's why we use a based tool retrieval almost like a sub aent that can perform the tool retrieval uh for us
and then after we retrieve the tools we will do reasonings. So here is basically you give a scientist you I wanted to do this and then you have these resources gener generate me a plan. So the the
plan here is like pre-process you know identify differential reging load the gene set of comparison perform enrichment analysis. But these are this
enrichment analysis. But these are this is this is like a standard binatic um analysis but this just give you a taste that you know the reasoning um kind of a
it's like basically making a plan and uh and here we are we are we because biological task is like there's so many of them it's impossible to iterate over and every one of them to create a
detailed plan uh so that's why here we are using like a the using ability to automatically generate the plan uh as it goes because we want to make it general
purpose. Um and the plan is actually
purpose. Um and the plan is actually tracked along the entire process with this kind of adaptive replanning. So
when once a plan is failed uh sorry once a step is failed it will adapt readapt and replan it um as well. And lastly um so how do we perform actions like so our
environment is actually pretty specialized in the sense we have a lot of softwares and we don't define functions for softwarees because you can you can arbitrary define 1,000 tools for
each package but we want to make it you know general purpose and very flexible.
So we actually in in the LM we just tell the LM that you have access to an environment that have all of this software in installed. So, so and also
the the databases and some data path as well and also the specialist tools and then the question is like if if we have that kind of formulation how do we
interle the software and databases and preser tool by intelligently and the code is definitely the natural way to do it right because you can write a piece of code that utilize one
software another software and then you can also use call the databases interle it and also use some specialist tools So, so that's why instead of using like
a function calling which is the most popular method uh for to to use uh we actually use code as a way to do general task solving and another benefit of
using code is actually it can it can also have like more complex logic. So it
can do for loop it can use like if and else and so on. Um so so this just gives us more uh efficiencies uh it takes
fewer actions um uh to to achieve the same level of task and then this yeah at the same time it can also create a place
for errors to happen right it actually but uh as we know that all like all these closed source come like frontier lab are trying to improve agents uh sorry improve the coding
ability of the LM like crazy and now it's like the zero shot coding is actually pretty pretty pretty good. And
also even if it has arrows, our agent can also like automatically selfdebug and and fix the code as well. Um and
yeah. Yeah. And and that's al also another benefit of code is it's actually like creating new functions on the fly kind of thing. It's like a auto build new tools if there's no tool available
uh for that. So there's just a a lot of benefit of using code instead of this function calling uh approach. Cool.
So, so this is the the framework that we're using right now. We call this uh A1. A stands for uh agent. Uh so this is
A1. A stands for uh agent. Uh so this is the first step that that we're are doing. There's of course a lot of place
doing. There's of course a lot of place to improve, but I think this is actually already pretty uh robust and pretty sufficient uh as well. So basically
after a retrieval it will perform some reasonings. It will the tool will use
reasonings. It will the tool will use all the all these resources by using the code um and then it will after it will run the code uh execute it and then it
will return the observations and then it will do further reasoning and to use by coding and observations and so on so and until in the end we'll get a get the answer.
So so after that we basically we're now uh uh in evaluation mode. So the for the next six months so this project takes around a year and the first six months
is around building and then the the last six months is around evaluation um and benchmarking. So the first step of
benchmarking. So the first step of evaluation is we systematically benchmark using using existing exam style uh Q&A questions. So like future
house has done a lot of amazing work on trying to create a a solid benchmark uh to evaluate LM capability for for biology research. So we we picked two
biology research. So we we picked two subtracts of feature house that are like tool dependent or like environment dependent because we want to test out our environment. So the first is the DB
our environment. So the first is the DB database query. The second is like
database query. The second is like sequence manipulations and we use like 45 questions from this data set just for tuning the prompt.
And then after tuning we adapt it to a test set uh like like like unseen test set and evaluate our model and we're able to see that the biom actually has
human level performance for both of these tasks. Um and uh I quickly show
these tasks. Um and uh I quickly show the baseline here which is the first one is the LM the green bar here and LM is I guess most what what the most biologists are using today like the chat GPTs and
cloud and things like that and then there's also the react which is like a but with this agentic framework but there's not much tools and then this is
the react plus code which is the coding agent so you can perform coding and react plus literature is the literature agent And then this is a coding plus
literature and biom react is basically an ablation of the bioni where it can only use um yeah so it only use like the this two function calling approach
instead of the coding approach.
Cool. So the future house is still like you know exam style uh and also it's like with some uh sorry so future house is still like we are using some
development set to to tune the model to tune the prompt uh slightly. So of
course you you are saying you know the test set performance will be great right so so then we really develop like have a completely unseen benchmark humanity last exam there's no development set
completely zero shot and we're able to see that the performance is actually really good as well so the actually has only 7% accuracy but biom has around 17%
accuracy so so there's a like a very significant boost of performance so humanity last exam is like as as the name suggests it's like they are like
really difficult um tasks uh like you know if the exam is solved the humanity is kind of a is great right so so so so that's why the accuracy is still not
like 100% but it it has shown significant improvement then uh after we have done this kind of exam style benchmarking we're also
thinking um but these are still not realistic right these are exams But real biologists they don't care about exam they care about real tasks. So here we
we we then we build around eight benchmarks like new benchmark that we created uh that really evaluate the agents capabilities uh uh for real
biomeatic research tasks. So this
include the v variant priorization. We
want to identify like a causal variant from a list of potential variants.
Similarly for causal gene detection given a a G was low side there's a like like a list of potential genes we want to identify what is the causal gene in in that neighborhood and then perturbation screen design is given a
cellular context of interest like T- cell exhaustion for example we want to design a gene panel of 50 genes for my crisper screens
um and then single cell orientation like I have a single cell data just annotated cell type for for each cell and microbiome disease Texa is identify the
taxa that is associated with the disease uh uh using this microbound private data sets and then for drug repurposing is I have a real disease given a list of potential drugs I want to identify the
most likely one for repurposing and for real disease diagnosis is I have patient with phenotype like a list of HPO terms and the WGS returns gene mutations blah blah blah like like a set of genes and
give me that diagnosis and similarly for patient gene prization it's giving a patient list of phenotypes and WGI return gene mutation what is the color of genes. So you can see on these eight
of genes. So you can see on these eight benchmarks they span across um you know genetics genomics microbiology drug drug and patient. So span across entire
subfields about bio medicine and we able to see across the board biom has the best performance much better than our um um much better than coding agent and and
slightly better than our you know this is an ablation. So it's it's still equipped with biom environment uh but just uh different meth different agent architectures uh to do it. So so we
actually we just perform it once because it's pretty expensive to do it. Uh we
only have budget for like just performing once but but we are we are really kind of happy because you know the performance is like really robust.
um it's probably statistically significant um um like for a single method that can outperform across eight new benchmark that is kind of a
completely zero shot.
So we're still not satisfied with these um benchmarking because benchmarking is like benchmark so it's like it's not real you know deployment. Um so then we
after we have that we then we collaborate with a few labs uh at Stanford trying to really see how it can impact the real world uh research uh for
scientists. So the first chunk of the um
scientists. So the first chunk of the um for first categories of the impact is on the web app protocol design u because uh as you probably know like for the wab scientist there's actually a lot of part
is also uh relatively digitalized which is designing the protocols um and you know which involves a lot of literature research and also involve a lot of tool
use as well. So here we work with Lung's lab. So Lung is a is a you know famous
lab. So Lung is a is a you know famous like crisper uh person. Their lab has done a lot of crisper stuff and uh for this experiment cloning is basically one
of the fundamental task uh uh for for web val scientist and here the one of the scientist PhD student Jerry um basically had this
question you know I have this plasma sequence I hope to clone a guy targeting a human B2M gene into this plasma can you give me a step-by-step guidance on
how should I perform cloning and the bion basically perform form after given this it given go into this kind of agent uh mode I just perform the plasma
analysis gang design or design or analing blah blah blah and then it'll get you the final plasmid map uh assem assembles as well so it overall takes
around 8 minutes uh to complete and then this is the output it's like a stepby-step uh wab cloning protocols uh it has the you know the single the gi
sequence uh it has very detailed steps and it also has the plasma map as well and then after that a scientist actually follows the protocol really to perform the cloning
and then we are able to see that actually the cloning is uh successful um the sequence is aligned and you know everything is good so so so and Jerry uh
is very happy about the results and you know because it currently it takes several years of training and even then bal just spent hours navigating pools like Snapgene, energy, primary blast
just to design one construct. But you
know, biom does it all in 10 minutes and and uh and it's it's a it's pretty uh sufficient and and since then we actually we further develop a new benchmark like
open-ended uh evaluation like human evaluations to really see the performance of biom for cloning kind of like experiment protocol design tasks.
So these are including you know so the the the realistic open-ended task that Jerry kind of encountered in the in his day-to-day uh biomedical research and on
these 10 open answer calling scenarios biom actually has similar accuracy as expert basically what we did here is like we have four kind of a I guess
person the first one is Alan which is the just raw chatbt I mean cloud in in this case and then we have Biomy and then we recruit a trainee
who is a who just graduated from master at Stanford. She also has done um a lot
at Stanford. She also has done um a lot of uh um kind of cloning experiment before but not a lot but has done some cloning experiment before but definitely not not an expert um and she's also
premed and she she's you know going to apply for PhD and things like that. Um
and then here the expert is a like a postto at Leon's lab who has done cloning for the past eight years like like it's like a real real expert in in cloning experiments and uh we asked
Jerry to do like a blind review like so for each each of these four it open-ended tasks and generate answer and then Jerry basically look at each one
and then give a score how complete is the colon protocol how accurate is the colon protocol and And then we were able to see that actually the biom has similar performance at as the expert and
much better than trainees and and of course much better than so just like hallucinate all the ways. Um
um but uh but uh you know in contrast biom has expert level performance and I think this gap is where the most interesting is thing happening because this means that biom is a great
educational tools uh because for typically for a trainee it usually when when they how to learn to do cloning is by engaging with the the high the I don't know 50 or PhD student and just
keep bugging them right and then try to learn how to do the this cloning experiment and Now you can actually just talk to agent that it will guide you and it will teach you how to do how to
perform this expert level uh cloning protocols. So so I think this is a a
protocols. So so I think this is a a great example of how agent is like a democratized this kind of a scientific expertise um like a education tool.
So that's for the well experiments and uh we also have conducted u uh like another chunk category is the dry lab uh experiments. So, so this include um for
experiments. So, so this include um for example in this case we work with Michael Snider's lab as you probably know Michael Snder has a lot of wearable
data wear variable sensors um so their their lab has gazillions of data in there just many most of them like just left unanalyzed uh because they don't have enough biophmentation to analyze
them so Ma is a bopmentician at Michael Standards lab and uh and she gave us this data set she actually performed the same analysis just two like a few weeks ago
um before the study is done here and then give given this data set it's a very complex it has a 458 like raw excel sheets like very complex files and then
it basically we asked the biom omni to analyze the data to process it and to generate me some interesting uh conclusions or or ideas um so so what
biom does it you know it goes through very detailed steps it actually takes around 35 minutes to get this set to get this task done and this is like
completely autonomous u there's no human interventions u at at at this task. So
it able to like analyze different temperature cross subject comparative analysis you know identify thermogenic timing characters characterizations and so on and so on and it generate a lot of
figures and plots and in the end it generate a report that summarize the key insights and and interestingly the key insight is also what Ma has found like
as one of the co key insight uh after she has performed it for like two three weeks uh with two to hours uh uh per day. And uh so this basically shows a
day. And uh so this basically shows a dramatic productivity boost uh for this uh routine and bformatic analysis. Um
especially in this complex data uh regime where this like 458 uh raw raw Excel sheets it just takes human forever just to chunk through them and the Asian
is like very easy to automate them.
So, so that's for the uh one dry lab kind of case study. Another dry case study is we also perform on this single RNA and single attackic data to try to generate some interesting hypothesis. So
this task is much harder uh because it requires much more domain expertise uh because single taxic data you need some specialist tools that you know a typical
um LM or typical scientist may not may not even know. So here henchan has done amazing uh just to because he he's an expert in this single cell analysis. So
he know exactly what is the tool that he wants to do. So for this data set he created a very com very detailed instructions like a detailed plan on how
to perform uh analyze the data. By
detail I mean you know this use this tool use that tool and you know to do this to do that but it's still kind of kept in a relatively high level. Um it's
just a it's it's much more detailed compared to like a just analyze it for me this kind of a this kind of a prompt.
It's like a it's there's like roughly 20 lines of of in instructions and we find that the longer the instruction prompt the more detail it is the agent just
perform much much better. Um so so I think uh in in this like very difficult u task um um we we need this kind of detailed instructions. So, so here it's
detailed instructions. So, so here it's able to basically follow the instruction and perform all all of these different steps and it gets you an answer and it also generate a lot of cool visualization and that makes sense and
in the end it generates some interesting ideas hypothesises although it's not discovered yet because it's we haven't really tested out but these are like interesting hypothesis that scientists
can like act on it. So for example like the novel transcription factors AOTS2 um ZFS3 and and things like that that show high regulatory activities across
multiple skelter lineages.
Um and u and these hypothesis are like not covered in original author's paper.
Um and then we ask experts like Jin and Pung who are like domain expert and they exam the entire trace they exam result and and they all think this is like making sense.
Cool. And bound me is just also more than just a few case study that we showed. And since then we all just play
showed. And since then we all just play around with it and it's able to do a lot of different crazy stuff like interpreted a variant um predicting the
admat properties of of a molecule. You
can download download some data sets from the web doing like literature research, crisper screen design, cloning protocol design, reg diagnosis, prodis
docking and so on and so on.
So motivated by that is we build a free web platform uh for scientists just to enable scientists to really play around uh wi with with it and to explore and al also to inspire us uh what is the real
capabilities what are some of the cool use cases um for biomly um so so so this is the this is the the web platform uh
it's completely free it's academic so we have it's saved at Stanford servers Um
and then and then yeah so maybe I can I can give a very quick lab demo on um on the on the on the interface. Uh but you can also explore more at this biome. EDU
you can register today and we usually in 23 hours we will kind of uh let you in right now.
So so here is the interface. Basically
you can ask uh like like a question after you log in. You can just type in the question here. You can upload your files and here is the is the cloning protocol design questions. Uh and then
it will basically goes to this tab where the agent is working. So this is the this is the interface that you talk to agent and this is the interface the agent is doing the work and so you can
see that it will make the plan it will you know execute a lot of different code um and uh it will perform it has different you'll have um you know run
the code automatically and generate this observations here and you'll plan the next step and write full like more code and uh your you know for this is the guy
design you'll generate create the guide design sequence um and then it will annotate the the annotate the plasmid um and um you know using golden gates to
design some protocols and so on so on and in the end after a few iteration it it finishes and it will generate a report here. So the report here it has
report here. So the report here it has um a detailed report. You can also download the the the data like faster files and the step-by-step coing
protocol design and also some the reference like how did how how does it arrive at this answer and then of course you can then even
like further download the data load from previous session. And here for example,
previous session. And here for example, you can download it. Um,
and then and here here's the the raw data. And and usually after it's it
data. And and usually after it's it finishes, you can there's also a um like a Jupyter notebook where you can basically reproduce the result and
modify the code and and so on.
Cool. So, so this is a very very quick um demo on boundary but I think you should just try it out and then and then see how it goes. Um and also let me know
if there's any feedback as well. So so I guess that's the bounty. Um I also just give one quick conclusion on what we're working on uh next what we're excited
about. So the first thing is definitely
about. So the first thing is definitely uh bound unified action space. There's
of course one direction is trying to expand the D direction action space and we and we wish to do it through this community effort where we're we're going to open source pretty soon. Um so we
want we want the community actually to build build up these kind of a standard tools and databases and softwares that then together we can just benefit all from it. Um so that's the action space
from it. Um so that's the action space but I think action space is not enough.
We also need the action uh we also need to learn how to use this action really intelligently and so this basically reasoning abilities uh uh I think
currently given a task human scientists have its own way of doing things but maybe there's actually a better way of doing things maybe there's a more optimal way it's just we don't we have not explored right now so this is very
very interesting and we're we're exploring the idea of training like reasoning agent like training the underlying LM to perform the reasoning and to use this action in biom really
well. Um such that it can intentionally
well. Um such that it can intentionally discover a new way of of performing a task. Um and uh and and this basically
task. Um and uh and and this basically require like using reinforcement learning and but there's also a gazillion of questions for example what are the verifiable task what is the
reward and how to efficiently scale up because there's very limited data points in in bio medicine and uh so on. So this
this is this is one direction that we're we're very excited about and another direction is definitely the wab the wab stuff like how do we really if we want
to build a truly AI biologist where there's no biologist only just need to wake up every day and just said help me do this and then and then in another day
just finished it we definitely need the wab component as well um so so this would require some interface between the wab and dry lab um and I think this is also a very exciting directions. It
could involve robotics or it could also involve like programming language of experiments. Um yeah, but these are
experiments. Um yeah, but these are exciting directions. Cool. So I guess
exciting directions. Cool. So I guess that's my talk and I guess I leave around 8 12 minutes for question answering. So if we have we have special
answering. So if we have we have special time and I want to thank my co-author again and uh here is also my contact information. Feel free to contact me.
information. Feel free to contact me.
Yeah, thank you.
Thanks so much Kessen for for the great talk. Right, there are a lot of
talk. Right, there are a lot of questions there. 30 34 questions and
questions there. 30 34 questions and then lot more questions on the on the chat as well. So let me try to get some
of the questions out. Um if given well what is the computing resource now looks like uh and and you said that it used at
uh you know the the computer at Stanford. Um yeah yeah yeah currently we
Stanford. Um yeah yeah yeah currently we have around like uh so we we're serving it on like AWS uh kind of cloud providers and then we have around like
100 instance um that can you know each user comes in it will randomly autobalance it to like one instance so I think currently it's kind of sufficient
right now um and also the cost um it's yeah it was pretty expensive for the LM cost uh but we're lucky to get to got a sponsor which we're actually going to
announce next week as well uh to help us cover the RM cost. Um
um but yeah, so overall cost relatively is now manageable. Um but indeed there's more and more user coming in every day.
So uh yeah, we I guess we'll deal with that until it's becoming a bit unmanageable. Yeah. Right. And
unmanageable. Yeah. Right. And
with this kind of limitation, would it be able to take on the raw files whole
genome sequences from like UK biioank or DB app and other things to to transform it into, you know, into uh structural
data. Yeah. So indeed the web platform
data. Yeah. So indeed the web platform has a lot of restrictions. So like each instance only has like 16 GB of the
memories. So of course like it's it's
memories. So of course like it's it's it's not much right. So um so if you want to apply on like very large scale genomic analysis it's not possible right
now uh in at least in the web platform but in the future potentially and also because we have the open source version.
So for the open source version it's it's it's like very kind of a you know uh you can just use your local computing
environment uh to do it right and what is the current AOM uh that are you you are using is it chat
GBT or is it anthropic or what what is it? Yeah. Yeah. We we use the the
it? Yeah. Yeah. We we use the the anthropic as the default model like the cloud because we find it has great coding ability. it has a lot of
coding ability. it has a lot of biomeical knowledge and it also has these um agentic kind of abilities. So
that's why I use cloud. Um yeah.
Right. Right. So the next question is um from uh from uh Tim Slidell. I
understood that you didn't use lang chain and implemented your own framework code act. Can you explain the pros and
code act. Can you explain the pros and cons of these different framework and approaches? Oh, so so actually we we
approaches? Oh, so so actually we we also use uh I mean we we use the langraph framework as as well um to implement the Kodak framework. Uh but
honestly um the Kodak framework is pretty straightforward. So did we
pretty straightforward. So did we actually don't don't need to use longchain um at all but it's just because originally we actually try various different kind of a uh biom
agent kind of a formulations and some of them are like really extremely complicated and these are where we use kind of long graph to manage it really well. Um but for the final one that we
well. Um but for the final one that we have right now it's actually really straightforward. So, so for that
straightforward. So, so for that although we currently still use the the long graph uh like the state and things like that but actually uh we we we also have a version that does not use long
graph at all um that that's kind of working. Yeah. Yeah. The question is uh
working. Yeah. Yeah. The question is uh what is your experience of having a successful way of connecting different software different backage and database?
Does it need to be more like a classical engineering linking database or does it like purely you rely purely on AOM understandable documents so that agents understand how
you know to connect? Yeah.
Yeah. I think we are on more flexible kind of approach. It's like basically the RM we just give the RM here is the list of data sets. You can write code to
explore the data set and understand it.
So it's like very free form. Um because
indeed the data set here is like extremely complicated different formats, different you know different columns and I think it's just too much to to to do
an engineering to make them into a single uh kind of a normalized format.
And the RM is actually really good at this kind of more flexible stuff. You
just need to give the file names, give the file description, some meta metadata about the file and it's able to basically figure out on its own.
Right. And the next question from Bang Angela. So um
Angela. So um when when also when working with biom are there any features to remove the blackbox effect of AI Asians trying to
understand the logic step resources trying to use to make decision.
Yeah. So you mean just understand the the intermediate steps? Uh yeah. Yeah.
So, so I think this is very very interesting because in the in the web interface you can actually see everything um like step by step what is the exact code and it's like completely
reproducible. So I think this is another
reproducible. So I think this is another benefit of the agent. Um it's like it's fully reproducible. Uh like previously
fully reproducible. Uh like previously if a scientist trying to do it because it it spans a couple of days you sometimes you create another file and then you create that file and then you forget about it and then you search some
search some databases. So you some browsers so a lot of implicit actions are not documented uh anywhere. Um but
for for for agent it's actually documented everywhere. So you can
documented everywhere. So you can actually go in backwards and then search how did I get this answer and it will pinpoint exactly the code that can give you the answer. So I think I think for
that um agent is this is another benefit of of the agent um that where you can examine the trace very clearly.
Right. So the next one is um about how if Biomi doesn't have access to recent develop tool in its environment,
how easy is it to how much engineering work does that need to in that need to involve? Yeah. Yeah. It's actually
involve? Yeah. Yeah. It's actually
pretty easy to to to include. We
actually in in the open source uh project which we plan to release is like um just a few lines of code where you can just add add a new tool and add add
a new software and add a new database.
So it's a pretty easy to use uh framework. Um and of course uh um if you
framework. Um and of course uh um if you want to add it so I think there's a one one one way is to add it permanently into the the the environment. So for for
that you will need to make a contribution and make a PR and that's also very very straightforward and another thing is you just want to use your own private assets that that's also
very easy to do you can just add a few lines of code as well right uh also adding one of uh the questions of one of
my friend team regarding the tools um and software available that you package when there's a new release of a versions
like scan by uh how you rebuild automatically or is there any way for doing that? Yeah, I think this is a very
doing that? Yeah, I think this is a very big engineering problem because it's like indeed uh ideally um because the because the version changes for
individual softwares can also potentially lead to the code changes. Um
and then and I think this knowledge about this recent update is just not that refreshed. Um so it make it make
refreshed. Um so it make it make problems. Um and also it created like the management problem as well like for example if you have using this version
for for this uh task but in two months actually our version is actually up bump it to like another new version. then
suddenly this old version become not not really reproducible. Yeah. So I think
really reproducible. Yeah. So I think there's a lot a lot of this kind of engineering stuff we need to make sure like everything is tracked everything is like traced and u to make it adaptive
but currently we just go through the easy way just set a set a version and don't change it um yeah but yeah that's the
and I I I I myself also find reproducing reproducible is actually also a problem of a because it depends on the way the
user brought maybe He miss a comma. He
miss he adds something. So, so you know depending on the mood we ask the questions maybe we got totally different answers from that. Exactly. Exactly.
Yeah. There just so much problems. Uh so that so there's a a questions from our our friend Sisha.
Have you tested by Omni on brand sale type annotations?
uh they are very convoluted and difficult to annotate. This is expert you know this is question from expert in in the field. Yeah. Yeah. Interesting.
Yeah. we we we we definitely haven't done the brain specifically, but we have done some cell type annotation tasks and uh I think it's able to do it relatively well, but if you want to talk about like
very specialized um tissues and things like that, you probably need to input a lot of the detail instructions like how you as a scientist will perform and I
think that will help the biomly perform like much much much better. um uh
because uh this way the agent can uh just follow follow like follow your behavior to get things done and then you just need to just give high level instructions. So it it still saves you a
instructions. So it it still saves you a a lot of time u but you just need to make like additional effort in the in the prompting part to make it specialized and you behave like an
expert. Yeah. Yeah. Uh the next question
expert. Yeah. Yeah. Uh the next question is from uh Fasad. Is there any possibility for adding human in the loop?
Uh yeah. So actually that's a yeah so in a web platform right now it's indeed like a human agent interactions depending on different kind of a granularities right you can also maybe
in in the intermediate step you just like stop it and then make changing the plan and uh even like write your own functions like you know in intermediate code you can actually write your own
code but I think you know for that we don't support it right now currently it's in like a like more high level human agent interactions You can follow up questions like tell me
more about this. You can actually also u change like a one step and then ask it to redo it but it cannot like you know like intercept in the middle and then modify it and then continues. I think
that's another feature that we have not had not uh constructed yet. Right. Okay.
There are a lot more questions you know the more I answer the more questions we have there. Now we start with about 30
have there. Now we start with about 30 now when we I mark answer and now it becomes 41. So I think it's time to stop
becomes 41. So I think it's time to stop uh the um the the talk and guess this is this is amazing. Thank you so much for
for this amazing talk. Cool. Thank Thank
you so much for having me and really enjoy the the discussion and yeah, let me know if there's more question you can send me by email. I can try to um answer it as well. Of course. Of course. Thanks
so much, Cassin. And thank you everyone.
Thank you. Bye-bye.
Loading video analysis...