Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future
By Lenny's Podcast
Summary
Topics Covered
- Highlights from 00:00-18:23
- Highlights from 18:13-39:02
- Highlights from 38:59-58:52
- Highlights from 58:40-76:45
- Highlights from 76:34-93:50
Full Transcript
In 2023 when I started, nobody said anthropic and claude and coding in the same sentence.
I want to go back to the beginning of anthropic. I remember dealing, man,
anthropic. I remember dealing, man, these guys have no chance. OpenAI is so far ahead.
At the time, I saw people were starting to use these models not just for code autocomplete, but actually writing long form code and [music] sat an opportunity for us to train Opus 3 to be better at.
That was the inflection. [music] I
always think about Opus 45 a year later during winter break when everyone was home able to code.
What was magical about Opus 45 is we also now not just had a model but a vehicle a great product experience like cloud code. Opus 45 wouldn't have had
cloud code. Opus 45 wouldn't have had that moment without a product like cloud code and cloud code wouldn't have had that type of adoption accelerated without opus 45.
I want to talk about how the product role is changing for my team. The way to drive user value is to figure out the right user feedback. The evals, we actually have a
feedback. The evals, we actually have a saying on the team of evals are the new PRDs.
Something Gary Tan's been talking about.
If you are willing to spend $100,000 a year right now in tokens, you are living the way somebody in 2028 is going to live.
You have to sweat the tokens as much as you sweat the pixels. You have to be using the models to come up with good and great and better ideas. And there's
no substitute for that. People need to be more ambitious with AI tools these days because they're just capable of so much.
One thing I ask the team is let's say Claude 8 comes around. What changes in what users do? What does that mean for how you're building today?
Today my guest is Diane Penn, head of product for the AI research and labs teams at Anthropic. She joined Anthropic as the first technical product manager over three years ago, which is a
lifetime [music] in AI time when the product team was just five engineers.
She's helped ship every model adropic from claw 2 through fable. She's also
helped incubate and launch claw code, MCP, skills, claw design, and also core capabilities like computer [music] use, tool use, and reasoning. It is always such a treat and so mind expanding to
get to talk to someone who's at the very center of AI and product management.
It's hard to imagine someone who has seen more of where things are going than the head of product for anthropics [music] research and labs teams. Before we get into it, don't forget to check out lenniesproass.com
for a year free of the hottest and most beautifully crafted AI products in the world available exclusively to Lenny's newsletter subscribers. With that, I
newsletter subscribers. With that, I bring you Diane Penn.
Diane, thank you so much for being here and welcome to the podcast.
Thank you, Lenny. It's so nice to see you again.
I want to go back to the beginning of Anthropic, uh, the early days. I
remember when Anthropic first launched, this was, I don't know, years, the first model when it launched years ago, three years ago, something like that.
It was three years. I remember just like
three years. I remember just like feeling that man these guys have no chance. Open AAI is so far ahead every
chance. Open AAI is so far ahead every like how what are they thinking? How is
this possible? Open AI has won. It's too
late. Uh things are very different now.
The latest number I saw was Anthropic was making like I don't know $50 billion in ARR. That's like what companies used
in ARR. That's like what companies used to go public at like very successful companies went public at 50 billion in valuation. Anthropic reportedly is
valuation. Anthropic reportedly is making that every single year. You
joined as one of the earliest PMs. There were something like five engineers when you joined. The model hasn't hadn't even
you joined. The model hasn't hadn't even [clears throat] launched when you joined. What was it like in those early
joined. What was it like in those early days of Anthropic? What's something that might surprise people about what it was like at the beginning?
I think a big part of what's made anthropic today actually has been very much the core of even the early days. So
I joined in 2023 like you said we had five product engineers. There was one engineer for the entirety of our API business if you if you believe. Um and I
think a big portion of it was the culture was really strong and I think this is something I emphasize for folks who are interested in the company. Um
really do walk the walk of um the mission and the culture and the values.
Um, and the energy was very much like a startup. And I think you're right. We
startup. And I think you're right. We
were very much trying to find our identity in the early years. Like I
think there's one piece around the technology, but how does that technology bring value to users, bring value to society, and what could it possibly be?
And I think the early years were us exploring that in different ways. Like
we did start with like cloud.ai I like another chat chat assistant and evolving into things like tool use. Um I think one of the moments where really we
started to get into our groove was shipping things like Golden Gate Claude.
I don't know if you like remember that.
No.
Um so this this was actually up for about 24 hours or so. Uh we had just published one of our um early
interpretability research in early 2024.
And one of the examples was essentially you could have what's called like features of the model within the layers which uh express certain types of uh
thematics. So one of the one of the
thematics. So one of the one of the themes that the researchers was able to identify was uh let's say bullet point writing. Another one was people and
writing. Another one was people and places. And one that really came up
places. And one that really came up frequently that uh resonated was the Golden Gate Bridge. And so when you actually uh essentially dialed up that
feature, Claude would obsess about the Golden Gate Bridge. So meaning in every one of its responses, it would come back and talk about the Golden Gate Bridge.
So if you said like, "Give me a recipe for making spaghetti." Uh it would say, "Here is a recipe, and the orange color is just like international red that the
Golden Bridge, Golden Gate Bridge looked like." Um, and so it was like really
like." Um, and so it was like really quirky and we we we very much wanted to in that situation just bring that user bring bring it to the masses and bring
it to people who are starting to use claude and uh so the entire uh experience actually we spun up on our cloud.ai I website within 24 hours and
that took like engineering, product, design, uh our like research teams all working together and we were really really proud of it. I think it maybe reach only 2,000 people to [laughter] be
honest. Uh but it it made us feel like
honest. Uh but it it made us feel like oh we can actually bring new user experiences, showcase our research in a way that's different and authentic to us
and in a very startupy like pace. That
to me was like one of those like maybe hidden inflection points of we were starting to find our identity that we could build products, build experiences that were different for what our
competitors had seen, what was already out there. And I think that obviously
out there. And I think that obviously labs, clog code, etc. Like we then started to identify ourselves as what we actually think the world uh how to think
about AI, how to bring that closer to the public. Um but it was a very bottoms
the public. Um but it was a very bottoms up culture. And so that entire
up culture. And so that entire experience was very bottoms up. I see
engineers, I see uh designers donating time to work on. Um, and so I I like to always use that as example of like what the day early days were like, but the culture and and and the values have very
much I think stayed the same since those early days.
This episode is brought to you by our season's presenting sponsor work OS.
What do OpenAI, Anthropic, Cursor, Versell Replet Sierra Clay and hundreds of other winning companies all have in common? They are all powered by work OS. If you're building a product
work OS. If you're building a product for the enterprise, you've felt the pain of integrating single signon, skim, arback, audit, logs, and other [music] features required by large companies.
Work OS turns those deal blockers into drop-in APIs with a modern developer platform built specifically for B2B SAS.
Literally, every startup that I'm an investor in that starts to expand upmarket ends up working with work OS.
And that's because they are the best.
Whether you are a seedstage startup trying to land your first enterprise customer or a unicorn expanding globally, work OS is the fastest path to becoming enterprise ready and unblocking [music] growth. It's essentially Stripe
[music] growth. It's essentially Stripe for enterprise features. Visit
workos.com to get started or just hit up their Slack where they have actual engineers waiting to answer your questions. Workos allows you to build
questions. Workos allows you to build faster with delightful APIs, comprehensive docs, and a smooth developer experience. Go to works.com to
developer experience. Go to works.com to make your app enterprise ready today.
What are some of the other um big inflection moments as you think about just Anthropic going from just this like lab that's trying to compete with this juggernaut of OpenAI at that point to what it is today? What are some moments
that stick out of like wow that really changed things? Definitely when we were
changed things? Definitely when we were training and uh testing uh Opus 3, I think that was the moment when the
company I think we were less than 200 people still at that point and it was very clear that we needed and wanted to
create a frontier model and a uh that was very important in terms of like our ability to reach like users, consumers
and uh to showcase our research.
And we were looking for ways for also why should somebody choose Claude? And
that was like a core question and that was a core question we were getting asked in the early days. And I think with Opus 3, you know, it launched I
think early March 2024. But there was many many months of various teams across inference across research fine-tuning
pre-training that rallied at different points and towards a common goal and uh I think everybody that was involved was
like really proud. I remember uh being the PM, us uh the research leagues, myself, we were all in our um this was around December, so we were all at home
in our various uh um parents' homes and seeing everybody's background of like their childhood room and everybody was working really hard uh to figure out that like what are we training the model
for? Is it showing up the right way? So
for? Is it showing up the right way? So
I think that was really powerful in terms of just building a lot of trust and a lot of our research leads have actually uh from that time are now like
leading reinforcement learning leading our character work alignment work. So
that that foundational trust I think also helped us work well now with any of our production models across product and research because we were working just so
much in the trenches together in the early days. And then I think there were
early days. And then I think there were things like identifying that coding was important. Right? In 2023 when I started
important. Right? In 2023 when I started um nobody said anthropic and claude and coding in the same sentence. I think
competitor models like GPT4 at the time was used a bit for coding but it was one of many use cases. And one thing that for example I saw was
people are starting to use code uh these models not just for code not just like code autocomplete but actually writing long form code and is that an
opportunity for us to train you know opus 3 to be better at and it ended up being a relatively smaller change from a training perspective but it ended up
helping us differentiate in the early days uh competitively for users. and
actually bring a lot of the very early cla enthusiasts and developers because we were uh providing a value that they didn't really think was possible at the time.
It's so interesting you talk about Opus 3 like that's so long ago and just like it's hard to think that was a big inflection and so this is really interesting to hear that that was internally a big milestone. It almost
feels like this confidence you all built that wow we could really ship a frontier model which is now today so not great if you compare it to what we've got today.
What I always think about is Opus 45 which was and interestingly like a year later also during winter break when everyone was home able to code. Uh was
that another big milestone?
Yeah. Um Opus 45 was definitely another large moment. I think what was magic
large moment. I think what was magic about magical about Opus 45 is we also now not just had a model but a vehicle
which is like a great product experience like cloud code. Um one thing we say a lot on the team is you need frontier products in order to have frontier
models and for people to feel the magic of frontier models. And I think you know we felt the magic of cloud code for very
for for uh for many months before that.
Uh but the fact that the model essentially got to a level of intelligence where at a very broad level users can experience
both frontier intelligence in new use cases allow it to run things end to end in an agent manner. I think that was the inflection. It was actually both. I I
inflection. It was actually both. I I
think Opus 45 wouldn't have had that moment without a product like Cloud Code and Cloud Code I think wouldn't have had that type of adoption accelerated
without Opus45.
So kind of speaking on on this on this thread uh Daario interestingly if you look back at all his predictions he's just like okay coding is going to be solved it 100% in like a year something
like that. He kept talking about how
like that. He kept talking about how we're going to do code like AI is going to do all our code. And I remember everyone uh being like, "There's no way.
This is way too complicated. How is how is AI ever going to get really good at this very complex thing that humans do?
No, this is going to be humans for a long time." He was completely right.
long time." He was completely right.
Something else that he talks a lot about is this exponential that we're now that we're on. That's the way he describes it
we're on. That's the way he describes it now. We're like, we're on the
now. We're like, we're on the exponential curve. I remember not long
exponential curve. I remember not long ago we were new models were being released and everybody was like, "Okay, we're done. There's no more upside. It's
we're done. There's no more upside. It's
plateauing. It's over. There's no more room to grow." Uh, and now it's like the opposite. Now we're inside, like if you
opposite. Now we're inside, like if you think about the curve of the exponential. We're like inside of the
exponential. We're like inside of the exponential now, which by definition means every improvement is a massive jump because we're like on that hockey
stick part. What's it like just being on
stick part. What's it like just being on the inside of this crazy historic moment when AI is improving so fast, so much is
being unlocked? uh what is it like and
being unlocked? uh what is it like and how should people prepare for the coming acceleration of more and more improvement from AI? One thing I like to
say on the team is most of us weren't like actively working yet when the internet transitioned from this novelty to something that everyone can use and
it feels like that's just taking humans uh I think analogies are helpful and so like the analogy of that is I think a couple of things um number one is
adaptability becomes very important um I Think we we have evals. We have you know on the
safety side safety testing red teaming on the capabilities and product side new prototypes products like cloud code tag and others
but it's very hard to predict the exact moment or the exact model and so the adaptability of when you're faced with new information how do you then make
better decisions versus keeping the same plan. And so like that agility is really
plan. And so like that agility is really important. I think another piece is with
important. I think another piece is with that how do you actually be thinking very first principles and reason through what's next? What's the so what? How do
what's next? What's the so what? How do
we invest in new products? How do we invest in explaining the differences to users? So a lot of the a lot of the
users? So a lot of the a lot of the experiences I think of being in that exponential is that pace understanding how you operate and make better
decisions and then applying that first principles thinking to then do something that maybe we pull up a plan that uh we
would were expecting a few months from now but now the model can actually do uh and work on and actually bring that to user. So this is things like co-work
user. So this is things like co-work skills tag, you know, as the it's a very positive self-reinforcing loop. And I I
I think a big part of it also is just having the like trust in each other like making sure we have like we're we're thinking through the right decision making. We're bringing folks along. Some
making. We're bringing folks along. Some
teams might see the exponential feel it faster than others. So how do we kind of have the grace to bring the organization, the growing organization and company along on that?
So what I'm hearing here is you almost don't know what will be possible with every model release. And so the important things to focus on is being
adaptable as things emerge. Uh to your point, the product itself has to stay up to has to catch up to what is possible.
To your point again, just like it can do so much, but people may not understand how to do it and may not be able to do it. So the product making it easy and
it. So the product making it easy and even just like telling you here's something you could do feels like an important part. Is that roughly what
important part. Is that roughly what you're describing?
I I think so. I think um there's some really interesting graphs in the original scaling law papers and I think folks are very familiar with the scaling
loss in in the lens of um as you add in more compute and data what's called loss aka the loss from next token prediction uh goes down. And so it's a very smooth
linear curve of like the models get more intelligent as you scale them up. What's
actually also interesting uh in that paper is there are these like very uh different emerging capability graphs.
And so for example uh as you add in more data and you train the models with more compute you essentially see these actually discontinuous emerging
capabilities jump. So the models go from
capabilities jump. So the models go from 1 + one being a thing that it can't calculate to a thing that it can reliably calculate. And so these
reliably calculate. And so these emerging capabilities, this like some nature of like predictability is is is not necessarily everyone knows the exact
moment like you need the ebells to be able to assess that has actually always been a part of uh how this technology works and also what makes like things
like safety harder because unless you have the eval unless you have the systems to test um these jumps might actually happen and you don't know M
that's so interesting that you may have developed this like AI brain that uh can do something you're not even aware of and so part of the job is just uncovering wow we just got really good at this thing what can we do with that
I think there's like product overhang and user overhang like to to maybe put it in our um PM language even on today's models and I think there's like a lot
that uh we could be exploring on like our current opuses and definitely with like Fable for example temple and that that discovery is actually another part
of what's been in the early days of anthropics DNA and I think is also continuing to be a big part of how we operate in product in
labs and and across research. This makes
me think about something Gary Tan's been talking about uh president of YC. I
don't know what his title is. uh he's he had this interesting point that if you're willing to spend $100,000 a year right now in tokens, you are living the way somebody in 2028 is going to live
because by then it'll be really cheap.
Everyone can work this way. But if you there's this alpha opportunity right now to just live in the future, go crazy on token spend. Uh and so there's a big
token spend. Uh and so there's a big opportunity for people to learn what the future's like and also just build much faster. Thoughts on this idea of and the
faster. Thoughts on this idea of and the value of token maxing, let's call it.
Yeah, I think I I take more of like a almost product lens. It's almost like token spin is more the input and really the output is what you described of experimentation
and I think if we were orienting like goals around experimentation. I feel
like that that might be the better framing of the outcomes and therefore there might be different ways of achieving that outcome. I will say internally some of the most creative
thinkers, the best like prototypers do spend a lot of time with Claude with every new version of a research model that we have. And so there is something
around you have to be like using the models to then come up with good then great than better ideas and there's no
substitute for that. um it's very hard to come up with a perfect strategy without touching the technology when it's moving this quickly.
At the same time, I think there's other things that we could be doing like so one thing that we um do a lot is actually working in public at internally
within anthropic. And so in the early
within anthropic. And so in the early days when we had less product surfaces, there was a slack channel where everyone almost the entire company was testing
early versions of Claude and trying different use cases. Like people were not calling them use cases, but you might be asking it to edit an essay or uh to come up with the right way to send
this email. Like they were all different
this email. Like they were all different use cases, but we all worked in public.
And then what you would see magically is different users or different different folks on the team coming up with an idea and then other people trying different
variations of that idea and then within maybe 10 or so requests there was something magical or potentially a new use case
that emerges. And I think there's a lot
that emerges. And I think there's a lot in not just individuals figuring out by themselves how to use this technology. I
think we could be doing more to actually bring like that communal discovery when we do experimentation. Like
experimentation is not always necessarily a individual sport.
It's so interesting. Yeah. This idea
that we're just we're not sure what this is capable of or what we could do with it and it takes all this poking around and people trying things, hearing what other people are trying to figure out
what's possible. such an interesting I
what's possible. such an interesting I don't know technology or just like okay here's what oh I figured out it could do this thing what are you gonna do with that I think at a broad theme we know right we know that the models to write great
essays or you can write long form writing but individual pain points of what can you actually solve with that and bring it to like a user level that people can use um I think is something
that is more exploration or experimentation uh based following this thread you uh you oversee product for the labs team which uh is extremely cool. We've had Ben man on the
extremely cool. We've had Ben man on the podcast, Mike Griger who whom both work on labs now talk about labs. What is
labs? What's come out of labs? Many
people have heard of these things and how do they work that enables them to create such innovative ideas outside of even the core anthropic product team.
The thesis of labs in many ways is identifying and pulling the thread on the thread of discontinuous large bets
that might not be in the core road map and figuring out is there a there there and also what is the 10x 100x a
thousandx of the there there and so for example uh things like cloud code um I think I've heard of [laughter]
uh things like cloud code uh things like uh skills and most recently cloud design MCP the thing that we really try to
emphasize within the teams is especially right now there are so many things that could be built what does it mean then to have a discontinuous bet and I think one
approach that we're taking this year is you can be very strongly held opinion about the theme or the area and then more weekly held about the exact
prototype. And so like there is a
prototype. And so like there is a culture of experimentation.
Um there's a lot of the bottoms up like engineers on the team are very selfable um self-driven to test out different
ideas and sometimes uh we have a thesis and it might not work yet and so we then might revisit it in one to two model generations. And so this idea of like
generations. And so this idea of like these prototypes that actually end up just helping us learn like that's also valuable even if it doesn't lead to
something immediately shipping. And so I think that allows the incubation and like the charter of labs to really accelerate and see around corners more broadly for anthropic. It's so funny to
think about a labs within an anthropic which is already so innovative and and creative and just you know shipping like crazy that there's value to still creating a labs team within anthropic.
What enables labs to work as well as it has because you listed all these products and it's let's like what else has anthropic shipped it like feels like all the biggest wins almost. I'm sure
there are many that I'm not thinking about right now. What's what's kind of core to creating a successful labs or within within a larger company? I think
that team culture like similar to broadly at anthropic I think that team culture is very valuable. I think Ben
sets an uh incredible vision and pushes people to think about the 10x 100x of the idea and you know our the teams the
pods within labs is small. Sometimes
these ideas start with one engineer, right? And I think uh sometimes when
right? And I think uh sometimes when there's almost really large teams pursuing very ambiguous large ideas, you
end up actually being slowed down because of that. Um so I think it's culture. I think you know we actually
culture. I think you know we actually also select for folks who actually want to do that zero to one experimentation and it's not
easy. There's a lot of bets that we end
easy. There's a lot of bets that we end up turning down or turning off. Um and
maybe you know we revisit them in the future. Uh but that's hard. That's hard
future. Uh but that's hard. That's hard
when you pour your heart and soul.
You're acting as a founder for a bet and it's not working yet. Um so I think it's like that type selecting for that type of personality folks who are really passionate and deep about the zero to one.
So you lead product for the research team. You work with the researchers at
team. You work with the researchers at anthropic. A lot of people kind of get
anthropic. A lot of people kind of get an sense of what is research what are research what researchers do. I think a lot of people don't totally understand these very valuable people uh at all the AI labs. Uh the way I think about it and
AI labs. Uh the way I think about it and I want to help people understand help me understand just what are researchers doing all day. What I imagine is they have a hypothesis for how to improve the
model. They find data, they tweak some
model. They find data, they tweak some algorithms, they check adjust how it's trained, and they test it, see how it did, keep iterating, and keep trying to find ways to improve the model. Is that
roughly right slash help us understand what researchers are doing all day?
That's really I I think that's a lot of uh maybe the the like the more day-to-day. I think one piece around uh
day-to-day. I think one piece around uh researchers and like research organizations like at anthropic is
there's also a vision of the future like more broadly. So for example things like
more broadly. So for example things like uh I think even at the founding of the company researchers were talking about how do we get cla to you use a computer
how do we get AI to like navigate a screen right so there's a lot of actually very founderlike energy is how I describe it within researchers or
really bold and ambitious researchers um and we have a ton of those at at anthropic so there's one layer of vision of what this technology can go
and then I think on this other side of the loop there's also now that this technology or cloud is in people's hands how do we make it better today so it's a
medium and long term and a lot of energy thinking about that lens of the future and also in the immediate and short term what are the improvement areas we can make and so like I think you're
describing a really good sense of how do we make iterative improvements on different versions of claude the way that like my team works with researchers is kind of being very
integrated and embedded in in those loops particularly areas where there's a lot of impact on users. So this is
things like vision, computer use, coding, agent coding, tool use, test time, compute, things where there's a
direct user impact and then figuring out what are the ways to uh bring the user feedback and ground it
in a level that is understandable for user uh for researchers and also actionable for researchers. And I think that's the second piece is actually a
big part of the job and sometimes a hard part of the job. So for example, we might get feedback on cla.ai. Claude
hallucinated.
It's very vague. If you bring that to a researcher and you say, "Please fix Claude from being hallucinated." It's
not very actionable. And so part of the time of the team is understanding, okay, what's the trajectory of why that user gave that feedback? And it's like
consented. And so we we we look at okay
consented. And so we we we look at okay what should Claude have called tools in that moment or from its current knowledge or it called the right it
looked at the right document but it looked at the wrong facts. In the first case that would have been a failure on tool use. On the second case it would
tool use. On the second case it would have been a failure on let's say search or knowledge and search and search synthesis or it could be something
around alignment. And so bring that
around alignment. And so bring that level of detail to researchers coming up with like is this a big enough problem figure out things like evals to
then describe how we've improved it like those are the levels of actionability and it's the day-to-day language of the researchers. And so we try to stay very
researchers. And so we try to stay very close to how to bring that in an actionable manner uh between users to to the core model training and the research development loop.
I was talking to someone the other day about how feels like research AI research is uh the place to be now if you want to be very successful in life.
What does it take to become a really successful researcher from what you can tell uh you know not everyone can get in not everyone's brain is going to work this way but just say people are like hey I want to explore this career path from what you've seen what does it take
to to make it there researchers generally are research and product managers working with research or both let's do both but uh the researchers like you know PM's working researchers
also going to be very successful but it feels like everyone's trying to you know poach all the top researchers across every company so just I I know you're not an AI researcher, but just from what you've seen, just like what does it take
to make it in that in that career path?
Yeah, I think a lot of the most successful researchers and research leadership at Anthropic are folks who are really strong first principles
thinkers about problems. Like they reason through problems really well. um
who are just passionate about their research area and have a bold description of what that could look like
and then who are actually close to the details and so uh you know our like leadership our chief scientists our
heads of like fine-tuning and like RL folks are actually really close to the training runs and actually look at things like how the training run is
eval looking at the underlying data. So
like actually staying really close and be excited to be in the details I think have been like a sign of like really strong researchers and developing taste.
And I think like another piece is just like their ability to think big over time and be like very ambitious, right? like the
Dario like we can transform software engineering and and and the and I think uh going in that direction you learn so
much you get you had to shoot for the stars in in in many ways across u your ideas I think in order to be a a successful researcher
I I love just this meme of just be more ambitious comes up so often now which is so hard like it's it's easy to say that it's hard to actually just like how big can you and how that's so much of what AI now
unlocks. Just be more ambitious.
unlocks. Just be more ambitious.
Yeah.
Yeah.
I think it's thinking through it once or twice and to end and then being I think stubborn
about the uh area and maybe more uh loose around the exact like approach. Um
it it is a question we challenge ourselves with. uh but the technology is
ourselves with. uh but the technology is moving so quickly and so how do you make sure what you're building is actually uh forward compatible
and so it's also actually part of like I think the core product development loop to think bigger right uh one thing I ask the team frequently or how I think about
when we're building a product is let's say claude 8 comes around what do what changes in what users do and then what should what does that mean for how
you're building today? Is it going to be forward compatible to that experience, right? So like just grounding it's I
right? So like just grounding it's I think um being ambitious is very broad and so trying to like ground it in in some ways of describing describing that and also yeah everything heading in a
direction that all is cohesive and makes sense versus just ambitious in a completely different direction. Speaking
of ambition and cloud8, uh, Fable Mythos recently feels like hit this very new kind of tipping point with models where it used to be you have an awesome model,
release it. Hey everyone, welcome.
release it. Hey everyone, welcome.
Opus45 is out, everyone can use it.
Mythos went in a very different direction. It got blocked. There was a
direction. It got blocked. There was a lot of scrutiny, a lot of concern about what it was capable of. Uh, all the companies had to go make sure it wasn't going to hack into all their systems.
And it feels like now every model because they continue to get better will now have a lot more scrutiny and there will be more restrictions on who can use them which feels like a big deal. How do
you think about that? How does that change the way you operate?
I'm going to maybe leave the policy and the export control side to to folks that um own that and work on that. Um I think the product question and how we interact
with these internally is I think as you mentioned as frontier models become more capable the safeguards and the ways of red teaming and testing and the
pre-release process uh also needs to evolve and adapt quickly to to address that. And so one example is you know
that. And so one example is you know before fable models we didn't have as strong of let's say fallback UX's and
systems because our our our goal was to make sure that like there is asymmetrical benefit for this technology and to minimize like the downside or
like a severe risk of of it. And so we ended up building like fallback systems so that users will still get a great
response from Opus 4.
And so I think there's a piece around uh as we evolve and like improve safety systems. How do we continue to develop and deliver great user experiences?
I think there's more that we can do on both sides. And so you'll see us
both sides. And so you'll see us innovating, improving on what we call now the model safeguards package uh more and more in the coming coming weeks and months.
What's really interesting and just like unexpected here is creates this really interesting advantage for anthropic where you have access to the latest stuff and this is going to happen at every lab. Everyone's going to keep
every lab. Everyone's going to keep improving and it's it creates this unfair advantage within the labs to have access to the best stuff that other people can't yet outside of your control. You'd prefer everyone use it.
control. You'd prefer everyone use it.
So it's a really interesting this new feedback loop that's going to start where models that are so advanced are only accessible to certain companies and that's going to be a whole new unexpect it's like a second order effect of of
all these restrictions. Our goal is to be uh to develop these systems and the models to be as inclusive as possible.
Um I think our goal is to not have that happen uh for the general purpose general use like technologies and to make it more accessible. I think, you
know, it this is like one of our top priorities right now to kind of reduce what we're seeing there.
Yeah, that makes sense. I would imagine you'd want as many customers if people using this thing as possible. This
episode is brought to you by Mercury, radically different banking, loved by over 300,000 entrepreneurs and now with command. I've been a customer of
command. I've been a customer of Mercury's for over 6 years. I have never once thought about leaving. Mercury is
basically what happens when banking is built by product people, not by bankers.
They make it so easy, dare I say fun, to send invoices, move money [music] around, set up virtual cards for folks on my team. Does your bank have an API,
a terminal native CLI or an AI ready MCP server? I don't [music] think so. And
server? I don't [music] think so. And
just recently, they launched Command, a conversational interface built directly into Mercury, which acts as your financial operator. I've been using
financial operator. I've been using command to transfer money around to figure out what categories I've been spending the most money in, analyze my cash flows, and just today I used it to find out how much I've made from a
specific sponsor over the past year. I
just ask, [music] "How much have I made from X over the past year?" 10 seconds later, I have an answer. It is so freaking cool. Visit mercury.com to
freaking cool. Visit mercury.com to learn more and apply online in minutes.
Mercury is a fintech company, not an FDIC insured bank. banking services
provided through choice financial group and column NA members FDIC. I want to talk a little bit about how the product role is changing and who who is doing
well in this new world uh now that AI is such a core part of uh of our life. When
you're hiring PMs, product people, when you're looking at people that do well in today's world, what are some things that you notice? What are you looking for
you notice? What are you looking for more most? What are you looking for
more most? What are you looking for more? What's kind like trending up in
more? What's kind like trending up in what you find is important and what's kind of trending down? We actually on my team have not changed our hiring loop uh
for three years now. Um
so what we actually look for and the traits and how we evaluate uh generalists like PM's generalist like research product managers have actually
been the same. Um so I think some of those traits number one is first principles thinking and this is really uh rather than
pattern matching what you used to do in let's say consumer product or B2B SAS um but actually figuring out in this
moment for this user group with this technology what what is the user value is there an example that a lot of people hear first principles thinking they're like yes I about it. I'm good at this.
What is what's an example of someone having really demonstrated really good first principles thinking?
I think one example is I think you think of a product manager as I own product strategy and delivering user value as
but I demonstrate day-to-day by writing a PRD or writing a product vision doc.
And for for my team as like research product managers, the way to drive user value is to figure out the right user
feedback, the evals, right? That then can be a
right? That then can be a personification of that user need. So
like we do write some product documents and PRDs, but we actually have a saying on the team of evals are the new PRDs, right? because in order to deliver that
right? because in order to deliver that user value uh it's not that exact artifact that people used to write in the last like one to two decades it's a
new way of working and so the first think principal thinking would be let me figure out what is the thing I should do to achieve my goals rather than here is
a set of activities that I've done and therefore I will continue to do so the idea here is used to be have kind of an idea create a PRD talk to people about it. Align on the plan, design it,
about it. Align on the plan, design it, build it, ship it, see how it goes, iterate. What I'm hearing here is it's
iterate. What I'm hearing here is it's like, okay, here's some feedback about something that's wrong or an opportunity. Step one is the eval is now
opportunity. Step one is the eval is now how you define what the work is versus a PRD.
Maybe maybe step one would be uh understanding the user painoint. And so
the way to even access that user painpoint is different, right? In the
past, we might do a user interview and I think if you go like deep enough, you you might have the user walk you through their user flow, the pixels. Here, you
have to sweat the tokens as much as you sweat the pixels. And so, one activity we have on the team is reading the transcripts and understanding
uh what was the trajectories that failed very deeply to then say was this like a hallucination? was this claw being
hallucination? was this claw being overconfident. So like the theme of the
overconfident. So like the theme of the failure actually has a lot of nuance and then that allows you to build a description
a like sustained description of that painoint.
Uh so that could be essentially in a new eval and is the eval on distribution right is it capturing both the positive situations where this is failing and
also areas when it should actually not fail and then bring that back to let's say research so then we can make the improvements and actually measure the
quality of okay when we have opus 5.5 is this area improving or not is claude now able to uh identify the right places in
the document uh and pull the right synthesis out. So it's just the
synthesis out. So it's just the actionability like and and shortening the distance to actionability um for for our stakeholders and partner teams like researchers um to take action
on.
Is there an example of something like this where you found an issue or opportunity and then wrote the eval? And
what is what is the eval looking like in in most cases? uh what when people want to picture an eval what is that what is what do they picture?
We actually uh pioneered this concept within anthropic. So uh one of the early
within anthropic. So uh one of the early examples is the early cloud models were not very good at following specific
schemas. So like things like outputs and
schemas. So like things like outputs and JSON and uh now that is fundamental to claude being able to be a good agent.
Right? if you can't output a certain format, you don't know how to like access APIs, you can't call tools, etc.
And so the initial uh end to end was I was hearing feedback around you know claude 2 days claude was not very good at following instructions. So then
digging in with users, what do you mean by claude is not good at following instructions? Give me what situations
instructions? Give me what situations this was happening like what's the exact like paragraph? what did you ask? What
like paragraph? what did you ask? What
was Claude's response? Going to like that level of detail. And what I saw was something like 80% of what people meant in the early days for this failure was
Claude would not write the right JSON.
And so then, okay, let's generate maybe to start just 30 to 40 examples of when Claude was not doing this thing correctly. And then that actually is
correctly. And then that actually is your eval set. And you could have essentially uh a prompt and a response.
And if that is not working uh in the right golden answer that you might have, then that means that the the eval essentially uh is beneficial because it's identifying a painoint
consistently. And so then we added that
consistently. And so then we added that to our um repositories for evals. And
when we have uh versions of claude, we actually run that eval and just check. I
think at this point it's always 100% or like 99.9. And so it's no longer a pain
like 99.9. And so it's no longer a pain point. Uh but in the early days was
point. Uh but in the early days was taking the user feedback, figuring out actually what they mean, can we reproduce it, is it consistent, is it a
big issue, and then figuring out how to uh standardize it in a way that can be consumable for researchers. It's
basically test-driven development for PMs is is the world we're living now. Uh
where you write the test first. So is
this just a core part of the product management job now at Enthropic writing bells?
I think so. I I also think it's um something I've talked to other Piana other companies about and I think it's also more and more of the skill set more
broadly because a lot of the products that we're building is at the intersection of models with harnesses with a set of contexts for a set of
users. And so having things like eval
users. And so having things like eval isn't is a way not just for uh folks working on models but generally within product uh to to get to better user
experiences because you can't improve what you can't measure and a lot of this is very still tactile based. It's still
very judgment based and so you have to stay close to the details and also very non-deterministic which is a big part of this just like it's not going to give you the same answer every time. So you got to describe it kind of
time. So you got to describe it kind of more broadly. It's not going to be yeah
more broadly. It's not going to be yeah an exact match. So this is a really interesting change in the way product happens and will happen is eval is is a big part of this. Do you guys
still do PRDS? Is there still like a one pager describing a problem or is it play? Okay, now you're shaking your
play? Okay, now you're shaking your head. Yes,
head. Yes, we we we are we do I think um when there's a very defined problem I think things like eval might be almost a shorthand. I think there's other cases
shorthand. I think there's other cases where PRDs are really valuable. Um, PRDS
are great vehicles for getting a very large group of people aligned on a set of sources of truth about experience and setup goals. So when we do have a model,
setup goals. So when we do have a model, we actually for every model we do have a PRD less necessarily for our researchers
but more for our growing product surfaces, for our engineering teams, for our um stakeholders like uh legal and
safety and others as just a source of truth of putting together what we're aiming to achieve so that a big group of people can row in the same direction.
The other place where I do think PRDS are valuable are on the more ambiguous problems and opportunities right so we if we haven't shipped a thing like
computer use we don't necessarily have a set of like user specific pain points always and I think there's value in the product vision portions of a PRD to
explore what could even if a technology is not yet ready to work for everyone how do you get it to work well for some group.
So you can explore the value, you can actually bring something that is uh coherent to a user group. So we do have PRDS. Um I think the application is a
PRDS. Um I think the application is a little different now.
Okay, this is great. There's I just had a uh Andrew for he's the head of the codeex app at OpenAI and he's you guys are aligned. Uh PD is not dead. Still
are aligned. Uh PD is not dead. Still
very useful for specific projects and ideas. Uh great. Okay, we've closed
ideas. Uh great. Okay, we've closed closed the book on purity is still kicking. Okay, so we've been talking a
kicking. Okay, so we've been talking a bit about just what kind of skills are kind of emerging for product people. Um,
is there anything else that you find is shifted in what patterns uh are common across people that are doing well in this new AI world in terms of product managers and folks on the
product teams? Is there anything else
product teams? Is there anything else that you're like, okay, does something you got to shift or something you look for more people? I think maybe specifically
uh for folks who might be midc career or folks who have been more in a managerial like product like leadership seat. Um, one
thing that I think I feel pretty strongly about is in order to be good managers of teams and PMs working with this technology,
you have to be really hands-on yourself and have spent not just time tinkering but actually shipping with
this technology and and and again being in the details and sweating the tokens along with your PMS and your engineer.
and your teams. And so even for folks that I hire who have more tenure PM experience, the onboarding plans are exactly the same as somebody
who is like more uh early career and it's around understanding users, reading like consented user feedback, talking to customers. I think there's something
customers. I think there's something around uh being able to like understand what to do with this, what what good looks like
and having developed that in a very hands-on manner. That's important. Um
hands-on manner. That's important. Um
it's not necessarily easy for someone to uh agree or be able to see what a what a good or great AI product or AI feature
could look like if they haven't kind of experienced building themselves. Um, so I think I think there
themselves. Um, so I think I think there is a I I I do feel pretty strongly that like, you know, if you're a manager, you have to be hands-on. You have to spend a portion of your time actually shipping.
You you have to kind of walk in the shoes of your teams. uh and and that's I I always try to carve out a portion of time uh to to actually like own one to
two work streams when we have models in order to keep like keep my theory of mind, keep my sense of how the models are moving, how quickly it's improving
uh uh so I can help the team make make decisions and and make better decisions.
So, what I'm hearing here is if you're not, no matter where you are in the ladder of hierarchy at a company, if you're not building yourself, if you're not actually talking to Claude, talking to Codex, building stuff, you're not going to make it.
And you should have fun working with his technology. I think that's the other
technology. I think that's the other piece. I think the folks that would be
piece. I think the folks that would be most successful regardless of their level are people who love working with AI and and are exploring and
experimenting and carving out the time not just for the experimentation but actually hands-on shipping end to end getting the user feedback I think has to be fundamental for everyone.
I 100% know what you mean there. Just
like me sitting on my newsletter and this podcast just talking about stuff and like yeah that sounds great. Like
every time I actually build something and I tinker with all kinds of little projects, you're just like, "Okay, I see what's happening here." And you just get so much more, it's like hard to exactly describe what you're what you what you experience actually working with the
models and building stuff, but it's like a whole different world of like, "Okay, I see. Here's where the here's what
I see. Here's where the here's what they're talking about computer use.
Here's what they're talking about with this limitation, this UX situation."
Yeah.
So, yeah. So, it's just like, and you made this really interesting point that you have to have fun with it, which is not easy for a lot of people because they're pushed to use AI or they just don't know exactly what to do with it.
For people that are just like, I don't know, it's just so annoying. I just have to do this. I don't know what's so like, I hate this freaking thing. Why do I have to work with this? Things are
changing so much. I'm tired. Uh, advice
for helping people find that find that joy in this work. I think maybe I'll reemphasize something I said earlier around just that experimentation is not
an individual sport. Like some of the moments where I think I've touched practically every version of research models across 20 versions of production clause at this
point and I think part of the joy comes from seeing other people discover use cases too. And so maybe one idea here would be
too. And so maybe one idea here would be pairing with somebody who is excited and seeing what
on a use case that you care about and and working together versus um uh identifying or trying to figure out the perfect use case yourself because that
might feel like work. Working with
others feels like joy a lot of the time.
And is there more that we could do to bring that bring other people along?
That's something like a lot of times internally we have somebody who is like very curious and them sharing an idea of a new prototype actually brings a ton more people who are like oh I didn't
know this could work now with claude and so there's just some virtuous cycles here um and and ways of yeah bring continue to have joy with with this technology.
That's such a good point. I think that's also why Twitter's so useful for a lot of this is you see other people sharing what they've done and it inspires you to come up with your own little ideas and also it's just like
fun to share your own thing that you've done.
So that's a really good point just like find other people to kind of play around with and look for use cases. The thing
I've also heard a lot is just find like a problem you want to solve in your life or work and just open up cloud cloud code tell it here's what I want to do and it's incredible how far you can get just with like a vague idea of a problem
you want to solve. Yeah. I think it gets hard in that there's so many different things that you could try.
Yeah.
And so you just like narrowing in on either pairing with someone, working with somebody who who is who have a lot of joy about this technology or figuring out something that you could immediately
find value. Like either of them those
find value. Like either of them those things allow you to go deeper rather than like more high level about too many things. Um I I find it hard to keep pace
things. Um I I find it hard to keep pace with the number of prototypes or products that are out there and so my lens has been how do I go deep in one to two of them
myself. That's uh that's so interesting
myself. That's uh that's so interesting you say that because that's exactly it.
We just had this survey uh that I I ran with uh my colleague Noam uh asking my readers just how they're feeling about all the things going on in the tech right now and AI and uh one of the most
interesting takeaways we had was uh to find that happiness is exactly what you said is go deep in a couple things versus trying to just ton of little things. find a couple things to really
things. find a couple things to really solve well and then go deep and that is a source because a lot of the happiness people feel is when they finally unlocked a way for AI to actually make their lives better versus just like a
couple messed up broken half working things.
Yeah, it's it's um how do you go from this being a check the box, right? And
so like us as product people, it's then a exercise of product prioritization of your time and your energy. And and if the goal is to experiment with joy, then
how do you what are the inputs that you need for that? Um, but yeah, I I I think a lot of the um I think the secret sauce
of anthropic is the culture and the bottoms of nature of how people work and this like experimenting in public.
Um, and by doing that, it's very much about how to bring other people along.
Um, that ends up being, I think, really valuable. Yeah, I've heard this so many
valuable. Yeah, I've heard this so many times from all the labs just like no no no one's exactly sure how some of this is going to be used and a lot of it is just putting stuff out early, seeing how people use it, seeing what it's what's
possible and then using that information to build the actual product to lead in.
Yeah. Yeah.
I'm curious how kind of on this thread of finding ways AI for AI to help you in your work in life. Are there any interesting ways you've been using Claude lately in your work as a as a PM?
I think there's a lot of things with um you know fable and things like tag. So
there there I think tag is um in in the very like early days I think there's something around how you work in a different paradigm of allowing this an agent to go off and work and then bring
back uh product and experiences to you.
I think one area that it's not very recent, but one that um I bring up a lot with the team and I think we could do
more on using AI is just like how to use it to also be more uh to have better conversations with each other to be better managers.
I don't think it's necessarily uh just about raising the IQ of like experiences we build, but also I used it a lot and actually like prepping for how
to have better conversations um in the moment during like crucial conversations. So, I love that book and
conversations. So, I love that book and so I actually have a skill that helps me figure out am I having am I going in the right level of detail given the the situation at hand and actually helping
me be a better manager and better supporter for the team. Um, so for for like managers on the team, that's actually a thing that I've been sharing
more with with a uh with our managers of okay, how how do you actually use use claude to to to make you a better coach because it's hard sometimes to find the right perfect words and the models have
a lot of perfect and right words and uh I think there is something about how how it can actually augment us from like an ET perspective in addition to you. Oh
man, there's so much interesting stuff there. So just to understand what you're
there. So just to understand what you're doing there. So you built a skill.
doing there. So you built a skill.
You're just like Claude build a skill pulling in lessons from Crucial Conversations the book which it knows enough about. You don't have to even
enough about. You don't have to even give it the content. And then you use that skill to talk to Claude. Hey, I
have this very difficult conversation coming up with a colleague.
Give me some tips on how to approach it.
Yeah. And it's it's a great uh it's almost like uh coaching like individualized personalized coaching of just how to make you and and there's so
much context switching that we do all day and having like Claude help me pair and help me and maybe there are times where I end up not using suggestions
from Claude. Uh but it actually is uh
from Claude. Uh but it actually is uh ends up being very helpful for for just coming up and brainstorming. Am I
thinking about reactions in the right way? How do I actually uh go a bit
way? How do I actually uh go a bit deeper faster? Build trust faster, uh be
deeper faster? Build trust faster, uh be more direct.
Yeah, man. I have so many questions here. This so interesting. Uh one is
here. This so interesting. Uh one is just like there's concern people are going to start talking the way AI writes because they're talking AI so much and it's going to be like Diane, it's not this, but it's that. Uh I know that
you're not doing that, but that's a you know, a concern people have. Let me just ask about that, I guess. Do you fear this? There's this, you know, brain rot
this? There's this, you know, brain rot atrophy stuff people talk about it.
We're just so reliant on AI now and we stop learning and thinking and, you know, overly AI thoughts on that being so close to it and being so integrated with with AI constantly.
A lot of actually thinking process and writing process are tied together for me personally. And so I think
personally. And so I think there are ways where I use claw to augment my thinking. But what I want to make sure and maybe this is what you're
describing is Claude doesn't take over all of my thinking for me. And so I think depending on the situation, depending on how much more personal
judgment I want to have in a situation, I might uh um come up with my own POV first and then work with Claude through
that. Um and making sure that like I
that. Um and making sure that like I maintain my sense and tone throughout. I
think there are then other things like updates right we have like monthly business reviews and then in those cases it's much more I want actually want it to be standard and I want it to be much
more like it gets a cris crystallized information in the right way and maybe and I have a skill and like we're augmenting and improving our skill for that but I want to get a to a place
where like the monthly business review the writing of that is potentially asymmetrically less valuable than the thinking and so how do
I get that piece delegated to claude fully and I'm more of a reviewer and a verifier of that information. So I think it depends on like what you're using Claude for and what you're trying to
convey and like is there is there asymmetrical value in in delegating more to Claude.
What I'm also hearing the first tip is really great which was think first have a point of view and then kind of use Claude as a sparring partner almost to evolve the idea push back on the idea.
Yeah. Yeah. And I think this is where things like actually our alignment research and safety research is helpful because it what you don't want is like a
AI that just agrees with you, right?
What you want is this technology to actually augment and grow and like get to a better outcome. And so sometimes it's having Claude push back makes me better.
And so that's great. like a co-orker, I want somebody to push back when my ideas are not fully formed.
I want to hear more about that. But I've
heard that when Ben man was on the podcast, he talked about the constitution that is built into Claude and how unintuitively the work and the focus on safety and alignment as you
said and this constitution that describes how Claude should think and operate that actually you would think that would limit the abilities of Claude
and make it less fun and interesting.
It's exactly the opposite. Claude is the most interesting personality. I hear
that constantly. It's just like I much prefer talking to a like open claw famously was built on claude and then people were forced to we won't get into it were forced to switch to and
they're like this is so bad this is not who I'm used to talking to. Uh so that is I think a really interesting point I just want to make sure we spend a little time on. Why is it why is that the case
time on. Why is it why is that the case just this focus on alignment safety having this clear constitution? Why does
that make Claude better and and more interesting to talk to? Also,
in order to make Claude as like intelligent and as capable as possible, being able to have Claude actually push
back in the right points and then add it's like a yes or no and actually helps you come to a better conclusion. So,
I've used Claude to help with things like, are we making the right pricing decision on the next version of Claude?
It's a little bit meta, but using a research version of Opus, asking it to figure out how it should price and being able to come out with better outcomes is
a goal at the end of the day. And so
having AI not just be an assistant, not just be a doer, not and being delegated task, but figuring out is it doing the right thing. That's actually very
right thing. That's actually very integrated with knowing when to push back, right? That's part of knowing when you
right? That's part of knowing when you should be proactive. Proactivity is not a necessarily always doing a thing that you are scheduled to do. It is knowing
when to come up with a new idea. And so
in order for Claw to be more useful, the general approach has to be that it knows when to push back. It's a core part of the characteristics together uh of the
models.
That is so interesting. It's so
interesting that that is what a big part of like it be it being less compliant is almost what makes it better and more useful because we need that. Like I've
had so many people where they're like, "Hey, like AI told me I was right." and
like no I wish I wish to other people.
Yeah. And it comes back to our earlier point around thinking, right? How do you protect your thinking?
Um if you have a AI that can be a thinking partner, a thinking partner doesn't just agree with you. It should
add to you and you should come away at the end of the day having better ideas because you worked with Claude. That
should be the hero goal, not just making your ideas 10% better. Yeah, I love this since like it used to be think 10x. I
used to be the the way you know founders push people like what if we 10x this and I love what I keep hearing is like it's like how do we go thousandx from this idea? What is the most ambitious version
idea? What is the most ambitious version of this? I want to come back to
of this? I want to come back to something that I I was thinking about as we were talking about uh talking to Claude constantly. Um it's very clear
Claude constantly. Um it's very clear when AI has written something still.
It's funny that it's a large language model. you would think of all things it
model. you would think of all things it would be very good at writing and interestingly just no AI is very good at writing it's always very clear this was AI written
do you think we'll get to a place where we will not know this was AI I think it depends on what's the
uh goal that you're looking to achieve by knowing yeah uh what's the eval um I actually do think there's more that we could be doing on making Claude write
better. There's actually very active
better. There's actually very active efforts um on on my team and on the research side about making Claude write better. Just generally I think it should
better. Just generally I think it should be clear where an idea is ident is being
led by you or by you Lenny or me Diane.
I think it really depends on uh what's the goal of that writing. like for
something like a monthly business review, I would actually love to have that end to end be written by Claude.
Uh, and obviously and not make it feel like it was written by a human. It's such an interesting point you're making like is it actually better for us to know that it's AI versus not.
Yeah. But but it's it's um but it's also for maybe the lens is more around like verifiability or who's verifying the output. Right. Right. like who's
the output. Right. Right. like who's
signing off. Uh maybe less around who's writing, but who's verifying who's signing off. That becomes like more what
signing off. That becomes like more what matters than who's writing it.
Why Why do you think AI is not great at writing? Like my guess is it has studied
writing? Like my guess is it has studied all of the best writing in all of humanity. It's figured out here's the
humanity. It's figured out here's the best way to write. And now that we and it's just there's only so many ways to to write. And so we've just recognized,
to write. And so we've just recognized, okay, this is what AI does. It has these tropes. Is that the core of it? Is there
tropes. Is that the core of it? Is there
something else that's keeping it from being a great writer? Ironically, being
a large language model of all things, you think it'd be really great at language.
I think part of it is also uh we need to invest more in training improvements to make AI continuously strong on areas
like writing. Um I think it's also
like writing. Um I think it's also like the technology is jagged edged like like we mentioned. So sometimes when the models were good at writing but not
agentic our our thesis is how do we make the models more agentic or call the right tools. Now that that's improved a
right tools. Now that that's improved a bit then it's well now these other areas actually become more of the rough edges.
And so I think we're in one of those moments with writing where uh we need to actually just focus and prioritize on training the models to be like great at
this area and like that is an active a very active area for us that you mentioned.
Okay. I'm glad I'm glad. And also uh it's going to be interesting once AI is so good we're like I don't know who wrote that but um to your point sometimes we actually want to know that it's AI. That's really interesting. I
it's AI. That's really interesting. I
never thought of it that way. The other
interesting part of this is that there's that comedian who was joking that we're like on a plane and the Wi-Fi is down and we're just like, "What the hell? The
Wi-Fi is not working on this plane. The
sucks. How dare you?" When you're like in a in a tube in the sky flying like a bird and uh how dare you complain that the Wi-Fi doesn't work. Like your point is there's so much advancement and so
much power. Uh we can't fix it all. We
much power. Uh we can't fix it all. We
can't make it all work the best possible. And so uh basically AI writing
possible. And so uh basically AI writing has been not the priority and it feels like there's more investment happening there.
Yeah, I think like tone and character is a priority. I think it's this
a priority. I think it's this advancement of the technology is a work in progress and so we made we we see a
leap or emergence of like a jump in agentic behaviors and so that is a new normal and then these other capabilities needs to continue like improving
and I think once we improve let's say writing and like tone and character uh we probably will say like how do we have Claude be even more proactive like productivity is an
opportunity and that's human nature like we want to make ourselves better. We
want to make this technology better. Um
so yeah I I think we're applying it to to AI which is the right thing. We
should be making it better.
I want to ask you a couple questions I'd like to ask folks working at the very center of the future of that is coming.
Um one is where do you think human brains will continue to be most valuable over the years? I know anthropic's mission and and vision is we'll reach a
GI a super intelligence. So in the future maybe nowhere but before we get there where do you think human brains will continue to be most valuable as we've approached that that timeline?
We started to talk about making claude and models better at judgment um especially in the last um year or so. I
think judgment is one and is an area where it's an accumulation of so much nuance and so much experience and these systems haven't experienced as much as
humans have and so I think that hard-earned like judgment is a a a area for for product leaders and just generally um
will continue to be really critical.
There are so many things AIs can build.
which one are the things that you know an or like lab should build right a lot of that requires like human judgment persistence so proactivity these are all
traits that are beyond just general capabilities but just behaviors and characteristics of like people at that level of like how do you get to the best
solutions how do you create the the best experiences so I think those types of traits are actually the tactile uh traits that I think will
uh continue to be important. Um
I think there is also uh still a lot of like capabilities and subject matter expertise as well. I
think you know software engineering has been really transformed by AI. I think
there's areas like uh biology, life sciences. These are all things that um
sciences. These are all things that um we're just kind of at like the foot of the exponential on like maybe software engineering. We're on the exponential on
engineering. We're on the exponential on some of these area other areas. We're
not quite there yet. And so um I think you're seeing us ship things like cloud science investing in these areas because
those are areas that um I think is just bring the this technology to society and having a positive benefit for society.
So I think there's a lot more to go there.
Another question I want to ask is um as someone with kids, how do you think about what you are encouraging them to learn? or do you think you're gonna
learn? or do you think you're gonna nudge them to be successful in this wild new world that we're entering?
I actually think it's a lot of the same traits like you and I probably grew up with, which is curiosity for learning, persistence,
believing in your own inner voice, developing, and then believing in your own inner voice. Like I have a four-year-old, I have a 8-year-old. It's
on us to help uh it's on me to help them develop their indoor voice and whether that's being opinionated and taking a
stance to me right and developing that encouraging that uh I think that those types of skill sets are things that um is important in the future and like
having their own individual voice.
That is so interesting. It's so related to the answer you had when I asked about how to avoid a brain rot essentially and overrelying on AI which is just keep focused on your own point of view and
your own perspective before you overly AI and just this idea you're describing of building that in kids is is really important. Uh that is so interesting and
important. Uh that is so interesting and I love how this all this kind of connects judgment persistence in a point of view of your own.
Yeah.
Both for kids and also adults.
Yeah. Anything we think about um for your Oh man. Well, like the question I'm
Oh man. Well, like the question I'm thinking about is just when to get them on like some AI thing, you know, when I have a three-year-old, so it's pretty early for that, but you know, how do you get how do you onboard them to this
crazy thing? I had I was at an event
crazy thing? I had I was at an event recently and bunch of parents were talking about how they think about AI in their kids and one person had a really interesting approach which is uh keep them on the very early models so that
they still have to struggle a bit and not get all the answers immediately.
thought that was interesting. Like an
open source local model, not stable.
Yeah.
Yeah.
Yeah. And curiosity is something uh I I keep mentioning Ben man, but his answer actually to this question has always stuck with me, which is um curiosity and also just like he's a big fan of Monosuri, which is what I'm we're
encouraging for our kids. So, there's
something there. Maybe a last question just along kind of along these lines, something Fiona Fun actually suggested to ask you uh who's recently on the podcast. How do you stay just recharged
podcast. How do you stay just recharged and not burn out being in the center of this crazy storm of AI as a mom uh working in, you know, we're seeing the
research work at Enthropic. Uh I just like we're living through the most unprecedented time working at just like being, you know, being on the outside of Anthropic. It's crazy. I don't even know
Anthropic. It's crazy. I don't even know what it's like to be on the inside. Um
what have you learned about avoiding burnout, staying recharged, staying sane during the middle of all this? In 2024,
we shipped four models for the in the whole year or four series of models and I think we did more than that volume in just Q2 of this year.
[laughter] I think I've been really lucky with uh the team that we grown and built both the stakeholders on the research side
and within our research product management team. Um I think that one of
management team. Um I think that one of the magical parts about approaching all of this is that it's not an individual
sport. Um there's like a sense of
sport. Um there's like a sense of radical ownership and team collaboration that I think
sometimes it does feel like a high performance sport because you're in very critical decisions. there's new
critical decisions. there's new information about users about training and you have to make recommendations and judgments and decisions very quickly and
nobody can do that sustainably by themselves. Um, and so I think what's
themselves. Um, and so I think what's really helped is having a team that is incredible, who looks out for each
other, who, you know, night before a launch, even if they're not the core DRRi on that model, will stay up and help the DRRi, who uh to review the blog
post and make edits and come up with better demos and knowing to be each other's sort of extra hand. I think it's very easy if you take all of this change
on your own shoulders to feel like you're alone and to feel like you have to do everything. Uh but I think one of the like magical parts of anthropic is
this ability for us to uh figure out what are those opportunities to help each other and actually then taking the next mile of like mindmelding. We called it like
like mindmelding. We called it like entering the hive mind. There was an article about this and I think like part of that is just that allows like the team to replenish. It's not that you I I
was just on PTO in June. It's not just that you can take PTO and you come back to like 3x the amount of things to do.
It's actually that you can take PTO and know the team can figure out the right things to do and that we individually can like watch out for each other. Um so
I think that's a big part. I'm really
lucky just personally um also my partner is really supportive um this is year six of me working in AI so Amazon and then anthropic and so he sees how much I just
love the technology and what this can do and that really helps I think also um from like a personal perspective as well.
I love I love how many of these answers connect. So what I'm hearing here is
connect. So what I'm hearing here is just the having other people, working with other people, relying on other people, helping each other out when things get crazy. Uh which is a similar
answer you had for just how to how to find the joy and and and fun in this work. Just get be inspired by other
work. Just get be inspired by other people, see what they're doing, work together.
Yeah. And it's interesting when Fiona was on the podcast recently, she I was asking her just like what's changed in the world of software engineering and she pointed out it's a lot lonier now because now we're working with agents instead of other humans. Teams are
smaller, people are have all these fleets they're talking to constantly.
And so this is just a reminder of just the power of just actual other humans around you.
We're we're asked to work and make decisions on really big things because you have more scale from the technology, right? And I think
right? And I think having individuals, having other folks more who can have some level of like mind meld with what you work on, how you
approach maybe not exactly every detail, but what are the first principles? What
are the assumptions you make then helps them uh you know back up for you or uh push your decision and sharpen your thinking. Um, so I think you know we
thinking. Um, so I think you know we really try to like I really try to look for that when like building the team, growing the team, hiring like is this person going to care about their own ego
and building out a big org or are they going to care about contributing to anthropic and contributing to the like impact of the team and orienting towards
folks who are like low ego team oriented. Um, I think that's,
oriented. Um, I think that's, yeah, it it's a big part of I think the sustainability.
Yeah, just always a lot of it always just comes down back to culture and hiring and and I know I've heard a lot just the reason Anthropic is able to move so fast. I remember that moment when like something shipped every day of
the month. There's like a calendar of
the month. There's like a calendar of launches and people were talking about how is this possible and what I heard a lot is just because everyone is so aligned around the mission and the values it allows people to make
decisions really quickly before we get to our very exciting lightning round. Is
there anything else Dan that you wanted to share? Anything else you wanted to
to share? Anything else you wanted to touch on? Anything you want to maybe
touch on? Anything you want to maybe double down on of things we've talked about?
This was actually really fun because I feel like your questions actually sharpen some of my thinking around how the dots kind of connect. I'm I'm your real human claude over here.
One thing that I really uh want to like convey or um have people take away is I think one in the ways of working, but
also just two that like this is a this is a lot of like growth and change and having the joy in using this technology and like if you're feeling
like in this moment you don't have as much of that feeling of initial joy, how do you find people who do uh if this is an area that that you're
excited and like want to work on and I think developing skill sets replenishing skill sets in many ways of things like thinking from a first principles manner
about what you solve I think fundamentally you didn't ask me this but there is this question in the community of do we still need PMS when the models
are so capable when engineers are leaning in um I think the role of people who are user centric who go into the
details of understanding what users are trying to accomplish bubbling that up in an actionable manner and doing the relentless work to do that like that to
me is a core of a product person and I actually think we need more of that.
I think we are becoming very technology layered driven and actually to make that impactful it's you have to go deep you have to be curious you have to be super
hands-on and those are things that I think are also traits that have I think helped anthropic from a product development and model development perspective and as part of the culture and hopefully that's valuable for others
as well.
Amazing. What an inspiring way to end it. Oh man. Yeah. And this is I've been
it. Oh man. Yeah. And this is I've been saying this too for a long time just now that building is easy the hard part part becomes as you said what should we build and is the thing we have built correct and good and worth
leaning into and to me that's what PMs do and what PMs are good at.
Yeah. Yeah. Yeah. And it's getting into the details of the user.
Yeah. Empathy. Okay. Great. PMs are
going to make it. Okay. PRD is not dead.
[laughter] All kinds of all kinds of uh important lessons here. Uh Dan, with that we've reached our very exciting lightning round. I've got five questions
lightning round. I've got five questions for you. Are you ready?
for you. Are you ready?
Yep.
First question. What are two or three books that you find yourself recommending most to other people?
One personal one I really like how to raise an adult.
So uh I'm a mom. I think a lot about what is the things that I want to instill in in in my kids. in that book is really
helpful for describing we're not trying to raise children, we're trying to raise adults. So just the framing of what does
adults. So just the framing of what does that mean and what does it mean? What
are the characteristics that we want to hone and like harness and foster in our kids? Um the other book that I uh was
kids? Um the other book that I uh was listening to on Audible recently is Incorable by Eric Reese. So the
Incorruptible Incorruptible Yes. Yes.
Yeah. His recent podcast guest.
Um Yeah. And I I I just I think the question of how to build great companies is important. I personally just been
is important. I personally just been most fascinated with how to keep great teams and great companies going further.
And it was very interesting to just kind of see his framing and reframing of the question. Um I loved some of the
question. Um I loved some of the examples around having metrics around culture. you if you can't if you only
culture. you if you can't if you only measure revenue and then that's kind of how you're going against but if you have other better metrics that's actually the
way uh to to to sustain the the values you care about. I've been kind of trying to think about how to actually bring that to the team level of like how do we better articulate right our norms a lot of the things we talked about on the
team. So I think [snorts] that's also a
team. So I think [snorts] that's also a really good read.
There you go. Uh that'll be your next watch everyone as you're listening to this the Eric Greece episode. Yeah.
Such a good episode. Yeah.
And his book just came out.
Incorruptible.
Yes.
And I think it was like a New York Times bestseller. Like it's actually doing
bestseller. Like it's actually doing incredibly well, which I was really happy to see.
Yeah, exactly.
Next question. Favorite recent movie or TV show you really enjoyed. Most people
at Antropic don't have time to do what to watch things, but I'm curious if you have an answer. I would say um during uh some time off last month, I did get to
like binge watch Fallout on Amazon Prime. So that was actually I kind of
Prime. So that was actually I kind of like um it's kind of uh Have you heard of it?
Yeah. Yeah, it's based on the video game.
Yes, it's based on the video game. Uh I
think it's a it was really um it's witty, it's humorous, it's also like super actionoriented. So highly
super actionoriented. So highly recommend.
Okay, next question. Do you have a favorite product you recently discovered that you really love?
I really do think like claw tag is very interesting in terms of a product experience. Um, we actually have like
experience. Um, we actually have like different versions of this uh within Anthropic and I I I think it's actually been really uh really really uh powerful tool.
Yeah, it feels like I think some people are like what's the big deal? The fact
that everyone at Anthropic is like raving about it tells me something important is going on here. And I'm
trying to actually get it working within my Slack community that I have for paid newsletter subscribers. How cool would
newsletter subscribers. How cool would that be?
Yeah.
Yeah. I'm trying to figure out how it works when it's not a company when it's just a bunch of people that don't know each other and how that might work. But
we're trying it out. Okay. Uh two more questions. Your favorite life motto that
questions. Your favorite life motto that you find yourself often coming back to in work or in life. So I was actually raised by my grandparents uh for the
first 10 10 years of my life and my parents were immigrant uh college and master students in the US.
Oh wow. And um my grandfather always says, "No matter how far you go, there's always another level, [laughter]
which uh um is I think um a really good way though, like a pretty uh intense way of describing uh his his life philosophy. But I go back to that
philosophy. But I go back to that whenever there's something new or unprecedented that we experience. And I
think you know first half of this year there was definitely a lot of that like there was a lot of new things that we were learning. I was learning um so just
were learning. I was learning um so just feeling like there's always like another mountain another uh opportunity to not good enough Dan we need to go better
we need to go bigger. Uh makes me think about actually another Ben man line from his podcast episode that this is the most normal it's ever going to be. It's
only going to get weirder and crazier.
Yeah. Yeah. No, we're good.
Okay, final question. Uh, I was poking around at your LinkedIn. You were a high yield bond trader, JP Morgan Chase early in your career. Uh, you had like uh you
have this like redacted uh hundred million dollar trading portfolio of some kind. Uh what did you learn from that
kind. Uh what did you learn from that time in your life that has stuck with you and or is there a crazy story from that period? It was four years of your
that period? It was four years of your life. I think I learned actually a lot
life. I think I learned actually a lot that I uh apply here uh at at Anthropic and other uh jobs thereafter. Um so when
I was at JP Morgan um the trading floor you could kind of envision like sort of Waffle Wall Street that's very different. Uh most traders I think are
different. Uh most traders I think are in front of a terminal. They're much
more doing analyses uh on their computers. Um, but it's still very, I
computers. Um, but it's still very, I would say, like male-dominated.
And so, uh, I was the only woman. I was
the only, uh, um, person with like my background, uh, on the trading desk. And
I learned that the it was a very good environment to kind of building one my sense of authentic self
and two uh that even if I was the most junior person, even if I may look different, uh that the best ideas
and having conviction in the best ideas uh irregardless of all of those other factors like is the most important thing. And so I think just bringing that
thing. And so I think just bringing that sense of um how I show up more at work.
Um I'm pretty vulnerable and authentic with my team. Uh I try to really make sure that regardless of people's levels or tenures, if they have a great idea,
how to help them pursue that and to do also the same. Um so to like put the idea out there to actually um have conviction in it to do the follow
through to do the like nitty-gritty work to make it happen. Um so those were all things that I learned from trading. Um
and yeah I think applies to any any job in many ways.
That is beautiful. Where can people find you online if they want to follow you and how can listeners be useful to you?
I don't have a large presence on like uh social. Uh I think the best way to uh
social. Uh I think the best way to uh find my work uh my team's work is really uh the anthropic blog and when we're
publishing new models, new product experiences I think in terms of uh useful uh for me I think the best thing
number one is your feedback like we actually if if you thumbs up or thumbs down on any of our product surfaces if you contact your salesperson with feedback back about the model, it will
make its way to me. Uh we actually with every like research model, I actually get pretty close into understanding favorability and feedback. Um so giving us that feedback, pushing Claude,
telling us where it's falling down, um those help us make Claude better. Uh the
other the other thing is like if you have folks in your network who seem like this type of profile of person that I just talked about I'm hiring the team is growing.
We really will love just people who love this technology who are deeply curious first principles thinkers who are fearless in questioning assumptions um
and who have like a tinkering hackery spirit.
Wow dream job. So basically open open PM roles add anthropic on the research team.
Yes.
And they apply I assume on the website the careers page.
Yes.
Holy moly. All right. Here we go. Enjoy
the flood of resumes you're about to receive.
Thank you. [laughter]
Uh Dan, thank you so much for being here.
Thank you so much for having me. Thank
you for um really helpful, thoughtprovoking questions um helping me even connect the dots on how how we work, how how this whole technology is coming together and being product people
in it.
I really appreciate that. But thank you Dan for real. Okay. Well, bye everyone.
Thank you so much for listening. If you
found this valuable, you can subscribe to the show on Apple Podcast, Spotify, or your favorite podcast app. Also,
please consider giving us a rating or leaving a review as that really helps other listeners find the podcast. You
can find all past episodes or learn more about the show at lennispodcast.com.
See you in the next episode.
Loading video analysis...