9.4s
started, Jeff? Sure. Sounds great. All right. Jeff, welcome. And again, thank you so much for being here. Especially I just got a cold and thank you for being here. Yeah, I'm afraid I've lost my voice. I don't normally sound quite like this, but we'll we'll do what we can. So, um, you built map reduce, big table,
10.2s
Sure. Sounds great.
11.6s
All right. Jeff, welcome. And again,
13.7s
thank you so much for being here.
14.9s
Especially I just got a cold and thank
17.3s
you for being here.
17.8s
Yeah, I'm afraid I've lost my voice. I
19.4s
don't normally sound quite like this,
20.6s
but we'll we'll do what we can.
24.1s
So, um, you built map reduce, big table,
27.2s
tensorflow, the TPU, Gemini. We could spend a whole hour on all the things you've done, but what I love is that you're still making bold predictions in public. Last year, yes, last year in May 2025 at AI Ascent, you said that AI is at the level of a junior engineer.
30.0s
spend a whole hour on all the things
32.1s
you've done, but what I love is that
34.9s
you're still making bold predictions in
36.9s
public. Last year, yes, last year in May
42.6s
2025 at AI Ascent, you said that AI is
46.1s
at the level of a junior engineer.
50.0s
That was about a year ago. It's been How close are we to that prediction? Yeah, I mean I feel like uh the models have been getting a lot better at sort of agent-based longer running coding tasks and it seems pretty clear that they are now actually pretty capable and depending on exactly your definition of
55.4s
close are we to that prediction? Yeah, I
58.1s
mean I feel like uh the models have been
60.2s
getting a lot better at sort of
61.7s
agent-based longer running coding tasks
64.5s
and it seems pretty clear that they are
66.7s
now actually pretty capable and
68.7s
depending on exactly your definition of
70.3s
of junior engineer it seems pretty spot-on I would say. What did you underestimate from that What did you underestimate from that prediction? Um I mean I think the the ability to do more and more complex tasks has been growing faster than I thought. Um and I also think uh outside of coding these these agent-based
72.6s
spot-on I would say.
75.2s
What did you underestimate from that
78.0s
What did you underestimate from that prediction?
79.8s
Um I mean I think the
84.2s
the ability to do more and more complex
86.6s
tasks has been growing faster than I
88.7s
thought. Um and I also think uh outside
92.5s
of coding these these agent-based
94.3s
systems are are really starting to shine in other domains and I I think uh you know uh that's that's going to be an important trend in the future. So snorts give us another bold prediction. What do you think is going to be the 2027 edition? Uh I think you will see a lot more
96.4s
in other domains and I I think uh you
99.5s
know uh that's that's going to be an
101.8s
important trend in the future.
104.2s
So snorts give us another bold
107.0s
prediction. What do you think is going
108.6s
to be the 2027 edition?
111.4s
Uh I think you will see a lot more
114.3s
automation of uh ML systems themselves. um basically getting ML systems to improve their capabilities by running lots of experiments, breaking things down into subpros, you know, running those subpros in a tight automatic experimentation loop, putting the results together and being able to then uh you know, get some improved system uh
118.9s
um basically getting ML systems to
120.7s
improve their capabilities by running
123.4s
lots of experiments, breaking things
125.0s
down into subpros, you know, running
127.2s
those subpros in a tight automatic
129.7s
experimentation loop, putting the
131.5s
results together and being able to then
134.1s
uh you know, get some improved system uh
137.4s
out from that uh sort of fully automated problem decomposition and and automated experimentation that I think that's going to be really exciting. M I think that also applies not just to ML but also to other fields of science and engineering. Um basically anything where you can have a measurable objective uh I
140.5s
problem decomposition and and automated
143.0s
experimentation that I think that's
144.4s
going to be really exciting. M
145.8s
I think that also applies not just to ML
148.6s
but also to other fields of science and
151.0s
engineering. Um basically anything where
152.7s
you can have a measurable objective uh I
155.5s
I think you can uh actually make a lot of progress these days. Now let's go back to a little bit in history. Back in way back in 2001 Google search used to run on hard drives. Yep. And you and Sanjay did the math and realized that at some point the whole
157.4s
of progress these days.
160.0s
Now let's go back to a little bit in
162.0s
history. Back in way back in 2001 Google
167.0s
search used to run on hard drives.
169.4s
Yep. And you and Sanjay did the math and
173.3s
realized that at some point the whole
175.1s
search index would finally fit in all of the RAM of all the computers you had the RAM of all the computers you had running and you made that radical realization and you basically in few days with Sanjay shipped in production a whole new search version that worked in RAM rather than hard drive and that was the thing
178.6s
the RAM of all the computers you had
182.0s
the RAM of all the computers you had running
183.7s
and you made that radical realization
187.3s
and you basically in few days with
189.5s
Sanjay shipped in production a whole new
192.2s
search version that worked in RAM rather
194.9s
than hard drive and that was the thing
197.3s
that got Google to be so fast. Google that got Google to be so fast. Google searches. So history tends to remix. What is the it fits the memory moment right now in 2026 that everyone in this room is still should be thinking about and designing? Yeah. Yeah, I mean it's a little
199.6s
that got Google to be so fast. Google searches.
201.4s
So history tends to remix.
205.8s
What is the it fits the memory moment
208.6s
right now in 2026 that everyone in this
211.9s
room is still
214.8s
should be thinking about and designing?
216.9s
Yeah. Yeah, I mean it's a little
218.6s
different, but I think uh you're going to see more and more uh uh high performance and um low energy uh inference hardware systems because I think everyone is now realizing that inference is the key to making you know these agent-based systems be available to more and more people and that latency is really important and that
221.0s
to see more and more
223.4s
uh uh high performance and um low energy
228.4s
uh inference hardware systems because I
231.4s
think everyone is now realizing that
233.3s
inference is the key to making you know
235.8s
these agent-based systems be available
238.2s
to more and more people and that latency
241.4s
is really important and that
243.1s
specialization of the hardware is a really key way you can make uh things that are more energy efficient and lower latency than more general purpose uh computational devices like say GPUs or computational devices like say GPUs or TPUs because I think all everyone here is used to waiting for responses on on models. clears throat
245.7s
really key way you can make uh things
247.9s
that are more energy efficient and lower
250.2s
latency than more general purpose uh
253.0s
computational devices like say GPUs or
255.0s
computational devices like say GPUs or TPUs
256.3s
because I think all everyone here is
258.1s
used to waiting for responses on on
261.3s
models. clears throat
262.2s
models. clears throat So waiting is no fun master speed. So you're saying what if we don't have to wait anymore? Yeah. I mean, I think we'll imagine what you could do with something where the latency is, you know, 50x better. Interesting thought. Now, what's one assumption that perhaps 6,000 people in this room hold that's
262.7s
waiting is no fun
265.3s
master speed.
267.7s
So you're saying what if we don't have
269.4s
to wait anymore?
271.0s
Yeah. I mean, I think we'll imagine what
272.9s
you could do with something where the
274.1s
latency is, you know, 50x better.
278.3s
Interesting thought.
280.3s
Now, what's one assumption that perhaps
283.4s
6,000 people in this room hold that's
286.1s
already false about AI? Yeah. Uh, that's that's a good question. I mean I think um probably one thing is people don't quite realize how possible it is to have you know agent-based systems that can run not just for an hour or two hours on a problem you care about but for some problem domains and with highly capable
290.1s
Yeah. Uh, that's that's a good question.
291.8s
I mean I think um
295.5s
probably one thing is people don't quite
297.6s
realize how possible it is to have you
300.9s
know agent-based systems that can run
302.9s
not just for an hour or two hours on a
305.8s
problem you care about but for some
308.0s
problem domains and with highly capable
309.8s
models underlying them you can get them to run for days or weeks and do really really complicated tasks and I think that's you know starting some people are starting to see inklings of this but I don't think everyone has really internalized this and that's going to be really uh a pretty big deal.
312.1s
to run for days or weeks and do really
315.1s
really complicated tasks and I think
317.3s
that's you know starting some people are
320.4s
starting to see inklings of this but I
322.5s
don't think everyone has really
323.7s
internalized this and that's going to be
325.6s
really uh a pretty big deal.
328.0s
What's a particular task that you have run that has run for weeks? What what was it? What did the tell what did you tell the agents to solve? Yeah, I mean I think uh you can tell agents to uh go off and implement um you know completely new versions of software in different programming
329.9s
run that has run for weeks? What what
332.5s
was it? What did the tell what did you
334.1s
tell the agents to solve? Yeah, I mean I
336.1s
think uh you can tell agents to uh go
341.0s
off and implement
343.1s
um you know completely new versions of
346.2s
software in different programming
347.9s
languages that might be you know have better safety properties or better performance properties uh that and then then they can go off and and actually do that in a you know pretty serious way. That's pretty cool. clears throat Now, one thing that you've been very well known for is you're really good at napkin math.
349.7s
better safety properties or better
351.0s
performance properties uh that and then
353.5s
then they can go off and and actually do
355.1s
that in a you know pretty serious way.
358.9s
That's pretty cool.
360.9s
clears throat Now, one thing that
362.3s
you've been very well known for is
364.5s
you're really good at napkin math.
367.7s
Sounds funny. So, one of the stories about you is that back in uh 2013 when speech recognition started to work at Google, you did the nap napkin math where if every Google user used their phone and talked to it and used the speech recognition system for three minute just three minutes a day, you
370.2s
about you is that back in uh 2013 when
374.3s
speech recognition started to work at
376.6s
Google, you did the nap napkin math
379.8s
where if every Google user used their
383.2s
phone and talked to it and used the
384.8s
speech recognition system for three
387.1s
minute just three minutes a day, you
389.4s
found that the system requires a Google server. you would have to double the fleet which would be really really expensive just to do speech translation. expensive just to do speech translation. Yeah. And instead you basically built a custom ship and that was the origin story of the TPU. Yeah. Yeah. I mean I I sort of had done
392.1s
server. you would have to double the
393.7s
fleet which would be really really
396.5s
expensive just to do speech translation.
399.3s
expensive just to do speech translation. Yeah.
400.4s
And instead you basically built a custom
404.1s
ship and that was the origin story of
407.4s
the TPU.
408.6s
Yeah. Yeah. I mean I I sort of had done
411.8s
you know we were starting to see really good uh quality results on this the sort of deep learning based speech systems uh speech models we were training. um but they were computationally expensive compared to the old speech system but they haved the error rate. So that was like the equivalent of 20 years of
413.4s
good uh quality results on this the sort
417.1s
of deep learning based speech systems uh
419.6s
speech models we were training. um but
422.1s
they were computationally expensive
423.6s
compared to the old speech system but
425.4s
they haved the error rate. So that was
427.4s
like the equivalent of 20 years of
429.2s
advances in speech recognition in just a few months of like fiddling with the model and getting scaling it up a bit and getting better data. And so we started to get worried that if speech worked a lot better, people would use it more. And so that that back of the envelope calculation was really about
431.8s
few months of like fiddling with the
433.8s
model and getting scaling it up a bit
435.8s
and getting better data. And so we
439.0s
started to get worried that if speech
440.9s
worked a lot better, people would use it
442.9s
more. And so that that back of the
445.0s
envelope calculation was really about
446.6s
that like well what if people start start to use speech recognition more to dictate emails or to talk to their phone or whatever. Um and yeah it turned out that um we realized that we needed some better solution than running on CPUs at the time. And so we came up with TPUs
449.1s
start to use speech recognition more to
451.3s
dictate emails or to talk to their phone
453.8s
or whatever. Um and yeah it turned out
457.2s
that um we realized that we needed some
461.9s
better solution than running on CPUs at
464.8s
the time. And so we came up with TPUs
467.2s
which are sort of very specialized for essentially low precision dense linear algebra which is at the heart of nearly all of the modern machine learning algorithms we we use today. And um if you build a specialized chip for low precision dense linear algebra and can't do anything else that turns out to be really useful for machine learning
470.1s
essentially low precision dense linear
472.5s
algebra which is at the heart of nearly
474.8s
all of the modern machine learning
476.1s
algorithms we we use today. And um if
480.1s
you build a specialized chip for low
483.0s
precision dense linear algebra and can't
484.5s
do anything else that turns out to be
486.6s
really useful for machine learning
487.8s
inference uh even though it can't run Chrome or Word or whatever. Uh, and so that system produced a chip a couple years later that was uh 30 to 80 times more energy efficient than CPUs and GPUs of the day and also much much lower latency like 20 to 30x lower latency which is incredible what the foundation
489.9s
Chrome or Word or whatever. Uh, and so
492.5s
that system produced a chip a couple
495.0s
years later that was uh 30 to 80 times
499.0s
more energy efficient than CPUs and GPUs
502.8s
of the day and also much much lower
505.1s
latency like 20 to 30x lower latency
508.2s
which is incredible what the foundation
510.6s
that TPU has become today. No way you would have predicted that TPU would be so foundational now with transformer architecture which was invented way later before you actually invented the TPU. Yeah, I mean that's sort of why we built a general purpose linear algebra system, which is what a TPU is really.
512.9s
would have predicted that TPU would be
514.6s
so foundational now with transformer
517.0s
architecture which was invented way
518.6s
later before you actually invented the
520.8s
TPU. Yeah, I mean that's sort of why we
523.4s
built a general purpose linear algebra
526.4s
system, which is what a TPU is really.
529.0s
Um, because we knew ML algorithms were still evolving and you didn't want to over specialize, but you wanted to specialize enough that you got the dramatic performance benefits of we could have very big multiplier units. uh we could have you know high-speed memory we could have high-speed interconnect or later TPUs that like brought many many
531.9s
still evolving and you didn't want to
533.8s
over specialize, but you wanted to
535.8s
specialize enough that you got the
537.8s
dramatic performance benefits of we
540.3s
could have very big multiplier units. uh
543.0s
we could have you know high-speed memory
544.6s
we could have high-speed interconnect or
546.7s
later TPUs that like brought many many
549.5s
chips to bear on the same problem efficiently and u you know we've continued to scale those up and and improve their performance uh for over many many generations now incredible napkin math so what's good napkins are good so actually what's a good napkin math that everyone here who wants to be a future founder
551.5s
efficiently and u you know we've
554.1s
continued to scale those up and and
555.9s
improve their performance uh for over
558.2s
many many generations now
560.4s
incredible napkin math so
562.8s
what's good
563.8s
napkins are good
565.4s
so actually what's a good napkin math
568.6s
that everyone here who wants to be a
571.7s
future founder
572.9s
should run tonight to potentially build something as consequential as the TPU. Yeah, I mean, uh, it's always hard to say. Um, clears throat I think, uh, think about what problems you see in whatever it is you're thinking about, what what bottlenecks you see, and are there very different ways of thinking of
575.6s
something as consequential as the TPU.
579.1s
Yeah, I mean, uh, it's always hard to
582.0s
say. Um, clears throat I think, uh,
586.8s
think about what problems you see in
589.7s
whatever it is you're thinking about,
591.1s
what what bottlenecks you see, and are
594.6s
there very different ways of thinking of
596.4s
the solutions to some of those problems that would get you, you know, an order of magnitude or two orders of magnitude better uh, performance or capability or whatever it is. Um, you know, because sometimes if you just squint at a problem and you think about not necessarily being anchored on exactly how that problem is solved today, but
598.5s
that would get you, you know, an order
601.0s
of magnitude or two orders of magnitude
602.7s
better uh, performance or capability or
605.6s
whatever it is. Um, you know, because
608.2s
sometimes if you just squint at a
610.8s
problem and you think about not
613.2s
necessarily being anchored on exactly
614.6s
how that problem is solved today, but
616.8s
how you would solve it from first principles, you can come up with really good ideas that are, you know, maybe not what other people are thinking about. That's a good tip. clears throat No. Um, for everyone here who doesn't know, years ago, Jeff wrote a very famous list called the latency numbers. every engineer should know
618.8s
principles, you can come up with really
620.9s
good ideas that are, you know, maybe not
623.1s
what other people are thinking about.
625.1s
That's a good tip. clears throat
627.5s
No. Um, for everyone here who doesn't
629.8s
know, years ago, Jeff wrote a very
633.0s
famous list called the latency numbers.
636.0s
every engineer should know
638.9s
and these are numbers around for example how long a cache miss takes uh disk seek a network package traveling let's say from California to Netherlands um lots of numbers like this about distributed systems and systems engineering clears throat and it's been sort of taped and become the bible for a lot of distributed systems engineers okay yeah
641.4s
how long a cache miss takes uh disk seek
646.2s
a network package traveling let's say
648.3s
from California to Netherlands
651.8s
um lots of numbers like this about
653.6s
distributed systems and systems
655.1s
engineering clears throat
656.2s
and it's been sort of taped and become
658.1s
the bible for a lot of distributed
659.6s
systems engineers
660.6s
okay yeah
661.8s
now fast forward that list is up for an update give us the AI edition for now update give us the AI edition for now 2026. Yeah, I mean I think if you looked at what is important in AI systems these days, you would want to know things like the bandwidth between you know your main
664.9s
update give us the AI edition for now
668.3s
update give us the AI edition for now 2026.
669.5s
Yeah, I mean I think if you looked at
671.0s
what is important in AI systems these
673.4s
days, you would want to know things like
677.2s
the bandwidth between you know your main
681.0s
memory system on your accelerator to the onchip memory to the um you know the multiplier unit or whatever. You want to know how much energy does it take to do a single multiplier operation. um you know uh what is the interconnect bandwidth between chips and how much does that uh how how many chips can you
683.5s
onchip memory to the um you know the
687.0s
multiplier unit or whatever. You want to
688.7s
know how much energy does it take to do
691.4s
a single multiplier operation.
694.2s
um you know uh what is the interconnect
696.7s
bandwidth between chips and how much
699.7s
does that uh how how many chips can you
702.2s
connect with that bandwidth and then if you go beyond that domain like what is the fall off in in uh network bandwidth when you need to talk to 10,000 strips instead of instead of uh 500 or something I think these are all really important numbers to to learn and and really affect how you think
703.8s
you go beyond that domain like what is
707.4s
the fall off in in uh network bandwidth
710.3s
when you need to talk to 10,000 strips
712.9s
instead of instead of uh 500 or
715.3s
something I think these are all really
717.0s
important numbers to to learn
719.8s
and and really affect how you think
722.2s
about solving particular kinds the about solving particular kinds the problems. H clears throat and one interesting thing that I've heard you talk about is that nowadays the unit that you measure everything is energy. everything is energy. Yeah. You pointed out that doing a calculation or math costs about one pico.
723.7s
about solving particular kinds the problems.
724.5s
H clears throat and one interesting
727.0s
thing that I've heard you talk about is
728.7s
that nowadays the unit that you measure
732.4s
everything is energy.
735.4s
everything is energy. Yeah.
735.8s
You pointed out that doing a calculation
738.0s
or math costs about one pico.
741.5s
Uh but moving the data and doing data IO costs thousand times that. Yeah. Just bringing it in from HPM on an accelerator into the processor so it can actually compute on it. Yep. That gap kind of quietly decides what products are possible and how these algorithms in AI are built. So what are
744.2s
costs thousand times that.
746.3s
Yeah. Just bringing it in from HPM on an
748.9s
accelerator into the processor so it can
751.4s
actually compute on it. Yep.
753.4s
That gap kind of quietly decides what
756.7s
products are possible and how these
758.6s
algorithms in AI are built. So what are
762.6s
the kinds of problems that founders keep calling model problems but are in fact actually energy or data IO problems? Yeah, I mean I think the the example you raised of a thousandx difference in bringing mo moving data versus actually computing on it uh in in terms of energy is is a pretty significant one and it
764.8s
calling model problems but are in fact
767.4s
actually energy or data IO problems?
771.2s
Yeah, I mean I think the the example you
774.0s
raised of a thousandx difference in
776.2s
bringing mo moving data versus actually
779.1s
computing on it uh in in terms of energy
782.5s
is is a pretty significant one and it
784.5s
shapes a lot of aspects of what we do in machine learning. Um because if you didn't have that thousandx difference then you know you wouldn't have to do batching but you have to do batching of you know many examples or maybe many tokens at once in order to amortize that data movement clears throat so that
786.3s
machine learning. Um because if you
790.0s
didn't have that thousandx difference
791.8s
then you know you wouldn't have to do
794.0s
batching but you have to do batching of
796.6s
you know many examples or maybe many
798.4s
tokens at once in order to amortize that
801.1s
data movement clears throat so that
802.9s
you can uh you know not pay a thousandx slowdown but pay a 1000x divided by batch size uh energy cost. Um and you know for for really low latency batching is not really very good. Um so I think these clears throat kinds of things and the energy uh behind various decisions in the computer hardware we
806.2s
slowdown but pay a 1000x divided by
808.3s
batch size uh energy cost. Um and you
812.7s
know for for really low latency batching
816.4s
is not really very good. Um so I think
819.8s
these clears throat kinds of things
820.7s
and the energy uh behind various
823.4s
decisions in the computer hardware we
825.2s
use really affects a lot of decisions we make in building higher level systems. A very concrete example is just how training models is done. There's this whole whole concept of batching the the data sets and running epochs. That's basically people perhaps may confuse that as a model problem, but it's really a systems data IO problem, right?
828.0s
make in building higher level systems.
831.8s
A very concrete example is just how
833.8s
training models is done. There's this
835.8s
whole whole concept of batching the the
839.6s
data sets and running epochs. That's
841.8s
basically people perhaps may confuse
843.8s
that as a model problem, but it's really
845.8s
a systems data IO problem, right?
848.3s
Yeah. Yeah. I mean, you have to assemble batches to get better efficiency in your hardware. You know, ideally you might do batch size one training, but uh you know, it's um not as not as good in terms of efficiency. So people use re pretty large batches these days. Do you think uh it's possible for uh I
850.6s
batches to get better efficiency in your
853.0s
hardware. You know, ideally you might do
855.3s
batch size one training, but uh you
857.9s
know, it's um not as not as good in
861.4s
terms of efficiency. So people use re
864.2s
pretty large batches these days.
866.2s
Do you think uh it's possible for uh I
868.2s
know you're you're well known for uh taking off uh on a long week or weekend and coming up with this brilliant solution. Is there such things of Jeff going and working on it for a couple weeks and clears throat getting batch size equals one training getting batch size equals one training done. Yeah, I've been thinking more about
870.2s
taking off uh on a long week or weekend
873.0s
and coming up with this brilliant
874.3s
solution. Is there such things of Jeff
876.5s
going and working on it for a couple
878.1s
weeks and clears throat
879.0s
getting batch size equals one training
881.2s
getting batch size equals one training done.
883.2s
Yeah, I've been thinking more about
884.6s
inference actually. So I think inference is a pretty interesting problem because you do want very low latency. You know training you don't necessarily need incredibly low latency. Um and I think there's a lot of room for specializing hardware more for inference than we are hardware more for inference than we are today. What are some of those interesting
886.3s
is a pretty interesting problem because
889.1s
you do want very low latency. You know
892.2s
training you don't necessarily need
893.8s
incredibly low latency. Um and I think
896.4s
there's a lot of room for specializing
898.9s
hardware more for inference than we are
900.8s
hardware more for inference than we are today.
902.5s
What are some of those interesting
903.6s
things that are on inference that you're really thinking a lot about? Um I mean just trying to minimize data movement. uh trying to think about movement. uh trying to think about incredibly uh low precision operations uh and maybe not supporting lots and lots of different kinds of precisions. Uh if you feel like you have a a good
905.7s
really thinking a lot about?
907.2s
Um I mean just trying to minimize data
909.9s
movement. uh trying to think about
912.4s
movement. uh trying to think about incredibly
914.1s
uh low precision operations
917.4s
uh and maybe not supporting lots and
919.1s
lots of different kinds of precisions.
921.1s
Uh if you feel like you have a a good
925.0s
answer for what kinds of precision you need, maybe just build that into the hardware and and um not much else. which I think it brings down to a core clears throat analogy I heard from famous computer scientists that really the whole process of u AI is a big compression problem because in order to
927.0s
need, maybe just build that into the
929.1s
hardware and and um not much else.
933.4s
which I think it brings down to a core
936.0s
clears throat analogy I heard from
938.9s
famous computer scientists that really
941.0s
the whole process of u AI is a big
944.7s
compression problem because in order to
948.0s
have the data to be f fully lossy and compress it and then restore it you basically need to understand it. Yeah, I mean if you truly understand the data, you should be able to compress it really you should be able to compress it really well and now transformer architecture is basically one of the ways that has
950.0s
compress it and then restore it you
952.2s
basically need to understand it. Yeah, I
954.7s
mean if you truly understand the data,
957.0s
you should be able to compress it really
958.4s
you should be able to compress it really well
959.6s
and now transformer architecture is
962.1s
basically one of the ways that has
964.9s
turned out to work really well. Yeah. Yeah, I would say working pretty well so far. Good work by my colleagues. laughter Yes. Now let's zoom out a bit. Um AI progress used to mean just better models. You could had more data trainer models with bigger parameters. But increasingly in the last years or so,
967.0s
Yeah. Yeah, I would say
968.8s
working pretty well so far.
969.8s
Good work by my colleagues. laughter
972.0s
Yes. Now let's zoom out a bit. Um AI
976.1s
progress used to mean just better
978.2s
models. You could had more data trainer
980.8s
models with bigger parameters. But
983.2s
increasingly in the last years or so,
986.2s
it's everything around the model. Not just the model size and number of parameters or more data. It's everything around things like retrieval tools, memory, agent tools, and it might kind of get consolidated into what people call uh context engineering, right? call uh context engineering, right? clears throat Yeah. I mean I think uh the model is
987.8s
just the model size and number of
989.4s
parameters or more data. It's everything
991.4s
around things like retrieval tools,
994.0s
memory, agent tools, and it might kind
997.3s
of get consolidated into what people
999.2s
call uh context engineering, right?
1002.2s
call uh context engineering, right? clears throat
1002.6s
Yeah. I mean I think uh the model is
1005.6s
really only one piece of what you're trying to do which is build an overall system that can solve really interesting problems and that involves you know a model that knows how to use various tools. It maybe knows how to retrieve relevant information, maybe has a, you know, a history of other uh information
1008.1s
trying to do which is build an overall
1009.7s
system that can solve really interesting
1012.0s
problems and that involves you know a
1015.6s
model that knows how to use various
1017.2s
tools. It maybe knows how to retrieve
1019.5s
relevant information, maybe has a, you
1022.1s
know, a history of other uh information
1025.1s
that it has retrieved for past problems and it can put information into the context of the of the model. And the nice thing about that is that information is really clear to the model, unlike the training data the model was trained on where it's all kind of like trillions of tokens stirred
1028.0s
and it can put information into the
1031.8s
context of the of the model. And the
1034.0s
nice thing about that is that
1035.4s
information is really clear to the
1037.3s
model, unlike the training data the
1039.4s
model was trained on where it's all kind
1041.1s
of like trillions of tokens stirred
1044.0s
together into a soup of of hundreds of billions or trillions of parameters, but it's all less clear than the actual context uh that the model sees directly for this particular problem or uses use case. And then I think being able to understand what tools are available, which ones are going to help me solve
1046.5s
billions or trillions of parameters, but
1048.6s
it's all less clear than the actual
1052.3s
context uh that the model sees directly
1054.7s
for this particular problem or uses use
1056.9s
case. And then I think being able to
1059.9s
understand what tools are available,
1063.0s
which ones are going to help me solve
1064.6s
the help the model solve this next you know phase of the problem, how to decompose a problem into a sequence of of tool calls. Maybe trying multiple approaches to solve the problem and seeing which ones work and being able to evaluate that. you know this is the whole um you know orchestration of complex agent and multi-agent systems
1067.3s
know phase of the problem, how to
1068.6s
decompose a problem into a sequence of
1070.7s
of tool calls. Maybe trying multiple
1072.8s
approaches to solve the problem and
1075.1s
seeing which ones work and being able to
1076.6s
evaluate that. you know this is the
1078.7s
whole um you know orchestration of
1082.8s
complex agent and multi-agent systems
1084.9s
that I think is going to be more and more important and uh super exciting times I would say and I think the fun thing about this particular problem domain set is actually something that everyone in this room can actually do because before to train a model you needed incredible amount of resources incredible amount of
1086.8s
more important and uh super exciting
1089.0s
times I would say
1090.3s
and I think the fun thing about this
1091.6s
particular problem domain set is
1093.6s
actually something that everyone in this
1095.3s
room can actually do because before to
1098.1s
train a model you needed incredible
1099.6s
amount of resources incredible amount of
1101.6s
access of to GPUs and data but for context engineering everyone here could do you have you just need the API to something like Gemini and then work on your own setup for your own retrieval your own tool calls and etc etc. So how does what are some tips for everyone here? How does everyone get better at
1104.7s
context engineering everyone here could
1107.1s
do you have you just need the API to
1109.4s
something like Gemini and then work on
1111.8s
your own setup for your own retrieval
1114.3s
your own tool calls and etc etc. So how
1118.8s
does what are some tips for everyone
1120.2s
here? How does everyone get better at
1122.9s
and become exceptional at context engineering? Yeah, I mean I think uh engineering? Yeah, I mean I think uh clears throat a really good way to do it is to use these models and and sort of harnesses and tools and so on to try to solve problems and then some sometimes you can actually see where the models are
1125.8s
engineering? Yeah, I mean I think uh
1129.2s
engineering? Yeah, I mean I think uh clears throat
1130.1s
a really good way to do it is to use
1133.0s
these models and and sort of harnesses
1135.4s
and tools and so on to try to solve
1137.3s
problems and then some sometimes you can
1140.0s
actually see where the models are
1141.3s
failing. And often you can actually make the model work better and succeed at that kind of problem by not just adjusting the model parameters which is hard to do from the outside but from you know creating better guidelines for the model you know writing skills for the model to know how to use different tools
1144.8s
the model work better and succeed at
1147.4s
that kind of problem by not just
1150.2s
adjusting the model parameters which is
1151.8s
hard to do from the outside but from you
1154.6s
know creating better guidelines for the
1156.7s
model you know writing skills for the
1159.1s
model to know how to use different tools
1161.4s
that would be incredibly useful for solving this particular class of problem. And I think as you do that, you end up on this kind of improving self-improving of the setup that you're trying to use to to solve things. Uh, and you know that that's a really good way to get better at understanding what
1162.8s
solving this particular class of
1164.7s
problem. And I think as you do that, you
1168.2s
end up on this kind of improving
1170.9s
self-improving of the setup that you're
1173.2s
trying to use to to solve things. Uh,
1176.1s
and you know that that's a really good
1177.9s
way to get better at understanding what
1181.3s
what additional information the model would want in order to become more would want in order to become more capable. Can you give an example of uh some context engineering you personally have done? um I don't know skills you wrote tools that really made a huge different in your in your workflow. Yeah, I mean I
1183.1s
would want in order to become more
1185.0s
would want in order to become more capable.
1186.5s
Can you give an example of uh some
1188.7s
context engineering you personally have
1190.3s
done? um I don't know skills you wrote
1192.5s
tools that really made a huge different
1194.6s
in your in your workflow. Yeah, I mean I
1197.2s
guess uh Sanjay and I were working a few weeks ago and we you know we often do some amount of like uh performance improvement for very low-level libraries and we have a microbenchmark library we've written at Google where you can write microbenchmarks of how how long different kinds of operations take or
1199.0s
weeks ago and we you know we often do
1203.2s
some amount of like uh performance
1205.3s
improvement for very low-level libraries
1208.0s
and we have a microbenchmark library
1210.8s
we've written at Google where you can
1212.6s
write microbenchmarks of how how long
1215.2s
different kinds of operations take or
1217.0s
how long does it take to populate this data structure whatever and sometimes those data structures are used on millions of processes across Google. So, it's actually pretty important to make sure they're high performance. And so, you can write microbenchmarks. Um, but then without an agent-based system, what you usually do is you measure what the
1218.7s
data structure whatever and sometimes
1220.7s
those data structures are used on
1222.8s
millions of processes across Google. So,
1225.5s
it's actually pretty important to make
1226.6s
sure they're high performance. And so,
1228.7s
you can write microbenchmarks. Um, but
1231.2s
then without an agent-based system, what
1233.8s
you usually do is you measure what the
1235.9s
current performance is on some benchmarks you care about. You make some modifications to improve the performance you hope. Then you rerun the the benchmarks, see where things improved. Um, you run a maybe a broader set of benchmarks, measure the cache footprint of things. And so we wrote a skill that basically taught the model how to do
1237.4s
benchmarks you care about. You make some
1239.7s
modifications to improve the performance
1242.9s
you hope. Then you rerun the the
1245.8s
benchmarks, see where things improved.
1248.2s
Um, you run a maybe a broader set of
1249.9s
benchmarks, measure the cache footprint
1252.3s
of things. And so we wrote a skill that
1256.0s
basically taught the model how to do
1257.8s
most of those things in in var in various sequences so that it could actually you know do self-improving uh benchmark measurement benchmark improve you know code changes measure the performance improvement and then iterate on that and that that seemed to work uh pretty well for some kinds of problems. And it really just is us giving the
1260.3s
various sequences so that it could
1261.9s
actually you know do self-improving uh
1264.4s
benchmark measurement benchmark improve
1266.8s
you know code changes measure the
1268.7s
performance improvement and then iterate
1270.6s
on that and that that seemed to work uh
1272.6s
pretty well for some kinds of problems.
1274.5s
And it really just is us giving the
1278.4s
approach we would use as people to the model in a form that it could use. Wow, that seems very impressive. So you're saying you have this skill that if someone got access to it, it could do perform optimizations like Jeff Dean. Seems like the world would love this and is worth infinite amount of money to
1281.0s
model in a form that it could use.
1283.8s
Wow, that seems very impressive. So
1285.4s
you're saying you have this skill that
1287.4s
if someone got access to it, it could do
1290.9s
perform optimizations like Jeff Dean.
1293.3s
Seems like the world would love this and
1296.2s
is worth infinite amount of money to
1298.4s
someone have access to this. Oh. Uh we actually published a document maybe a few months ago called performance hints that Sanjay and I wrote that's like a 30-page document about you know various kinds of performance tricks and some people have taken that and then given it in summarized form to various models and
1299.7s
Oh. Uh we actually published a document
1301.8s
maybe a few months ago called
1303.1s
performance hints that Sanjay and I
1305.4s
wrote that's like a 30-page document
1307.4s
about you know various kinds of
1309.0s
performance tricks and some people have
1311.2s
taken that and then given it in
1313.2s
summarized form to various models and
1315.4s
seen that they that model can now get you know better at uh per reasoning about performance issues in code. So you heard it all here. You could actually get your own optimize your own code like Jeff Dean if you take this this paper that you published when performance hints. Yep. It's all free available, so you should
1317.8s
you know better at uh per reasoning
1319.5s
about performance issues in code.
1321.5s
So you heard it all here. You could
1322.9s
actually get your own optimize your own
1325.3s
code like Jeff Dean if you take this
1327.2s
this paper that you published when
1328.6s
performance hints. Yep.
1330.2s
It's all free available, so you should
1331.6s
all try it. Very cool. Very cool. Yeah. Now, you're talking about agents. Um, everyone here is probably building one or built one at some point. And I'm sure everyone has seen your agent go off the rail at perhaps like step 30 or 40. Like agents are great for like up to step, I
1332.9s
Very cool.
1333.5s
Very cool. Yeah.
1334.0s
Now, you're talking about agents. Um,
1336.0s
everyone here is probably building one
1337.6s
or built one at some point. And I'm sure
1340.4s
everyone has seen your agent go off the
1343.1s
rail at perhaps like step 30 or 40. Like
1346.2s
agents are great for like up to step, I
1348.2s
don't know, 10 or something and then gets shaky at step 50. What do you think is the constraint today? Is it like is the constraint today? Is it like context evaluators or just errors that compound because it's basically a openloop because it's basically a openloop system? Yeah, I mean obviously we want agents to
1349.7s
gets shaky at step 50. What do you think
1352.7s
is the constraint today? Is it like
1354.5s
is the constraint today? Is it like context
1356.0s
evaluators or just errors that compound
1358.6s
because it's basically a openloop
1360.2s
because it's basically a openloop system?
1361.5s
Yeah, I mean obviously we want agents to
1364.6s
be able to run for very long periods of time because that's how they're going to solve more and more complicated problems. Um but as you as you observe today, you know, they sometimes stop working after, you know, 10 10 interactions with the tools and so on. Um, and sometimes that's because the
1366.3s
time because that's how they're going to
1367.4s
solve more and more complicated
1369.1s
problems. Um but as you as you observe
1372.2s
today, you know, they sometimes stop
1375.4s
working after, you know, 10 10
1379.1s
interactions with the tools and so on.
1381.8s
Um, and sometimes that's because the
1384.3s
model is trying to do something it doesn't have a lot of experience doing. So it's been trained on a whole set of things and as soon as you get a little bit off the distribution of things it knows how to do then like most machine learning models it will you know its performance will suddenly will start to
1385.8s
doesn't have a lot of experience doing.
1387.8s
So it's been trained on a whole set of
1389.4s
things and as soon as you get a little
1391.4s
bit off the distribution of things it
1393.2s
knows how to do then like most machine
1395.5s
learning models it will you know its
1398.0s
performance will suddenly will start to
1400.0s
degrade and the farther you get off the comfort zone of what it knows how to do the the more likely it is to to not work as well. Um so there's a bunch of things you can do. So one is you know give the model skills and hints that kind of tend
1403.0s
comfort zone of what it knows how to do
1405.3s
the the more likely it is to to not work
1407.7s
as well. Um so there's a bunch of things
1410.7s
you can do. So one is you know give the
1413.4s
model skills and hints that kind of tend
1415.8s
to keep it in in the uh sort of more brightly lit path of things it does know how to do. Um, I think you know having multi- aent systems where you have multiple agents trying different approaches and you can evaluate you have maybe another model or another agent that's evaluating which ones of those
1418.8s
brightly lit path of things it does know
1420.6s
how to do. Um, I think you know having
1423.4s
multi- aent systems where you have
1425.6s
multiple agents trying different
1427.2s
approaches and you can evaluate you have
1429.0s
maybe another model or another agent
1431.4s
that's evaluating which ones of those
1433.3s
seem promising is another way to kind of in some sense search the path of pos search the space of possible solutions and stick to the ones that seem most promising and discard the ones that that didn't seem to work or maybe that went off the rails. or whatever. Um, and that's a very very useful general
1438.2s
in some sense search the path of pos
1441.0s
search the space of possible solutions
1443.6s
and stick to the ones that seem most
1446.4s
promising and discard the ones that that
1449.0s
didn't seem to work or maybe that went
1450.9s
off the rails. or whatever. Um, and
1453.1s
that's a very very useful general
1455.3s
technique is you know inference time compute to perform search over plausible ways of solving the problem that can get much much higher performance or much more reliability in longunning agent more reliability in longunning agent flows. How are some ways you implemented this particular workflow for your agents particular workflow for your agents internally?
1458.0s
compute to perform search over plausible
1460.6s
ways of solving the problem that can get
1463.6s
much much higher performance or much
1465.4s
more reliability in longunning agent
1467.9s
more reliability in longunning agent flows.
1470.1s
How are some ways you implemented this
1472.2s
particular workflow for your agents
1474.0s
particular workflow for your agents internally?
1475.6s
Yeah, I mean we have uh you know harnesses and then we have a whole set of skills uh particularly in the internal Google development environment. We have skills so that the agents can know how to use lots of our internal tooling for coding or for code reviews or for you know measuring performance or
1477.7s
harnesses and then we have a whole set
1479.8s
of skills uh particularly in the
1482.0s
internal Google development environment.
1483.8s
We have skills so that the agents can
1485.9s
know how to use lots of our internal
1488.1s
tooling for coding or for code reviews
1490.8s
or for you know measuring performance or
1494.4s
you know fetching log files. And um those are just skills that you can add to make the base model more capable even though it hasn't necessarily been trained on exactly the way that you know Google internal uh engineers would fetch log files from our you know proprietary system with the right kind of skill uh
1497.8s
those are just skills that you can add
1499.5s
to make the base model more capable even
1501.8s
though it hasn't necessarily been
1503.0s
trained on exactly the way that you know
1506.2s
Google internal uh engineers would fetch
1509.7s
log files from our you know proprietary
1512.6s
system with the right kind of skill uh
1515.4s
definition you can actually get it to definition you can actually get it to work uh and that that improves the usefulness of the agents. Now let's talk about uh where startups can can win. This section is one that I personally care a lot about because also everyone here in this room needs to decide what to build in
1516.9s
definition you can actually get it to work
1517.8s
uh and that that improves the usefulness
1519.9s
of the agents. Now let's talk about uh
1523.1s
where startups can can win. This section
1527.1s
is one that I personally care a lot
1529.0s
about because also everyone here in this
1531.7s
room needs to decide what to build in
1533.5s
the future of your future founder. So the thing about Google is you co-design everything on the system from the processors to the products. um which are the layers that someone like Google would keep building and compounding being better and and where does a two three person team can still does a two three person team can still win?
1537.3s
the thing about Google is you co-design
1539.2s
everything on the system from the
1541.6s
processors to the products.
1544.9s
um which are the layers that someone
1547.3s
like Google would keep building and
1549.6s
compounding being better and and where
1552.1s
does a two three person team can still
1555.6s
does a two three person team can still win?
1557.1s
Yeah, I mean I think obviously Google and and our Gemini models and and our hardware infrastructure are really trying to build very general models that can do almost anything. But in in a lot of cases that means that we don't have a lot of attention on particular domains where perhaps a really well-designed
1560.1s
and and our Gemini models and and our
1562.7s
hardware infrastructure are really
1564.6s
trying to build very general models that
1566.4s
can do almost anything. But in in a lot
1571.4s
of cases that means that we don't have a
1574.2s
lot of attention on particular domains
1576.3s
where perhaps a really well-designed
1579.7s
surface that and maybe a model and set of skills or maybe a specialized model that uh isn't in sort of a general mix of of things that our models do well can actually have a significant advantage because you can build something delightful and you know really high accuracy. really high quality for a
1582.7s
of skills or maybe a specialized model
1585.8s
that uh isn't in sort of a general mix
1588.5s
of of things that our models do well can
1591.4s
actually have a significant advantage
1593.7s
because you can build something
1594.8s
delightful and you know really high
1597.2s
accuracy. really high quality for a
1600.2s
domain that you are really passionate about. And I think that's that's where you know the two or three people in a room uh building that that they're really excited about can have an advantage. Um but I I would also caution that the general models are definitely getting better at a broader and broader
1601.8s
about. And I think that's that's where
1604.8s
you know the two or three people in a
1606.5s
room uh building that that they're
1609.1s
really excited about can have an
1610.9s
advantage. Um but I I would also caution
1615.0s
that the general models are definitely
1617.3s
getting better at a broader and broader
1619.3s
range of things. So you have to figure out, you know, is that thing you're working on, is that going to be a durable thing or do you think the models uh at the forefront are going to get better at that in the next six months or 12 months or is it something they're not
1620.6s
out, you know, is that thing you're
1623.6s
working on, is that going to be a
1624.8s
durable thing or do you think the models
1627.4s
uh at the forefront are going to get
1629.4s
better at that in the next six months or
1631.7s
12 months or is it something they're not
1633.2s
going to be able to do for a couple years or three years? And you know, you you want to weigh that as you're as you're deciding what to work on. So let's uh dive deeper into this. So the general models of course you're going to keep working on and keep making them all better.
1635.5s
years or three years? And you know, you
1637.9s
you want to weigh that as you're as
1639.9s
you're deciding what to work on.
1642.2s
So let's uh dive deeper into this. So
1644.6s
the general models of course you're
1646.2s
going to keep working on and keep making
1647.6s
them all better.
1649.6s
And how should the audience reason about what are those areas that uh it doesn't I mean h how should founder think about things to pick on and work on. Yeah. I mean I mean the most important thing is to pick something you're super excited about and want to build and you think would be useful in the world,
1651.8s
what are those areas that uh it doesn't
1655.1s
I mean h how should founder think about
1658.2s
things to pick on and work on.
1660.7s
Yeah. I mean I mean the most important
1662.7s
thing is to pick something you're super
1664.2s
excited about and want to build and you
1666.8s
think would be useful in the world,
1668.5s
right? So if you do that um that that's you're already way ahead uh than if you wake up and you're like oh I don't really want to do this or whatever or you're going to build something that is actually not that useful to to the world or to to many people. Um so I think
1671.4s
you're already way ahead uh than if you
1674.1s
wake up and you're like oh I don't
1675.4s
really want to do this or whatever or
1677.6s
you're going to build something that is
1679.0s
actually not that useful to to the world
1681.6s
or to to many people. Um so I think
1685.3s
that's the number one selection criteria I try to apply for what problem should I work on next. Um, second, I think you want to look at what the current more general models can do in that problem domain, right? You can you can test them with like, are they able to do this thing very well? And if they're
1687.2s
I try to apply for what problem should I
1690.2s
work on next. Um, second, I think you
1694.2s
want to look at what the current more
1696.9s
general models can do in that problem
1699.7s
domain, right? You can you can test them
1702.0s
with like, are they able to do this
1703.7s
thing very well? And if they're
1706.7s
completely failing, that's probably a good sign. If they're kind of able to do some of it but not very well, that's maybe not a great sign because that's a probably a a sign that the capability is starting to be present in those models and with more training data or larger scale models or or whatever it's likely
1708.3s
good sign. If they're kind of able to do
1710.8s
some of it but not very well, that's
1713.5s
maybe not a great sign because that's a
1715.3s
probably a a sign that the capability is
1718.0s
starting to be present in those models
1721.1s
and with more training data or larger
1723.7s
scale models or or whatever it's likely
1725.8s
to get better. So um you know look for something where the model succeeds 0 or 1 of the time not not 20. How do you find those? I mean are those things effectively uh out of distribution from the training set and what exactly is the problem shape that fits that? Yeah, I mean I think uh sometimes it's
1728.7s
something where the model succeeds 0 or
1731.5s
1 of the time not not 20.
1734.8s
How do you find those? I mean are those
1736.4s
things effectively
1738.8s
uh out of distribution from the training
1740.6s
set and what exactly is the problem
1742.8s
shape that fits that?
1744.5s
Yeah, I mean I think uh sometimes it's
1748.4s
uh a product that you build that might have access to particular kind of data that the underlying model might not the a general model. So it might be you're building something to help users organize all their own personal information and the model won't necessarily have access to that. And so there you can have a big advantage
1751.3s
have access to particular kind of data
1753.4s
that the underlying model might not the
1755.6s
a general model. So it might be you're
1758.2s
building something to help users
1759.5s
organize all their own personal
1761.0s
information and the model won't
1762.9s
necessarily have access to that. And so
1764.9s
there you can have a big advantage
1766.2s
because all of a sudden your model has visibility or your product has visibility into important data. Um it could be some incredibly hard problem where if you get the right training data and you can train a more specific model than a general purpose one, you can actually do that in a very affordable
1768.7s
visibility or your product has
1771.0s
visibility into important data. Um it
1775.0s
could be some incredibly hard problem
1778.1s
where if you get the right training data
1780.6s
and you can train a more specific model
1782.5s
than a general purpose one, you can
1784.5s
actually do that in a very affordable
1786.0s
way. you maybe it doesn't take that much compute to train a a niche model for this particular problem, but you can get something that's highly accurate. That can sometimes be a a really good uh building block for for solving a important problem that is maybe not handled very well by the general model.
1788.2s
compute to train a a niche model for
1791.0s
this particular problem, but you can get
1792.6s
something that's highly accurate. That
1794.6s
can sometimes be a a really good uh
1797.2s
building block for for solving a
1800.4s
important problem that is maybe not
1803.0s
handled very well by the general model.
1805.0s
I think that's interesting. I think there are basically two paths. The first path is uh a little bit funny is uh you guys are organizing the world's guys are organizing the world's information. information. Yeah, that's probably kind of well covered. Yeah. but organizing your personal information that's open which is funny. information that's open which is funny. Yeah.
1806.0s
there are basically two paths. The first
1807.7s
path is uh a little bit funny is uh you
1810.7s
guys are organizing the world's
1812.2s
guys are organizing the world's information.
1813.0s
information. Yeah,
1813.3s
that's probably kind of well covered.
1815.5s
Yeah. but organizing your personal
1817.7s
information that's open which is funny.
1821.0s
information that's open which is funny. Yeah.
1821.6s
And then the second path um you talked about more specialized models in certain domains. Can you tell us more about what are some of these domains? Yeah. I mean I think like if you look at uh my colleagues work on say alpha fold that was a very specific model for uh protein folding and it was highly
1824.2s
about more specialized models in certain
1827.1s
domains. Can you tell us more about what
1829.1s
are some of these domains?
1830.9s
Yeah. I mean I think like if you look at
1833.0s
uh my colleagues work on say alpha fold
1835.4s
that was a very specific model for uh
1838.1s
protein folding and it was highly
1840.2s
protein folding and it was highly successful um and was able to really handle that domain quite well so that all of a sudden you now have this amazing tool and model that can give you answers to questions about proteins and their structure um really effectively um but it's not a general model it's a very
1841.8s
um and was able to really handle that
1844.5s
domain quite well so that all of a
1846.2s
sudden you now have this amazing tool
1848.2s
and model that can give you answers to
1850.6s
questions about proteins and their
1852.6s
structure um really effectively um but
1855.9s
it's not a general model it's a very
1858.1s
specific one and there are other I domains where that kind of approach can work really well. Uh maybe in material science or chip design or things like that that uh will enable you to leverage the capabilities of a very accurate but but niche model uh to do things that are hard today.
1861.2s
domains where that kind of approach can
1863.4s
work really well. Uh maybe in material
1865.7s
science or chip design or things like
1867.4s
that that uh will enable you to leverage
1871.4s
the capabilities of a very accurate but
1874.6s
but niche model uh to do things that are
1878.1s
hard today.
1879.3s
That's a good example. So if some of you find a problem that's similar shape like alpha fold could be a good problem to work on. Now let's assume you found a problem to work on. We're going to talk a bit about how do you become a AI native founder? How do you really become
1881.0s
find a problem that's similar shape like
1883.0s
alpha fold could be a good problem to
1884.6s
work on. Now let's assume you found a
1886.7s
problem to work on. We're going to talk
1888.9s
a bit about how do you become a AI
1890.6s
native founder? How do you really become
1892.5s
good at it? Uh you in the past said that managing a fleet of agents, it's like 50 or 100 agents is all about writing really good crisp design docs or specs. And how do people get good at that? What what do those look like? Yeah, I mean I think uh you you'll have a lot more success when
1895.2s
managing a fleet of agents, it's like 50
1897.4s
or 100 agents is all about writing
1900.5s
really good crisp design docs or specs.
1905.0s
And how do people get good at that? What
1907.8s
what do those look like? Yeah, I mean I
1910.2s
think uh
1915.6s
you you'll have a lot more success when
1918.1s
working with your virtual agents if you can clearly specify what it is you want. And the clearer you are on what it is you want, the more the agent will have sort of guidelines and sort of rules of, you know, an outline of what it is trying to accomplish. Um whereas if you
1920.7s
can clearly specify what it is you want.
1923.0s
And the clearer you are on what it is
1925.0s
you want, the more the agent will have
1927.9s
sort of guidelines and sort of rules of,
1930.5s
you know, an outline of what it is
1932.6s
trying to accomplish. Um whereas if you
1935.4s
don't specify very much stuff, the agent has to sort of infer what it is you meant. And in many cases, it might infer things that are different than what you imagined. So we've always told computer scientists from the very beginning that really it's really important to specify what it is, what's the software that
1937.4s
has to sort of infer what it is you
1939.7s
meant. And in many cases, it might infer
1942.2s
things that are different than what you
1943.8s
imagined. So we've always told computer
1946.2s
scientists from the very beginning that
1947.9s
really it's really important to specify
1949.4s
what it is, what's the software that
1951.8s
you're writing is trying to accomplish before then going and writing it. And so now we actually have agent-based systems that can do the writing, but the importance of specifying what what it is you want has actually gone up because before you'd be handing it off to a very intelligent human who maybe has context
1954.3s
before then going and writing it. And so
1956.8s
now we actually have agent-based systems
1958.7s
that can do the writing, but the
1960.7s
importance of specifying what what it is
1962.6s
you want has actually gone up because
1965.4s
before you'd be handing it off to a very
1968.1s
intelligent human who maybe has context
1970.6s
or can ask you follow-up questions. Um, and agents can sometimes do that, but I I think clear specifications is is a really good idea. Um, and to give you an example of a a a use of a coding agent that works extremely well is you can ask today's models to translate software from one computer language to another
1973.0s
and agents can sometimes do that, but I
1974.8s
I think clear specifications is is a
1978.3s
really good idea. Um, and to give you an
1980.6s
example of a a a use of a coding agent
1984.0s
that works extremely well is you can ask
1987.5s
today's models to translate software
1990.3s
from one computer language to another
1993.0s
very effectively because in that case you actually have a incredibly detailed specification. You have the whole software that says what the system is supposed to do. And so if you have a Python implementation of something and you want a Go implementation of it, you know, that is something that the models seem incredibly capable at doing these
1995.8s
you actually have a incredibly detailed
1998.4s
specification. You have the whole
1999.5s
software that says what the system is
2001.8s
supposed to do. And so if you have a
2003.9s
Python implementation of something and
2005.8s
you want a Go implementation of it, you
2008.2s
know, that is something that the models
2009.6s
seem incredibly capable at doing these
2011.4s
days because you can it can sort of take all the tests that are in Python, make sure they pass in the Go version, translate the tests to Go, um you know, compare uh behavioral differences between the implementations until there aren't any um and be you know, highly effective because that spec is so clear.
2014.5s
all the tests that are in Python, make
2016.7s
sure they pass in the Go version,
2018.5s
translate the tests to Go, um you know,
2021.0s
compare uh behavioral differences
2023.3s
between the implementations until there
2025.3s
aren't any um and be you know, highly
2028.6s
effective because that spec is so clear.
2031.0s
Hm. Now let's assume now every founder gets good at running hundreds of agents at the same time and all the code is written for them by the agents. What becomes the scarce skill? Yeah, I mean I think it's really having incredibly good taste in what you ask your agents to work on, right? That is
2033.6s
gets good at running hundreds of agents
2035.3s
at the same time and all the code is
2037.4s
written for them by the agents. What
2039.9s
becomes the scarce skill?
2043.8s
Yeah, I mean I think it's really having
2047.4s
incredibly good taste in what you ask
2049.5s
your agents to work on, right? That is
2051.8s
the the crux of you know from my background uh a research problem. You know, a researcher can have all the tools and all the techniques, but often most of the battle is what problem are you gonna spend your time on? And if you pick the problem well and you succeed in
2055.2s
background uh a research problem. You
2057.9s
know, a researcher can have all the
2059.2s
tools and all the techniques, but often
2062.8s
most of the battle is what problem are
2065.6s
you gonna spend your time on? And if you
2067.9s
pick the problem well and you succeed in
2070.5s
in in solving it, that's way better than if you, you know, uh, delightfully execute a research investigation into a rather boring problem. And so that high level wisdom of what to work on, I think is incredibly important. And I think models are not necessarily going to be that good at it. So you're going to have people steering
2073.3s
if you, you know, uh, delightfully
2076.5s
execute a research investigation into a
2079.9s
rather boring problem. And so that high
2082.5s
level wisdom of what to work on, I think
2084.9s
is incredibly important. And I think
2086.5s
models are not necessarily going to be
2087.8s
that good at it. So you're going to have
2090.6s
people steering
2094.0s
uh a lot of AI assisted computation in order to accomplish great things and more quickly. Um but that essence of of what it is you want your models to do is the the the key thing you should focus the the the key thing you should focus on. So let's talk a bit more about taste
2097.0s
order to accomplish great things and
2099.3s
more quickly. Um but that essence of of
2103.0s
what it is you want your models to do is
2105.4s
the the the key thing you should focus
2107.4s
the the the key thing you should focus on.
2108.3s
So let's talk a bit more about taste
2109.8s
because it gets talked a lot about right now in this current era with agent coding. How do you exactly build taste and do that? I mean, yeah, that sounds so esoteric. How do you make it so esoteric. How do you make it concrete? Yeah, I mean it it is a difficult thing. It's not like there's a measurable
2112.0s
now in this current era with agent
2114.2s
coding. How do you exactly build taste
2117.7s
and do that? I mean, yeah, that sounds
2120.5s
so esoteric. How do you make it
2122.8s
so esoteric. How do you make it concrete?
2123.4s
Yeah, I mean it it is a difficult thing.
2126.5s
It's not like there's a measurable
2128.2s
objective of of taste in a lot of cases. Um, I think some of it is from experience. You know, working on a lot of different problems in the past kind of teaches you about what kinds of problems might be interesting in the future or what kinds of things might be just barely possible by cobbling
2131.8s
Um, I think some of it is from
2133.7s
experience. You know, working on a lot
2136.1s
of different problems in the past kind
2138.2s
of teaches you about what kinds of
2140.6s
problems might be interesting in the
2142.2s
future or what kinds of things might be
2145.8s
just barely possible by cobbling
2148.6s
together these previous approaches and then some open problems you might have to work on in order to get to something kind of magical or or you know, highly useful. Um, another way you can get more experience for yourself is to just write down a bunch of things you think might be important in the next 12 months. And
2151.3s
then some open problems you might have
2153.2s
to work on in order to get to something
2155.4s
kind of magical or or you know, highly
2157.6s
useful. Um,
2160.4s
another way you can get more experience
2162.4s
for yourself is to just write down a
2166.0s
bunch of things you think might be
2167.2s
important in the next 12 months. And
2170.3s
maybe you pick one of them to work on, but go back and evaluate in 12 months of these other things, which ones actually seemed important or which ones did other people in the world go out and and create and which ones did they did not seem to do yet. um that can give you a
2172.6s
but go back and evaluate in 12 months of
2175.8s
these other things, which ones actually
2178.0s
seemed important or which ones did other
2180.6s
people in the world go out and and
2182.3s
create and which ones did they did not
2184.9s
seem to do yet. um that can give you a
2187.3s
lot more samples for your own sort of taste creation uh capability. Um and and that's an important skill to have. I think a third way we were talking earlier was doing very crazy thought earlier was doing very crazy thought experiments. Oh yeah, that's another good way. I mean I think uh sometimes it's good to
2189.6s
taste creation uh capability. Um and and
2193.9s
that's an important skill to have.
2195.9s
I think a third way we were talking
2197.8s
earlier was doing very crazy thought
2201.7s
earlier was doing very crazy thought experiments.
2202.6s
Oh yeah, that's another good way. I mean
2204.8s
I think uh sometimes it's good to
2209.0s
not take as a given things that most people seem to take as a as a given. Um so I was doing a crazy thought experiment with some colleagues the other day about you know for 60 years the whole silicon uh chip design industry uh design and fabrication industry have you know done tremendous work to make
2212.3s
people seem to take as a as a given. Um
2215.7s
so I was doing a crazy thought
2217.0s
experiment with some colleagues the
2218.2s
other day about you know
2222.4s
for 60 years the whole silicon
2226.0s
uh chip design industry uh design and
2229.3s
fabrication industry have
2232.6s
you know done tremendous work to make
2235.0s
smaller and smaller scale transistors that are uh very low error rate right like because what what the assumption that we want is that every chip we manufacture of the same design should be identical to every other chip. You don't want any bits to flip. You don't want any bits to flip. Everything
2237.7s
that are uh very low error rate right
2242.1s
like because what what the assumption
2244.0s
that we want is that every chip we
2245.8s
manufacture of the same design should be
2248.6s
identical to every other chip.
2250.0s
You don't want any bits to flip.
2251.3s
You don't want any bits to flip. Everything
2251.7s
no bits should flip. There's all kinds of things you there's all kinds of error margins built into you know memories have ECC memory these days. um you know at the at the macro scale we don't make that assumption when we're building large scale distributed systems right we we build reliable large scale distributed file systems out of
2253.8s
of things you there's all kinds of error
2255.7s
margins built into you know memories
2258.1s
have ECC memory these days. um you know
2261.7s
at the at the macro scale we don't make
2265.0s
that assumption when we're building
2266.3s
large scale distributed systems right we
2269.1s
we build reliable large scale
2271.8s
distributed file systems out of
2273.8s
unreliable parts right like individual discs can fail but your data should be safe and so we have mechanisms at a higher level to enable us to have um you know three copies of the data on three different machines and three different racks so that if any rack switch or individual ual machine or disk fails,
2276.2s
discs can fail but your data should be
2278.6s
safe and so we have mechanisms at a
2280.3s
higher level to enable us to have um you
2284.2s
know three copies of the data on three
2286.3s
different machines and three different
2287.5s
racks so that if any rack switch or
2289.8s
individual ual machine or disk fails,
2291.7s
you still have your data. We have read Solomon encoding techniques. Um, but we don't seem to do this at a really extreme level in the uh sort of transistor level scale of of the technology we're working on. So what would h basically a interesting thought experiment is what would happen if you
2293.7s
Solomon encoding techniques. Um, but we
2297.2s
don't seem to do this at a really
2299.9s
extreme level in the uh sort of
2303.4s
transistor level scale of of the
2306.5s
technology we're working on. So what
2308.2s
would h basically a interesting thought
2310.2s
experiment is what would happen if you
2312.0s
tried to build a system out of transistors that might have you know 20 errors per day. Oh my god. rather than one every million years, right? That would be a very different design point and might be might enable you to do really interesting things in the fabrication side of things. You have very different
2314.3s
transistors that might have you know 20
2317.5s
errors per day.
2318.8s
Oh my god. rather than one every million
2320.9s
years, right? That would be a very
2323.2s
different design point and might be
2325.4s
might enable you to do really
2326.6s
interesting things in the fabrication
2328.1s
side of things. You have very different
2330.2s
kind of design methodologies because if you want to get a signal from here to there, you and you have these super unreliable transistors. You might have very different ways of signaling. You might send it along multiple redundant paths uh in order to make sure that it gets along one of them. Um, and I think
2332.8s
you want to get a signal from here to
2334.3s
there, you and you have these super
2336.3s
unreliable transistors. You might have
2338.5s
very different ways of signaling. You
2340.3s
might send it along multiple redundant
2342.0s
paths uh in order to make sure that it
2344.5s
gets along one of them. Um, and I think
2347.1s
that would be a pretty interesting set of thought experiments. I'm not saying we should go do this, but you know, that's the kind of thing where you do want to, you know, occasionally question assumptions. Now, oftentimes these thought experiments don't work out because there are very good reasons that, you know, for the last 50 years,
2348.9s
of thought experiments. I'm not saying
2350.6s
we should go do this, but you know,
2352.3s
that's the kind of thing where you do
2353.7s
want to, you know, occasionally question
2357.8s
assumptions. Now, oftentimes these
2360.8s
thought experiments don't work out
2362.2s
because there are very good reasons
2364.1s
that, you know, for the last 50 years,
2366.0s
we've done this thing this way and not that way. But it it's good to kind of revisit those every so often. That is so wild. Well, I mean, it's starting to rhyme a lot with neuromorphic computing or the human brain and and how nature works. I mean, exactly like signals in our brain are not especially reliable from
2368.2s
that way. But it it's good to kind of
2370.6s
revisit those every so often.
2372.2s
That is so wild. Well, I mean, it's
2373.8s
starting to rhyme a lot with
2375.8s
neuromorphic computing or the human
2377.8s
brain and and how nature works.
2379.8s
I mean, exactly like signals in our
2381.5s
brain are not especially reliable from
2383.9s
getting one place to another. And so, I think in brains when there are really important things you need to get from one place to another, there are multiple pathways that that enable you to sort of do that. What is uh in I mean, you have such an impressive career. What is one of these
2386.4s
think in brains when there are really
2388.9s
important things you need to get from
2390.3s
one place to another, there are multiple
2391.8s
pathways that that enable you to sort of
2394.4s
do that.
2396.6s
What is uh in I mean, you have such an
2399.0s
impressive career. What is one of these
2401.1s
crazy assumptions that you threw out of the window that actually built a consequential system in the past? that worked out actually. Yeah, I mean I think uh well TPUs is a good example like being able to specialize hardware for a very niche problem domain before that problem domain seemed as important as it is
2402.9s
the window that actually built a
2404.6s
consequential system in the past?
2410.9s
that worked out actually.
2413.3s
Yeah, I mean I think uh well TPUs is a
2415.4s
good example like being able to
2417.7s
specialize hardware for a very niche
2420.5s
problem domain before that problem
2422.2s
domain seemed as important as it is
2424.1s
today uh is one thought experiment. Um you know I think the the origin of map produce is another good example. So we had worked the you know my Sanjay and myself and a number of other colleagues had worked on various iterations of the crawling and indexing system at Google and you know we'd sort
2428.3s
you know I think the
2430.7s
the origin of map produce is another
2433.6s
good example.
2435.0s
So we had worked the you know my Sanjay
2437.8s
and myself and a number of other
2439.6s
colleagues had worked on various
2441.6s
iterations of the crawling and indexing
2443.4s
system at Google and you know we'd sort
2447.0s
of written lots of hand parallelized code with lots of checkpointing to make sure it would be robust and reliable if it was running on a 100 computers or a thousand computers and some of those died. Um, but that code tended to be intermixed with the actually relatively simple thing you often were trying to do
2449.3s
code with lots of checkpointing to make
2451.5s
sure it would be robust and reliable if
2453.4s
it was running on a 100 computers or a
2455.8s
thousand computers and some of those
2457.4s
died. Um, but that code tended to be
2462.2s
intermixed with the actually relatively
2464.3s
simple thing you often were trying to do
2466.8s
like I just want to like look at all the contents of all the web pages and then compute on the side a mapping from URL to you know what language is this page in the text of this page. Um, and it would get obscured by all this kind of other code for parallelization and
2469.7s
contents of all the web pages and then
2471.6s
compute on the side a mapping from URL
2474.2s
to you know what language is this page
2476.4s
in the text of this page. Um, and it
2480.4s
would get obscured by all this kind of
2482.5s
other code for parallelization and
2484.2s
reliability. And so we sort of remembered our training in functional languages and realized we could squint at those problems and developed this map produce abstraction that you could have above this implementation and then below the implementation you could put all the checkpointing and reliability mechanisms into that lower level library that everything could then build on. And so
2487.2s
remembered our training in functional
2489.1s
languages and realized we could squint
2491.4s
at those problems and developed this map
2493.8s
produce abstraction that you could have
2496.3s
above this implementation and then below
2499.6s
the implementation you could put all the
2502.1s
checkpointing and reliability mechanisms
2505.4s
into that lower level library that
2507.8s
everything could then build on. And so
2509.5s
that became a hugely successful way of of dealing with very large scale computations at Google in a robust and reliable way. From that thought experiment of like well if we squint at it could we find lots of problems that fit into this abstraction. That's impressive. So this thought experiment led you to create map reduce. Yeah. Awesome.
2512.1s
of dealing with very large scale
2513.7s
computations at Google in a robust and
2515.8s
reliable way. From that thought
2518.0s
experiment of like well if we squint at
2520.2s
it could we find lots of problems that
2521.9s
fit into this abstraction.
2523.9s
That's impressive. So this thought
2526.4s
experiment led you to create map reduce.
2528.3s
Yeah. Awesome.
2529.9s
Now let's go back to you talked a bit about um about your interest right now working on a lot of customized hardware. So right now alpha chip lays out chips. Now you also got alpha evolve that proposes solutions, evaluates them and keeps all the ones that work. Seems like you're starting to
2531.8s
about um about your interest right now
2535.2s
working on a lot of customized hardware.
2537.4s
So right now alpha chip
2539.7s
lays out chips. Now you also got alpha
2541.8s
evolve that proposes solutions,
2544.3s
evaluates them and keeps all the ones
2546.3s
that work. Seems like you're starting to
2547.7s
build all these system that can compound and build AI that builds AI. Yeah. I mean I think more generally there's a there's this sort of the foundation of the scientific method of you propose an experiment you implement what you need to run the experiment and you evaluate the experiment and then you get results from
2550.0s
and build AI that builds AI.
2553.0s
Yeah. I mean I think more generally
2554.8s
there's a there's this sort of
2559.2s
the foundation of the scientific method
2561.4s
of you propose an experiment you
2564.3s
implement what you need to run the
2566.2s
experiment and you evaluate the
2568.4s
experiment and then you get results from
2570.2s
that and I think there are more and more problems that are now possible to implement where that whole loop of running you know not just a few experiments but running many many experiments because you're able to automate that loop and make the latency of that loop extremely low is going to be really really important. It's going
2572.9s
problems that are now possible to
2575.3s
implement where that whole loop of
2578.6s
running you know not just a few
2580.5s
experiments but running many many
2581.9s
experiments because you're able to
2583.2s
automate that loop and make the latency
2585.8s
of that loop extremely low is going to
2588.4s
be really really important. It's going
2589.6s
to enable us to tackle you know lots of different problem domains in science and engineering and machine learning uh model design itself and also in engineering tasks like designing chips. And so if you can actually do those things in an automated way and have some orchestration framework that can take very high level objectives and break
2592.2s
different problem domains in science and
2594.2s
engineering and machine learning uh
2597.5s
model design itself and also in
2600.7s
engineering tasks like designing chips.
2603.5s
And so if you can actually do those
2605.0s
things in an automated way and have some
2608.6s
orchestration framework that can take
2611.2s
very high level objectives and break
2613.6s
them down into subpros and each of those subpros can be one of these automated loop that is exploring the best way to solve that sub problem and then a orchestration framework that can put together subpros solutions into a you know the overall solution for the higher level problem that's going to be really impactful and it's really really
2615.8s
subpros can be one of these automated
2618.3s
loop that is exploring the best way to
2620.7s
solve that sub problem and then a
2623.8s
orchestration framework that can put
2626.1s
together subpros solutions into a you
2630.2s
know the overall solution for the higher
2632.3s
level problem that's going to be really
2634.6s
impactful and it's really really
2636.4s
important and I think it'll enable us to do you know accelerate machine learning progress it'll enable us to accelerate science and enable us to accelerate engineering and I think that's that's going to be amazing that sounds awesome I mean it sounds like a lot of fields basically where you can have very good evaluators and maybe
2638.9s
do you know accelerate machine learning
2641.3s
progress it'll enable us to accelerate
2643.8s
science and enable us to accelerate
2645.9s
engineering and I think that's that's
2647.5s
going to be amazing
2649.1s
that sounds awesome I mean it sounds
2650.6s
like a lot of fields basically where you
2652.3s
can have very good evaluators and maybe
2655.6s
adjacent to basically things that can be formally verified right those are ripe for AI systems that can self-improve Yeah, I think in a lot of cases sometimes your evaluators need to be made much faster. Mhm. So as an example, my colleagues did some work maybe a decade ago on um some uh problems in quantum chemistry where
2657.4s
formally verified right those are ripe
2660.0s
for AI systems that can self-improve
2662.8s
Yeah, I think in a lot of cases
2665.7s
sometimes your evaluators need to be
2667.5s
made much faster. Mhm.
2669.0s
So as an example, my colleagues did some
2672.0s
work maybe a decade ago on um some uh
2676.2s
problems in quantum chemistry where
2678.2s
you're trying to understand the properties of a particular molecule and you can you know generate some molecule configuration and then you want to understand what properties it has. And so you can run a very computationally intensive density functional theory simulator which is something that might take like a a night of computation to
2679.2s
properties of a particular molecule and
2681.7s
you can you know generate some molecule
2684.2s
configuration and then you want to
2685.5s
understand what properties it has. And
2688.2s
so you can run a very computationally
2690.1s
intensive density functional theory
2692.7s
simulator which is something that might
2694.7s
take like a a night of computation to
2697.4s
tell you the answer for one thing. Um but what my colleagues did was take a bunch of output from those simulation runs the input molecule configurations and the outputs of the the expensive simulator and then use it to train a neural approximation to the simulator. So this is now a validation device, but instead of it taking a
2700.6s
but what my colleagues did was
2704.7s
take a bunch of output from those
2707.2s
simulation runs the input molecule
2709.5s
configurations and the outputs of the
2711.3s
the expensive simulator and then use it
2714.1s
to train a neural approximation to the
2716.6s
simulator. So this is now a validation
2719.8s
device, but instead of it taking a
2722.6s
night, they made something that was 300,000 times faster. 300,000 times faster. Wow. And nearly as accurate as running the full scale simulator. So now that completely changes how you would do science, right? Because now you have 10 million things to screen. you know, you could do that while you go to lunch rather than it being a six-month
2725.0s
300,000 times faster.
2726.6s
300,000 times faster. Wow.
2727.0s
And nearly as accurate as running the
2728.9s
full scale simulator. So now that
2732.2s
completely changes how you would do
2733.8s
science, right? Because now you have 10
2735.8s
million things to screen. you know, you
2738.2s
could do that while you go to lunch
2740.5s
rather than it being a six-month
2742.2s
endeavor where you could try to scrape together enough compute to to run all these simulations. And I think there's a lot of room in a lot of domains for much faster validation models, possibly learned valu validation models that can uh you know get you a a approximation to the true answer much much more rapidly.
2744.2s
together enough compute to to run all
2746.4s
these simulations. And I think there's a
2750.1s
lot of room in a lot of domains for much
2753.2s
faster validation models, possibly
2755.0s
learned valu validation models that can
2759.2s
uh you know get you a a approximation to
2762.2s
the true answer much much more rapidly.
2764.5s
And that changes how those experimental loops can be thought of and how quickly you can go around those loops. What are some of the spaces and problems that you're super excited that this super sped up scientific method is going to solve or achieve? What particular problems or achieve? What particular problems or spaces?
2766.4s
loops can be thought of and how quickly
2768.2s
you can go around those loops.
2770.2s
What are some of the
2773.0s
spaces and problems that you're super
2774.6s
excited that this super sped up
2777.8s
scientific method is going to solve or
2780.6s
achieve? What particular problems or
2782.3s
achieve? What particular problems or spaces?
2783.6s
Yeah, I mean I think uh one, right? So can we have a model that is able to recursively self-improve itself by running lots of experiments and you know if you think about how models are improved today in large research teams you know what usually happens is people think of some ideas they run a bunch of smallcale
2790.5s
one, right? So can we have a model that
2792.6s
is able to recursively self-improve
2795.0s
itself by running lots of experiments
2797.6s
and you know if you think about how
2799.4s
models are improved today in large
2801.8s
research teams you know what usually
2804.6s
happens is people think of some ideas
2806.9s
they run a bunch of smallcale
2808.3s
experiments they see if those small scale experiments worked out well if so they take the most promising ones of those they try them at larger scale and that gets then evaluated and then the results get integr ated together into you know a new recipe for your model. Um but I think there's no uh you know real
2810.4s
scale experiments worked out well if so
2813.0s
they take the most promising ones of
2814.5s
those they try them at larger scale and
2816.9s
that gets then evaluated and then the
2820.0s
results get integr ated together into
2822.6s
you know a new recipe for your model. Um
2826.0s
but I think there's no uh you know real
2829.7s
impediment to making that be a much more automated loop where the model itself decides it's going to explore or maybe with a nudge from some people uh at the various highest level like oh why don't you try some new ideas around model architectures that incorporate this and then it will go run lots of experiments
2831.9s
automated loop where the model itself
2835.8s
decides it's going to explore or maybe
2837.8s
with a nudge from some people uh at the
2840.6s
various highest level like oh why don't
2842.5s
you try some new ideas around model
2844.9s
architectures that incorporate this and
2847.6s
then it will go run lots of experiments
2850.1s
uh see which ones work and then those will get incorporated at a much more rapid rate and uh you know effectively you want to optimize you know your discoveries per unit of compute input. Very cool. Very cool. Yeah. Now going back to the room as all of you will become at some point founders or
2851.8s
will get incorporated at a much more
2853.6s
rapid rate and uh you know effectively
2856.8s
you want to optimize you know your
2859.8s
discoveries per unit of compute input.
2864.0s
Very cool.
2864.9s
Very cool. Yeah.
2865.7s
Now going back to the room as all of you
2869.8s
will become at some point founders or
2871.8s
start your careers you will probably collect lots of rejections. That will happen. Uh it has happened to you too Jeff. I mean there's a story that in Jeff. I mean there's a story that in 2014 you with Jeff Hinton and Oral Fin wrote a paper on distillation which has to do with taking a big
2874.2s
collect lots of rejections. That will
2876.6s
happen. Uh it has happened to you too
2879.8s
Jeff. I mean there's a story that in
2883.0s
Jeff. I mean there's a story that in 2014
2885.0s
you with Jeff Hinton and Oral Fin wrote
2889.4s
a paper on distillation
2892.3s
which has to do with taking a big
2895.5s
teacher model to train a much smaller and more efficient model that's a lot cheaper to compute less model parameters and it has become a trick that everyone is using right now in industry. Yeah. is using right now in industry. Yeah. And the thing is this paper got rejected at the thing is this paper got rejected at Europe.
2898.3s
and more efficient model that's a lot
2901.4s
cheaper to compute less model parameters
2904.5s
and it has become a trick that everyone
2907.8s
is using right now in industry. Yeah.
2910.7s
is using right now in industry. Yeah. And
2912.4s
the thing is this paper got rejected at
2914.8s
the thing is this paper got rejected at Europe.
2916.1s
Yeah. I mean Yeah. I mean I think I don't fault the program committee because you know a lot of times a paper gets three reviews and someone will look at one of the reviewers will look at it and in this case they said oh it's unlikely to have significant impact. Unlikely to have significant impact. But
2919.1s
don't fault the program committee
2921.0s
because you know a lot of times a paper
2924.0s
gets three reviews and someone will look
2926.3s
at one of the reviewers will look at it
2928.6s
and in this case they said oh it's
2930.1s
unlikely to have significant impact.
2932.0s
Unlikely to have significant impact. But
2933.8s
you know I think you know when we wrote the paper we actually saw this was a super important problem because we knew making cheaper highly capable models from larger scale models was something we desperately wanted to do because we wanted to serve models to more and more people in many different domains like
2935.4s
the paper we actually saw this was a
2937.2s
super important problem because we knew
2940.2s
making cheaper highly capable models
2943.3s
from larger scale models was something
2945.0s
we desperately wanted to do because we
2947.2s
wanted to serve models to more and more
2949.3s
people in many different domains like
2951.0s
speech or vision. Um but you know sometimes the reviewer maybe didn't have that that experience because maybe they're not thinking about you know largecale AI services and are thinking about you know is this a fundamental advance um so so you know it gets rejected every so often that's fine we put it on archive people read it people
2953.5s
sometimes the reviewer maybe didn't have
2955.1s
that that experience because maybe
2957.0s
they're not thinking about you know
2959.4s
largecale AI services and are thinking
2962.6s
about you know is this a fundamental
2964.2s
advance um so so you know it gets
2967.6s
rejected every so often that's fine we
2969.8s
put it on archive people read it people
2971.7s
use it it's all good uh and you know we do use it in making our flash models for example from our larger scale pro model that's partly why our flash models for example in Gemini are so capable uh relative to their size and and speed. They're some of the best in the benchmark for their model size class.
2974.9s
do use it in making our flash models for
2977.8s
example from our larger scale pro model
2980.2s
that's partly why our flash models for
2981.8s
example in Gemini are so capable uh
2985.0s
relative to their size and and speed.
2987.4s
They're some of the best in the
2988.4s
benchmark for their model size class.
2990.2s
Yeah. Just impressive. And I think part of the lesson is that even if you get rejected, keep going. Yeah. That's that's the lesson I would distill from that. distill from that. laughter Um no, I think the fun thing is that you basically join when you when you join Google as a 20 person startup back in
2992.7s
of the lesson is that even if you get
2994.5s
rejected, keep going.
2996.4s
Yeah. That's that's the lesson I would
2998.0s
distill from that.
3000.6s
distill from that. laughter
3001.0s
Um no, I think the fun thing is that you
3004.0s
basically join when you when you join
3005.5s
Google as a 20 person startup back in
3008.0s
Google as a 20 person startup back in 1999. Now, if you were to take the young Jeff Dean from way back then to teleransport him to now today. him to now today. Yeah. In this era with your skills. I'm feeling so vigorous and and young I'm feeling so vigorous and and young now.
3009.9s
Now, if you were to take the young Jeff
3012.7s
Dean from way back then to teleransport
3016.2s
him to now today.
3018.0s
him to now today. Yeah.
3018.6s
In this era with your skills.
3020.2s
I'm feeling so vigorous and and young
3023.2s
I'm feeling so vigorous and and young now.
3024.4s
Um what would you do? Do you join a frontier lab, start a company? I don't know what what would you do? The c the Jeff theme today 25-year-old Jeff theme. Yeah. I mean, it's always hard to say and it's a very personal choice of what it is you want to spend your time on. Um, to me, some
3026.9s
frontier lab, start a company? I don't
3029.8s
know what what would you do? The c the
3034.0s
Jeff theme today 25-year-old Jeff theme.
3036.6s
Yeah. I mean,
3039.6s
it's always hard to say and it's a very
3041.2s
personal choice of what it is you want
3042.7s
to spend your time on. Um, to me, some
3046.3s
of the most important questions are, are you going to work on something you really care about, will you're working on that? And if you're able to make progress on it with a bunch of colleagues you like working with uh if you're able to make collectively solve it or make progress on it, will that
3049.8s
are you going to work on something you
3052.3s
really care about, will you're working
3055.8s
on that? And if you're able to make
3057.8s
progress on it with a bunch of
3059.6s
colleagues you like working with uh if
3062.4s
you're able to make collectively solve
3064.6s
it or make progress on it, will that
3067.1s
make a difference in the world in some positive way, right? Like will you suddenly be able to do something and offer that service to you know offer that service to you know partically it's a broader thing. It'll help programmers or it will help all consumers. uh on the internet or or
3068.6s
positive way, right? Like will you
3070.6s
suddenly be able to do something and
3072.8s
offer that service to you know
3075.4s
offer that service to you know partically
3081.5s
it's a broader thing. It'll help
3082.8s
programmers or it will help all
3084.8s
consumers. uh on the internet or or
3087.8s
other things. Um what you you know what you should strive to do is to have impact in the world that is positive and to work with people you enjoy working with and to you know uh work hard and and do your best. Um so in terms of say the particular trade-off you offered
3092.0s
you should strive to do is to have
3093.9s
impact in the world that is positive and
3096.7s
to work with people you enjoy working
3098.2s
with and to you know uh work hard and
3102.0s
and do your best. Um so in terms of say
3106.0s
the particular trade-off you offered
3108.2s
joining a frontier lab versus say starting a company with just one or two or three of you you and your close friends. Um, I think those are different experiences, right? In a in a large established organization, you have some structure. You have lots and lots of amazing colleagues who know lots of
3110.2s
starting a company with just one or two
3112.3s
or three of you you and your close
3114.3s
friends. Um, I think those are different
3117.3s
experiences, right? In a in a large
3119.4s
established organization, you have some
3122.1s
structure. You have lots and lots of
3125.0s
amazing colleagues who know lots of
3126.8s
things you don't. Um, you have lots of interesting problems that uh you can work on and h you already have a platform for impact by your work, you know, influencing lots and lots of people in the world already. Um and then as a very small startup, you know, you have to have something
3130.0s
interesting problems that uh you can
3132.3s
work on and h you already have a
3135.6s
platform for impact by your work, you
3138.7s
know, influencing lots and lots of
3140.2s
people in the world already. Um and then
3143.0s
as a very small startup,
3146.0s
you know, you have to have something
3148.5s
you're passionate about and there's a lot of risk in taking on, you know, working on that particular problem in a way that uh you're going to succeed and you're going to grow a, you know, an endeavor in order to do that. But that can also be incredibly rewarding, I would imagine. So I I think um you know
3151.8s
lot of risk in taking on, you know,
3154.4s
working on that particular problem in a
3156.1s
way that uh you're going to succeed and
3158.7s
you're going to grow a, you know, an
3160.5s
endeavor in order to do that. But that
3162.6s
can also be incredibly rewarding, I
3164.9s
would imagine. So I I think um you know
3168.4s
it's really up to personal taste but but at the very least regardless of what path you take ask yourself if I work on this problem and the best possible outcome happens you know will the world be a lot better in some way or will the world go eh that's kind of cool but
3170.9s
at the very least regardless of what
3173.2s
path you take ask yourself if I work on
3176.5s
this problem and the best possible
3178.5s
outcome happens you know will the world
3181.5s
be a lot better in some way or will the
3183.7s
world go eh that's kind of cool but
3185.9s
world go eh that's kind of cool but whatever. Uh that's not the kind of thing you should spend your time on. Now let's talk a bit about more about that second path of working with people that you really like in a small team. You've been able to be an incredible mentor and manager to many many
3186.8s
Uh that's not the kind of thing you
3188.6s
should spend your time on.
3190.7s
Now let's talk a bit about more about
3192.6s
that second path of working with people
3195.1s
that you really like in a small team.
3198.0s
You've been able to be an incredible
3200.6s
mentor and manager to many many
3202.6s
engineers and you've been able to build huge systems and what are some some of the lessons for everyone here on how to get the most and how to work with smart people or find smart people? Yeah, I mean, you always want to find people who have really good skills in some some area
3205.3s
huge systems and what are some some of
3209.8s
the lessons for everyone here on how to
3212.5s
get the most and how to work with smart
3214.2s
people or find smart people?
3216.7s
Yeah, I mean,
3219.4s
you always want to find people who have
3222.9s
really good skills in some some area
3225.1s
that's needed in, you know, a team you're trying to form, whether that's inside a company or uh starting a company. Um, but you also want to find people that are people you delight being around, right? because you're going to spend a lot of time around people working on really hard problems and you
3227.8s
you're trying to form, whether that's
3230.4s
inside a company or uh starting a
3232.7s
company. Um, but you also want to find
3235.7s
people that are people you delight being
3239.1s
around, right? because you're going to
3240.5s
spend a lot of time around people
3242.8s
working on really hard problems and you
3245.9s
want people who are low ego that are team players that you know have complimentary skills to your own perhaps um I always find working in a small team where people know things that I don't know and where maybe I have some skills that other people don't have as much of you know is super fun because you're
3248.4s
team players that you know have
3250.8s
complimentary skills to your own perhaps
3253.8s
um I always find working in a small team
3256.6s
where people know things that I don't
3258.6s
know and where maybe I have some skills
3260.6s
that other people don't have as much of
3262.7s
you know is super fun because you're
3264.8s
collectively building something or working on something that none of you could maybe do individually. ually, but in the process of working on that, you actually gain a lot of new knowledge and new skills uh for yourself and so do they. And you you kind of want to view your engineering or research career as
3266.5s
working on something that none of you
3268.4s
could maybe do individually. ually, but
3271.0s
in the process of working on that, you
3274.1s
actually gain a lot of new knowledge and
3275.8s
new skills uh for yourself and so do
3279.0s
they. And you you kind of want to view
3281.6s
your engineering or research career as
3284.8s
you have an amazing tool belt of techniques. And you always want to be adding new tools to that tool belt because you never know when you might come across a problem where you need these four specialized tools rather than these three. And adding more tools makes it more likely that the problems you you
3286.5s
techniques. And you always want to be
3288.5s
adding new tools to that tool belt
3290.7s
because you never know when you might
3293.1s
come across a problem where you need
3295.1s
these four specialized tools rather than
3297.5s
these three. And adding more tools makes
3300.2s
it more likely that the problems you you
3302.7s
encounter in the future will be solvable by you. Now, one last thing. I'm pretty sure someone in this room or multiple people will eventually build something as consequential as you've done with map reduce, TPU, distillation, etc., etc. What problem do you hope they would be working on? Oh yeah. I mean I I think there's a lot of
3304.7s
by you.
3307.3s
Now, one last thing. I'm pretty sure
3309.8s
someone in this room or multiple people
3313.3s
will eventually build something as
3315.5s
consequential as you've done with map
3318.3s
reduce, TPU,
3321.3s
distillation, etc., etc. What problem do
3324.4s
you hope they would be working on? Oh
3328.0s
yeah. I mean I I think there's a lot of
3330.6s
interesting problems in the world and I'll just rattle off a few. This is not exhaustive because the world is a very big place and full of problems. You know I'm particularly excited about new approaches to hardware. You know we that thought experiment there was kind of you know a you know a indication of that or much
3332.3s
I'll just rattle off a few. This is not
3334.1s
exhaustive because the world is a very
3336.3s
big place and full of problems. You know
3339.0s
I'm particularly excited about new
3341.4s
approaches to hardware. You know we that
3343.1s
thought experiment there was kind of you
3345.3s
know a
3347.3s
you know a indication of that or much
3350.6s
more efficient inference hardware. You know, I think there are radically different kinds of algorithms for machine learning that might be much much more data efficient than the approaches we're using today. If you think about our large scale models today, they probably see a thousand times as much data as a human does by the age of 18.
3352.7s
know, I think there are radically
3354.2s
different kinds of algorithms for
3356.6s
machine learning that might be much much
3358.2s
more data efficient than the approaches
3360.2s
we're using today. If you think about
3362.0s
our large scale models today, they
3364.2s
probably see a thousand times as much
3365.9s
data as a human does by the age of 18.
3369.2s
Yet, the human by the age of 18 is better in a lot of things and, you know, on par uh with those frontier models that have seen way more data. So could you come up with much more data efficient systems that can learn continuously learn from their own actions? Uh continual learning is a
3371.8s
better in a lot of things and, you know,
3373.5s
on par uh with those frontier models
3376.1s
that have seen way more data. So could
3377.9s
you come up with much more data
3380.3s
efficient systems that can learn
3382.6s
continuously learn from their own
3384.2s
actions? Uh continual learning is a
3386.8s
really interesting thing. I think multi- aent interactions is an interesting thing. Um you know I think you know creating ways of having better discourse among people in the world uh could be interesting. Are there ways to have much more civil conversations and you know helping people meet other people are all over the world that they should know
3389.0s
aent interactions is an interesting
3391.0s
thing. Um you know I think you know
3395.0s
creating ways of having better discourse
3397.2s
among people in the world uh could be
3399.8s
interesting. Are there ways to have much
3401.8s
more civil conversations and you know
3405.0s
helping people meet other people are all
3406.9s
over the world that they should know
3408.2s
based on their interests. You know these are kind of interesting things. I think there there's lots of cool things in the world and we should all go and strive to make even cooler things occur. That sounds wonderful. Thank you so much Jeff Dane. That's all we have today. Appreciate it. Thank you all.
3410.4s
are kind of interesting things. I think
3412.2s
there there's lots of cool things in the
3414.4s
world and we should all go and strive to
3417.0s
make even cooler things occur.
3419.4s
That sounds wonderful. Thank you so much
3421.2s
Jeff Dane. That's all we have today.
3423.4s
Appreciate it.
3425.2s
Thank you all.