All Videos
Post-Training Is How You Keep Your Taste | Fireworks CEO Lin Qiao

Post-Training Is How You Keep Your Taste | Fireworks CEO Lin Qiao

Read the full transcript of "Post-Training Is How You Keep Your Taste | Fireworks CEO Lin Qiao" by Sequoia Capital. Practice English listening and reading wi...

Channel: Sequoia Capital Duration: 28 min Sentences: 83
When should a company move from prompting to post-training its own models? Fireworks AI co-founder and CEO Lin Qiao lays out the full progression at Sequoia Capital’s Own Your Intelligence event, from prompting and RAG to supervised fine-tuning, preference tuning, reinforcement learning, and distillation. And she explains which technique solves which problem. Lin also covers the pitfalls teams hit along the way: prioritizing data quantity over quality, relying on vibes instead of systematic evals, sloppy RL environments, and reward hacking. She shares how companies like Cursor have used post-training to compete at frontier quality, and why the right moment to start is after product-market fit, when production data becomes your fuel and owning your stack can cut serving costs 5–10x. 00:00 Introduction 00:37 What Fireworks sees across thousands of AI applications 02:47 Off-the-shelf APIs and the problem of keeping your taste 03:58 What "owning your intelligence" actually means 05:43 The progression: prompting → RAG → SFT → Preferences → RL 07:20 Why this mirrors how humans learn 09:03 Matching the technique to the problem you actually have 10:46 Where teams get stuck: data quality and vibe evals 12:28 Reward hacking: the model that wrote zero lines of code 13:59 Training-to-serving alignment (and why quality silently drops) 15:55 Post-training in healthcare and security 17:31 From coding to every co-work domain 19:35 Incumbents, cost burden, and not scaling into bankruptcy 21:26 How much control do you want? 23:24 Q&A: What makes a good reward signal 25:00 Q&A: When to start thinking about post-training
Watch original video on YouTube →
Start Learning with Interactive Transcript

Full Transcript

2.5s close friend and collaborator of Brendan's. Um Lynn, uh actually show of hands, who here who here does post training? Okay, good amount of the room. And who here uses Firework? Okay, good amount of the room, too. So, Lynn, you have you have some friendlies in the audience. Um and so for this next talk, what we're
3.8s Brendan's. Um Lynn, uh
6.8s actually show of hands, who here who
8.2s here does post training?
10.4s Okay, good amount of the room. And who
11.9s here uses Firework? Okay, good amount of
14.3s the room, too. So, Lynn, you have you
15.7s have some friendlies in the audience.
17.8s Um and so for this next talk, what we're
19.9s going to do is we're going to focus on post training. Um Lynn, si- similar setup to to Brendan's talk. We're going to do 15 minutes on how to approach the problem of post training, um and then 15 minutes or so for QA. Uh and Lynn, really delighted to have you here. Thank you for joining us.
22.1s post training.
23.6s Um Lynn, si- similar setup to to
26.3s Brendan's talk. We're going to do 15
27.7s minutes on how to approach the problem
29.7s of post training, um and then 15 minutes
32.4s or so for QA. Uh and Lynn, really
35.5s delighted to have you here. Thank you
36.6s for joining us.
37.3s Thanks for having me. Uh hi everyone, good morning. I'm Lynn, I'm CEO and co-founder of Firework. So, as we bring up the slide, today I'm going to do a little bit uh deep dive on post training and why post training could be very relevant to uh you building your own business. So, first of all, a little
39.2s good morning. I'm Lynn, I'm CEO and
41.2s co-founder of Firework. So, as we bring
44.0s up the slide, today I'm going to
47.0s do a little bit uh deep dive on post
49.2s training and why post training could be
50.8s very relevant to uh you building your
52.9s own business. So, first of all, a little
54.7s bit like prior, I think um this year we have seen So, first of all, I'm a Firework weather specializing intelligence platform. There are tons and tons of application built on top of us, all the way from startups to digital native and enterprises. So, we get uh um the fun part of my job is we get to
56.4s I think um this year we have seen So,
59.3s first of all, I'm a Firework weather
61.5s specializing intelligence platform.
63.2s There are tons and tons of application
65.1s built on top of us, all the way from
67.2s startups to digital native and
69.0s enterprises. So, we get uh
71.2s um the fun part of my job is we get to
72.8s see a lot of patterns. Um what are the innovation people are uh building on top of us and uh and how they are what are challenges that are uh they're facing and what are trend um people being developers are um building on top of us. So, uh one of the things, especially in the past 1 year,
74.5s Um what are the innovation people are uh
76.9s building on top of us and uh and how
79.9s they are what are challenges that are uh
82.0s they're facing and what are trend um
84.6s people being developers are um
87.2s building on top of us. So, uh one of the
89.5s things, especially in the past 1 year,
91.5s software development and application development has been somewhat disrupted because it goes fine in the past, um you have good idea, you want to implement it, scaling production, it requires a team of tens of very strong product engineers, PMs working together, multiple quarters to deliver that. And right now, with one person, a few weeks,
93.7s development has been somewhat disrupted
96.3s because it goes fine in the past, um you
100.4s have good idea, you want to implement
102.5s it, scaling production, it requires a
105.0s team of tens of very strong product
107.1s engineers, PMs working together,
109.2s multiple quarters to deliver that. And
111.4s right now, with one person, a few weeks,
114.8s uh without understanding how to write a single line of code, you can do that. So, that collapsing of resource required both in terms of timeline and deep expertise is shifting how how competitive the application space is, and it's shifting people from thinking about building on top of off-the-shelf off-the-shelf black box API to build a
117.1s single line of code, you can do that.
119.1s So, that collapsing of resource required
122.2s both in terms of timeline and deep
124.7s expertise is shifting how how
128.1s competitive the application space is,
130.6s and it's shifting people from thinking
133.2s about building on top of off-the-shelf
135.4s off-the-shelf black box API to build a
138.1s much deeper mode. So, that they can build a much durable business. As also you have seen lot of discussion, especially in the past 1 week, between open model and the closed model and all the rallying and support across the open the rallying and support across the open model. The depth of that
140.2s build a much durable business.
143.5s As also you have seen lot of discussion,
145.9s especially in the past 1 week, between
148.1s open model and the closed model and all
151.2s the rallying and support across the open
153.4s the rallying and support across the open model.
154.4s The depth of that
157.1s that alliance is because we believe I mean the industry is much deeper because look at the whole entire industry, there are so many companies, right? There are so many company, all of you are building your own company. Every single company exists for a reason because they focus on solving a unique problem in a special way.
160.6s mean the industry
162.4s is much deeper because
164.0s look at the whole entire industry, there
165.3s are so many companies, right? There are
167.2s so many company, all of you are building
168.6s your own company.
171.0s Every single company exists for a reason
173.1s because they focus on solving a unique
176.2s problem in a special way.
178.5s And that means they carry their own judgment, taste, and determination, conviction into that product, and that's why company exists. Today, if you build on top of off-the-shelf API, then you really need to think about how you keep that special taste judgment and unique part forward. And we believe one approach for every company to build
181.9s judgment, taste,
184.0s and determination, conviction into that
186.4s product, and that's why company exists.
189.4s Today, if you build on top of
191.2s off-the-shelf API, then
194.1s you really need to think about how you
196.1s keep that special taste judgment
199.4s and unique part forward. And we believe
202.8s one approach for every company to build
205.7s a build durable business is to actually bake your judgment, taste, and customer deep understanding into the intelligence you build on top of instead of just a off-the-shelf API. So, that's kind of a little bit context of what's happening in industry, what we are seeing, and why post training could be very very relevant to you.
208.7s bake your judgment, taste, and customer
211.2s deep understanding into the intelligence
213.4s you build on top of instead of just a
216.4s off-the-shelf API. So, that's kind of a
218.5s little bit context of
220.5s what's happening in industry, what we
222.2s are seeing, and why post training could
224.2s be very very relevant to you.
236.0s owning intelligence not rent. And what does owning your own intelligence mean? It actually means many things. So, first of all, it start from data. Intelligence is derivative of data. Intelligence is derivative of data. And obviously the the foundation labs, the all the foundation model we're using are building on top of the public data and
238.8s And what does owning your own
240.0s intelligence mean? It actually means
241.8s many things. So, first of all, it start
244.3s from data.
245.8s Intelligence is derivative of data.
249.4s Intelligence is derivative of data. And
250.7s obviously the the foundation labs, the
254.7s all the foundation model we're using are
256.7s building on top of the public data and
259.8s the label data that is has solved common tasks, but all of you are solving a specific task. That's why you you're building a business, you're building a company, and be able to curate production data with high quality and even generate synthetic data to enrich your production data is one step of owning your own intelligence.
262.5s is has solved common tasks, but all of
266.0s you are solving a specific task. That's
268.2s why you you're building a business,
270.2s you're building a company, and be able
272.9s to curate production data with high
275.8s quality and even generate synthetic data
278.4s to enrich your production data is one
281.2s step of owning your own intelligence.
284.5s And then after you have data, you will start to kind of use that data turn into a model, build on top of existing model and own the weights. And there are a collection of techniques you can use to get there. Those techniques are tailored to solve different kind of problem you could possibly have. And those techniques can
286.7s start to kind of use that data turn into
288.6s a model,
289.8s build on top of existing model and own
292.0s the weights. And there are a collection
294.4s of techniques you can use to get there.
297.1s Those techniques are tailored to solve
299.3s different kind of problem you could
300.8s possibly have. And those techniques can
303.2s also interoperate with each other for you to build a reach your final goal. And then after you have a great model belong to yourself solving your specific problem really well, and then you need to work on serving it. You probably first will do a some AB testing to make sure it really move the
305.7s you to build a reach your final goal.
308.4s And then after you have a great model
310.3s belong to yourself solving your specific
312.3s problem really well, and then you need
313.8s to work on serving it.
316.6s You probably first will do a some AB
318.2s testing to make sure it really move the
320.2s needle for your product metrics, and then goes back in in this loop. into post training right away. And and there are different phases you will go into. First, prompt everyone start from prompt, use the model as is. It's few shots. If quickly you can use that to test your ideas. And then you use rag to ground
323.1s and then goes back in in this loop.
331.1s into post training right away.
333.7s And and there are different phases you
336.3s will go into.
337.7s First, prompt everyone start from
339.9s prompt, use the model as is. It's few
342.6s shots. If quickly you can use that to
344.5s test your ideas.
346.3s And then you use rag to ground
349.0s And then you use rag to ground um, the usage of AI into your own data and you can start to do a lot of contest engineering from that. Those are from minutes of interaction to hours of interaction. And then you first progress into hey, I have some data. I want to
349.9s the usage of AI into your own data and
353.0s you can start to do a lot of contest
354.7s engineering from that. Those are from
357.0s minutes of interaction to hours of
359.1s interaction. And then you first progress
361.7s into hey, I have some data. I want to
363.6s see how my data is going to reflect and make uh, the model work better with my product. So you will start to do supervised fine-tuning that will take you a few hours to um, hey, um, the the model I want to reflect in personalized taste. It is it is very unique uh,
365.6s make uh, the model work better with my
368.8s product. So you will start to do
370.5s supervised fine-tuning that will take
372.4s you a few hours to um, hey, um, the the
376.9s model I want to reflect in personalized
379.6s taste. It is it is very unique uh,
382.6s choice of my products. So so therefore you want to start to use preferences uh, information where you collect from user interaction and uh, whether thumbs-up, thumbs-down, a lot of those kind of information and uh, help the model learn your product taste. Um, and and finally, you want to build a model towards understand carrying your expertise in
385.8s you want to start to use preferences uh,
389.2s information where you collect from user
390.8s interaction and uh, whether thumbs-up,
394.0s thumbs-down, a lot of those kind of
395.2s information and uh, help the model learn
398.2s your product taste. Um, and and finally,
401.8s you want to build a model towards
404.0s understand carrying your expertise in
406.2s that domain. Whether that expertise is that domain. Whether that expertise is across uh, legal, finance, healthcare, customer support, recruiting, marketing, sales, you name it. Even in one industry vertical, there are so many subdomains. So all of that is unique and special towards the product you're building. So usually um, you probably heard a lot of
409.2s that domain. Whether that expertise is across
410.4s uh, legal, finance, healthcare, customer
413.4s support, recruiting, marketing, sales,
416.4s you name it. Even in one industry
419.1s vertical, there are so many subdomains.
421.7s So all of that is unique and special
424.4s towards the product you're building. So
426.2s usually um, you probably heard a lot of
428.9s reinforcement learning and that is to to actually build towards a specialty. Um, so this progression is very similar to how we human being learn knowledge over time. Uh, for example, uh, we actually run uh, learn a lot of knowledge by reading uh, reading literature, right? In the literature it will say, hey, what is correct, what is
431.1s actually build towards a specialty.
433.8s Um, so this progression is very similar
436.2s to how we human being learn knowledge
438.8s over time. Uh, for example,
441.8s uh, we actually run uh, learn a lot of
444.4s knowledge by reading uh, reading
446.6s literature, right? In the literature it
448.4s will say, hey, what is correct, what is
450.0s it not correct? So this is a very similar to supervised fine-tuning. Um, and uh, as we grow, we develop our own taste and judgment of how we want to conduct the specific way of we want to approach a problem, and that is preference or DPO. Um and over course of time, we learn to
451.6s similar to supervised fine-tuning. Um,
454.6s and uh, as we grow, we develop our own
457.5s taste and judgment of how we want to
459.8s conduct the
461.2s specific way of we want to approach a
463.7s problem, and that is preference or DPO.
467.5s Um and over course of time, we learn to
469.9s be a doing a really good job at certain area, whether um hey, I want to be accounting accountant, and I would really know how to kind of um build into the financial financial data, or I want to be a specific set of um I want to be a dentist, and you learn the details of how to operate
472.2s area, whether um hey, I want to be
474.6s accounting accountant, and I would
476.5s really know how to kind of um
479.3s build into the financial financial data,
482.1s or I want to be a specific set of um I
484.6s want to be a dentist, and you learn the
486.5s details of how to operate
488.6s um with dentistry. So, all this is very similar exactly um to how we human being acquire knowledges. different techniques are there to solve different kind of problems. Um for example, if the model doesn't know um the fact, and the fact actually is very dynamic, the facts of your data is uh in in your product is very
491.5s similar exactly um to how we human being
494.3s acquire knowledges.
500.5s different techniques are there to solve
502.5s different kind of problems.
504.3s Um for example, if the model doesn't
507.0s know um the fact, and the fact actually
509.6s is very dynamic, the facts of your data
512.4s is uh in in your product is very
514.6s dynamic, and then typically you use rag to solve that problem. Um however, if your model output all the Um however, if your model output all the behavior uh or um or the structure is off, then you you give you curate data and do supervised fine-tuning to correct that. Um if your model's answer, the quality
516.9s to solve that problem.
518.5s Um however, if your model output all the
522.0s Um however, if your model output all the behavior
523.7s uh or um or the structure is off, then
528.0s you you give you curate data and do
531.0s supervised fine-tuning to correct that.
533.7s Um if your model's answer, the quality
537.3s is is personal or kind of is specific to the taste of your product, um and then you use preference and um and then you use preference and tuning. Or the model is quite weak on the special problem you're trying to solve, then you use RL uh reinforcement solve, then you use RL uh reinforcement learning.
541.0s the taste of your product,
543.1s um and then you use preference and
545.3s um and then you use preference and tuning.
546.4s Or the model is quite weak
549.2s on the special problem you're trying to
551.2s solve, then you use RL uh reinforcement
554.4s solve, then you use RL uh reinforcement learning.
555.6s And the end result of the model is too slow or too expensive to for you to serve in production, and then you use distillation to um go um let the teacher teach the student much smaller student model, so it can be more much more performing or economical. Obviously, distillation also means different things for LLM, VLM, or uh
558.7s slow or too expensive to for you to
561.4s serve in production, and then you use
563.4s distillation to um go
566.5s um let the teacher teach the student
568.6s much smaller student model, so it can be
570.8s more much more performing or economical.
573.7s Obviously, distillation also means
575.4s different things for LLM, VLM, or uh
578.2s image generation models. Um if you are interested, we can talk more about that.
580.8s interested, we can talk more about that.
587.8s teams start to hey, I want to kind of try those technology, and um and they may not be happy uh of the experience because they can burn money um and the time, but may not reach their uh ideal results. So, so here are the areas that you can possibly uh feel frustrated. Uh for example, um
589.7s try those technology, and um and they
592.8s may not be happy uh
595.0s of the experience because they can burn
597.4s money um and the time, but may not reach
600.5s their uh ideal results. So, so here are
602.8s the areas that you can possibly uh feel
605.6s frustrated. Uh for example, um
609.0s when it comes to data, and data is the essence of tuning, it the quantity is not the most important. Actually, quality is the most important. But, sometimes just by throwing uh tons and tons of data into a training process may not leading to a great outcome. So, you really need to control the data quality
610.7s essence of tuning, it the quantity is
614.5s not the most important. Actually,
616.0s quality is the most important. But,
617.7s sometimes just by throwing uh tons and
620.3s tons of data into a training process may
622.9s not leading to a great outcome. So, you
625.7s really need to control the data quality
627.9s and have And usually, who's the best judge of data quality? It's actually your product team. Um so, that's where we see the convergence of um before we see the convergence of um before GenAI, product team and the research team or ML team, they're separate organizations. They're separate team. And they work hand-in-hand to make things happen. And
630.5s judge of data quality? It's actually
632.3s your product team. Um so, that's where
634.4s we see the convergence of um before
636.9s we see the convergence of um before GenAI,
638.2s product team and the research team or ML
640.5s team, they're separate organizations.
642.2s They're separate team. And they work
643.9s hand-in-hand to make things happen. And
646.2s a lot of time these days, when um people a lot of time these days, when um people post-train uh GenAI models, we see kind of the convergence of the product team needs to make judgment call of the data quality and start to be deeply involved in the process. So, uh they can ensure uh the best result.
649.0s a lot of time these days, when um people post-train
650.2s uh GenAI models, we see kind of the
652.3s convergence of the product team needs to
654.8s make judgment call of the data quality
656.9s and start to be deeply involved in the
658.6s process. So, uh they can ensure uh the
661.4s best result.
662.9s Um and the second is you going to have evals. Um I know all of you are very busy, and uh part into launch quickly, and a lot of evals is a vibe evaling. and a lot of evals is a vibe evaling. laughter And the founders uh really look at the
665.3s evals. Um I know all of you are very
668.3s busy, and uh part into launch quickly,
671.3s and a lot of evals is a vibe evaling.
674.0s and a lot of evals is a vibe evaling. laughter
674.3s And the founders uh really look at the
676.5s result and feel hey, it is it right or not? Actually, this is judgment. This is you put your judgment of uh of end result and decide whether it's good or not. And that judgment you convert into systemic evaluation. So, this is no different from traditional software development where you have your unit test, integration test to ensure
679.2s not? Actually, this is judgment. This is
681.4s you put your judgment of uh of end
684.3s result and decide whether it's good or
686.1s not. And that judgment you convert into
689.1s systemic evaluation. So, this is no
691.6s different from traditional software
693.4s development where you have your unit
695.5s test, integration test to ensure
697.0s quality. Similarly, if you want to if you think about doing post-training and then have a way to build the eval and build your own judgment into a repeatable process, it's extremely important. Uh and then there could be when you do RL, there could be sloppy RL environment uh where you build a simulation and the
699.5s if you want to if you think about doing
701.3s post-training and then have a way to
703.0s build the eval and build your own
704.5s judgment into a repeatable process, it's
707.1s extremely important.
708.6s Uh and then there could be when you do
710.6s RL, there could be sloppy RL environment
714.2s uh where you build a simulation and the
716.2s simulation is not really reflecting the reality and then the model can heel climb on on kind of the a bad um simulation and I mean it it it will it will also could possibly do reward will also could possibly do reward hacking um and all kind of weird stuff. So, a fun story about reward hacking is
718.3s reality and then the model can heel
721.4s climb on on kind of the a bad um
725.0s simulation and I mean it it it will it
727.1s will also could possibly do reward
729.1s will also could possibly do reward hacking
730.2s um and all kind of weird stuff. So, a
732.5s fun story about reward hacking is
735.0s um we have been asking a model to uh to generate to do this is coding example uh to create um to to generate code that minimize uh the the error compilation error. Okay. So, guess what the model did? The model generate zero line of code. Okay, there's no compilation error, but that's absolutely not what you want. So,
737.8s generate to do this is coding example uh
740.8s to create um
742.8s to to generate code that minimize
746.0s uh the the error compilation error.
748.2s Okay. So, guess what the model did? The
750.6s model generate zero line of code.
753.5s Okay, there's no compilation error, but
755.0s that's absolutely not what you want. So,
756.9s this is a kind of one example of reward hacking, it's very common because model was very smart, it'll try in all different way to uh to get um your goal, but it may not be what you want. So, uh pay attention to all these details and uh and try not to let the model outsmart you. Um obviously,
758.6s hacking, it's very common because model
760.6s was very smart, it'll try in all
762.1s different way to
763.6s uh to get um your goal, but it may not
766.0s be what you want. So, uh pay attention
768.1s to all these details and uh and try not
771.0s to let the model outsmart you.
773.1s Um obviously,
775.1s between um your experimentation, so think about um your development process as hey, you're actually doing a lot of experimentation from training the model to uh training the model is not the end of of your experiment. The judge is the final judge is whether product matches is moving or not. So, you need to bring the final model into
777.2s think about um your development process
780.2s as hey, you're actually doing a lot of
782.1s experimentation from training the model
784.4s to uh training the model is not the end
787.2s of of your experiment.
789.0s The judge is the final judge is whether
791.9s product matches is moving or not. So,
794.5s you need to bring the final model into
796.8s into serving tier and do AB testing. Um and that transition is extremely important because um the quality can drop if you move from one training stack to a serving stack without aligning these two. Because think about the model is tons and tons of calculation, math, and the matrix multiplication. Um and the way to do multi matrix
799.3s and that transition is extremely
800.8s important because um the quality can
803.4s drop if you move from one training stack
806.5s to a serving stack without aligning
808.8s these two. Because think about the model
812.0s is tons and tons of calculation, math,
816.2s and the matrix multiplication. Um and
819.2s the way to do multi matrix
821.0s multiplication and calculation, if you use different library, different use different library, different numerics, um and different optimization, it will lead different results. And therefore, lead different results. And therefore, the the end result of a training may not be replicated or even uh you lose precision during serving. So, that alignment is very important. I can give you more
823.3s use different library, different
824.9s use different library, different numerics,
826.5s um and different optimization, it will
828.0s lead different results. And therefore,
830.0s lead different results. And therefore, the
831.1s the end result of a training may not be
833.2s replicated or even uh you lose precision
836.5s during serving. So, that alignment is
839.0s very important. I can give you more
840.0s examples of that. Um and there are a few other um challenges you will run into and happy to talk with you more about in details to talk with you more about in details offline. So, there have been many pioneers uh we worked with post-training. I would say Cursor is one of the few. They have
842.1s Um and there are a few other um
844.1s challenges you will run into and happy
846.8s to talk with you more about in details
849.7s to talk with you more about in details offline.
851.5s So, there have been many pioneers
854.8s uh we worked with post-training. I would
857.8s say Cursor is one of the few. They have
860.4s started onboarding getting onto this train from the beginning of last year. Uh there are multiple reasons. One is they really want to control their destiny of the supply of the model. And destiny of the supply of the model. And obviously, um they have a lot of they have a lot of user engagement. They understand um how
863.2s train from the beginning of last year.
865.8s Uh there are multiple reasons. One is
867.8s they really want to control their
869.1s destiny of the supply of the model. And
872.1s destiny of the supply of the model. And obviously,
873.5s um they have a lot of they have a lot of
876.0s user engagement. They understand um how
879.3s what is customer preferences, a lot of data. So, that becomes the beginning of their journey. And they do um pretty deep mid-training to post-training. And you have heard them announce um continuously newer models every few month. Um and uh and the result is great. Um their inspiration is to compete at the frontier quality. Uh very
881.2s data. So, that becomes the beginning of
883.9s their journey. And they do um pretty
886.3s deep mid-training to post-training. And
888.7s you have heard them announce um
890.4s continuously newer models every few
892.7s month. Um and uh and the result is
896.0s great. Um their inspiration is to
899.1s compete at the frontier quality. Uh very
901.8s bold aspiration. And they're getting there. So, very impressive result from composer two, composer 2.5, um on par with um They're always trying to be on par or They're always trying to be on par or beat um the frontier labs quality at the same time. So, um they are kind of one of the
903.8s there. So, very impressive result from
906.8s composer two, composer 2.5, um on par
910.0s with um
911.3s They're always trying to be on par or
913.5s They're always trying to be on par or beat
914.9s um the frontier labs quality at the same
917.3s time. So, um they are kind of one of the
920.1s examples, a very vibrant example, is completely doable as long as you stay focused and have the right tool and cursive build on top of us. Um there's another there huge huge variety of examples from healthcare. For example, Doximity is one example where they do clinic AI where they basically let doctors ask deep medical questions
923.1s completely doable as long as you stay
925.4s focused and have the right tool and
927.3s cursive build on top of us.
929.1s Um there's another there huge huge
932.0s variety of examples from healthcare. For
935.4s example, Doximity is one example where
938.7s they do clinic AI where
941.3s they basically let doctors ask deep
944.0s medical questions
945.7s matching symptoms to medication and side effects and have well-rounded research around everything in medical space. They actually also train on Fireworks and they top a very important benchmark which is Stanford Harvard clinic safety which is Stanford Harvard clinic safety benchmark. And we are very proud of that result and they they keep working on this training
948.8s effects and have well-rounded research
951.8s around everything
953.8s in medical space.
955.6s They actually also train on Fireworks
958.4s and they top a very important benchmark
961.2s which is Stanford Harvard clinic safety
963.1s which is Stanford Harvard clinic safety benchmark.
964.7s And we are very proud of that result and
967.8s they they keep working on this training
970.7s loop. Um Factory, that's another Sequoia loop. Um Factory, that's another Sequoia company. They build on top of Fireworks and especially focus on security part of the especially focus on security part of the coding. That is a very hard topic because security is not high tolerance. You need to get it high tolerance. You need to get it right.
974.5s loop. Um Factory, that's another Sequoia company.
975.8s They build on top of Fireworks and
977.9s especially focus on security part of the
980.1s especially focus on security part of the coding.
981.0s That is a very hard topic because
982.8s security is not
985.0s high tolerance. You need to get it
986.6s high tolerance. You need to get it right.
987.8s They tune a model that also top the kind of benchmarking security top the kind of benchmarking security area. So you can see all these examples is a demonstration. They solve all the companies are solving a very unique problem in a special way and they have been able to kind of build their own challenges use their own
991.3s top the kind of benchmarking security
993.1s top the kind of benchmarking security area.
994.5s So you can see all these examples is a
996.6s demonstration. They solve all the
998.5s companies are solving a very unique
1000.4s problem in a special way and they have
1003.6s been able to kind of
1005.0s build their own challenges use their own
1006.8s data by building on top of open model. And those are the details. Obviously, as you know, last year 2025 is the year of coding. There Fireworks we support all the coding companies build on top of us. A lot of success in mid-train to post-train very strong models. As you can see those benchmarks are very
1011.1s And those are the details. Obviously, as
1014.3s you know, last year 2025 is the year of
1016.4s coding. There
1018.4s Fireworks we support all the coding
1020.2s companies build on top of us. A lot of
1022.4s success in
1024.7s mid-train to post-train
1026.4s very strong models.
1028.6s As you can see those benchmarks are very
1030.3s impressive and all these coding companies are inspired to get on par or beat Frontier Labs in the coding space, but it's not coding. And we start This is the year we start to see very interesting development in all different kind of co-work space from general purpose co-work to specialized domain specific co-work as I mentioned.
1032.5s companies are inspired to
1034.6s get on par or beat Frontier Labs in the
1037.5s coding space, but it's not coding. And
1040.8s we start This is the year we start to
1043.1s see very interesting development in all
1045.7s different kind of co-work space from
1047.9s general purpose co-work to specialized
1050.4s domain specific co-work as I mentioned.
1053.0s There are like legal, finance, marketing, recruiting, sales, customer support co-work space. They are all starting to post train and own their own intelligence power their product. own intelligence power their product. Here, Jen Spark is one of the generic co-work application and they build deep research for professionals and slide generation. As you see,
1055.8s finance, marketing, recruiting, sales,
1059.1s customer support co-work space. They are
1061.0s all starting to post train and own their
1064.1s own intelligence power their product.
1066.6s own intelligence power their product. Here,
1068.7s Jen Spark is one of the generic co-work
1073.6s application and they build deep research
1076.0s for professionals and slide generation.
1078.8s As you see,
1080.2s this they compete with a frontier model and it's slightly better, but the cost is significantly lower. So here we're talking about five to 10 10 times cost reduction. So as a startup, you can once you hit a problem like if it you can scale quickly into a doable business. And I guess another health care example and health
1084.3s and it's slightly better, but the cost
1087.4s is significantly lower. So here we're
1089.8s talking about five to 10
1091.5s 10 times cost reduction. So
1094.4s as a startup, you can once you hit a
1096.7s problem like if it you can scale quickly
1099.6s into a doable business. And I guess
1102.6s another health care example and health
1105.0s care data or use case is not well care data or use case is not well captured captured in in frontier labs model. So you can see the quality of tuned model is significantly better than the state of art closed model.
1107.2s care data or use case is not well captured
1108.4s captured in
1109.6s in frontier labs model. So you can see
1112.7s the quality of tuned model is
1115.5s significantly better than the state of
1118.1s art closed model.
1124.7s developers build on top of fireworks to do post training. So what what are these do post training. So what what are these types? So we have seen many frontier agent So we have seen many frontier agent builder. So frontier agent builder, they usually build bespoke customized harness. They don't use common harnesses. And to
1127.5s do post training. So what what are these
1129.4s do post training. So what what are these types?
1130.5s So we have seen many frontier agent
1133.5s So we have seen many frontier agent builder.
1134.7s So frontier agent builder, they usually
1137.0s build bespoke customized harness.
1140.4s They don't use common harnesses. And to
1143.7s co-optimize their harness with their systems and all different kind of tools they want to build on top of often time requires post training. Uh so, we have we have seen a lot of repeatable success on doing that. repeatable success on doing that. And obviously, we have seen a lot of big
1146.6s systems and all different kind of tools
1148.6s they want to build on top of
1151.5s often time requires post training. Uh
1154.0s so, we have we have seen a lot of
1156.4s repeatable success on doing that.
1159.5s repeatable success on doing that. And
1160.8s obviously, we have seen a lot of big
1162.5s obviously, we have seen a lot of big companies put to post training because interestingly, those incumbents have huge amount of traffic. For them to deploy AI features across the board to all the all their customers is a huge cost burden. Huge cost burden. Uh so, if we say we we need to be careful not
1163.9s put to post training because
1165.1s interestingly, those incumbents have
1167.0s huge amount of traffic.
1169.2s For them to deploy AI features across
1171.6s the board
1172.9s to all the all their customers is a huge
1176.5s cost burden. Huge cost burden. Uh so, if
1180.0s we say we we need to be careful not
1183.1s scale into bankruptcy, it's not just for startups. Uh it's actually for incumbents. They literally their CFO is blocking their AI feature launch because of the cost. And the post training is way to remediate that. Uh and we have seen a lot of specialized model operator, uh those could be uh new
1185.4s startups. Uh it's actually for
1187.3s incumbents. They literally their CFO is
1190.0s blocking their AI feature launch because
1192.8s of the cost. And the post training is
1194.6s way to remediate that.
1196.7s Uh and we have seen a lot of specialized
1199.6s model operator, uh those could be uh new
1203.1s model operator, uh those could be uh new uh new labs and a lot of kind of cutting-edge model develops also um do significant post training. So, uh this is a quote we have heard repeatedly uh that after part of market fit, post training becomes the vehicle to for many, many companies to build specialized intelligence. We do believe
1203.9s new labs and a lot of kind of
1206.0s cutting-edge model develops also um do
1209.6s significant post training.
1215.4s So, uh this is a quote we have heard repeatedly
1216.9s uh that after part of market fit, post
1220.0s training becomes the vehicle to for
1222.4s many, many companies to build
1223.6s specialized intelligence. We do believe
1225.6s we do see the trend that in the future there could be millions of specialized models. One application per use case. It it it it will be millions. That's how uh I believe in that. Um And obviously, there are Um And obviously, there are um we have seen kind of different way to
1227.9s there could be millions of specialized
1230.9s models. One application per use case. It
1234.1s it it it will be millions. That's how
1236.9s uh I believe in that.
1239.2s Um And obviously, there are
1242.6s Um And obviously, there are um
1243.5s we have seen kind of different way to
1245.4s convert the data into into model. Um and the the signal from the reward and how to build those reward function is very important um part of the story. Um so, our observation is Um so, our observation is no, start early, start early, and the start to do experimentation a lot more to do experimentation a lot more iteratively
1249.9s the the signal from the reward and how
1252.9s to build those reward function
1255.6s is very important um part of the story.
1258.7s Um so, our observation is
1262.4s Um so, our observation is no,
1263.4s start early, start early, and the start
1266.8s to do experimentation a lot more
1269.7s to do experimentation a lot more iteratively
1271.6s and they get hands-on experience and we we have seen people on board so quickly. There there's actually not the barrier as many people feel because this this whole reward engineering is very similar to software engineering. The the the logical reasoning and mindset is very very similar. So so yeah so get your hands dirty and and start to
1274.3s we have seen people on board so quickly.
1277.1s There there's actually not the barrier
1279.4s as many people feel because this this
1282.6s whole reward engineering is very similar
1285.0s to software engineering.
1286.7s The the the logical reasoning and
1289.0s mindset is very very similar. So
1292.6s so yeah so
1294.1s get your hands dirty and and start to
1297.3s kind of test it out. I think Yeah so those are the kind of key Yeah so those are the kind of key takeaways and uh I think we snorts we we want to acknowledge there are different specialties or different knowledges you have right now in this space. For example, we work with the
1303.6s I think
1306.2s Yeah so those are the kind of key
1307.4s Yeah so those are the kind of key takeaways
1309.0s and uh
1310.8s I think we snorts
1312.6s we we want to acknowledge there are
1316.3s different specialties or different
1319.6s knowledges you have right now in this
1322.3s space. For example, we work with the
1325.4s companies like Cursor and Cognition. They have deep researchers. They want to control every single knob as much as possible to get extreme results. We give them the lowest API. That means we just have RL rollout that they can directly interact with and they fully control the interact with and they fully control the trainer.
1327.4s They have deep researchers. They want to
1329.1s control every single knob as much as
1331.7s possible to get extreme results.
1335.3s We give them the lowest API.
1337.7s That means
1339.3s we just
1340.6s have RL rollout that they can directly
1343.9s interact with and they fully control the
1345.8s interact with and they fully control the trainer.
1347.0s That gives them the best result. But many of the company you have product engineers or machine learning engineers. You have good enough information and knowledge about post training. You want to have some control but you do not want to have the lowest level control. So that's where we have our training SDK
1349.5s But many of the company
1351.5s you have product engineers or machine
1354.1s learning engineers.
1355.6s You have good enough information and
1358.4s knowledge about post training. You want
1360.8s to have some control but you do not want
1362.5s to have the lowest level control. So
1365.8s that's where we have our training SDK
1369.7s iterate very closely with all of you to get to to get you going and we also are happy to kind of deploy our own researchers. That's where we will see, um, help hands-on like training your team to to get on board. So, those are the kind of different way we see different teams want to engage and and
1372.5s get to to get you going and we also
1377.7s are happy to kind of deploy our own
1379.4s researchers. That's where we will see,
1382.0s um, help hands-on like training your
1385.5s team to to get on board. So, those are
1387.8s the kind of different way we see
1389.5s different teams want to engage and and
1392.5s get familiar with the whole entire get familiar with the whole entire process. Cool. Uh, I'm going to pause here and see if you have any questions. I'll jump um People are most most successful if they have data and reward signal already, maybe from their product. What are the what are the best examples of reward
1394.4s get familiar with the whole entire process.
1395.6s Cool. Uh, I'm going to pause here and
1397.8s see if you have any questions.
1399.5s I'll jump um
1401.1s People are most most successful if they
1402.7s have data and reward signal already,
1405.2s maybe from their product. What are the
1406.8s what are the best examples of reward
1408.8s signals you've seen with very successfully trained very quickly? So, um, usually the rewards you can think about rewards actually is code. You write rewards in code. And uh, usually think about rubrics of rewards. Um, and think about uh, you want to grade the result in multiple dimensions. Uh, snorts for example, if you are building a recruiting
1410.6s very successfully trained very quickly?
1413.3s So, um, usually the rewards you can
1415.9s think about rewards actually is code.
1418.7s You write rewards in code. And uh,
1421.2s usually think about rubrics of rewards.
1424.4s Um, and think about uh, you want to
1427.7s grade the result in multiple dimensions.
1431.7s Uh, snorts for example, if you are
1433.7s building a recruiting
1435.5s building a recruiting agent, uh, and then you can think about, Hey, how do we evaluate candidate selection? Um, and a different company, I guarantee you you have different criterias. Uh, for example, we want to grade a for example, we want to grade a candidate um, um, aptitude. Are they really hungry? They do not take
1436.7s uh, and then you can think about, Hey,
1438.2s how do we evaluate candidate selection?
1441.4s Um, and a different company, I guarantee
1443.9s you you have different criterias. Uh,
1446.0s for example, we want to grade a
1447.9s for example, we want to grade a candidate
1449.2s um, um, aptitude.
1451.4s Are they really hungry? They do not take
1453.6s no as an answer. They they will break down walls. And that's one matrix. The second is they're really fast in kind of second is they're really fast in kind of um, building things and making progress. And and so on. So, then you have a blend of score to blend to merge this. So, so
1455.5s down walls. And that's one matrix. The
1458.5s second is they're really fast in kind of
1462.0s second is they're really fast in kind of um,
1462.7s building things and making progress. And
1465.1s and so on. So, then you have a blend of
1467.0s score to blend to merge this. So, so
1470.0s think about reward in different dimensions of rubrics. And then different company have a different way to blend those. And that's that's your um, unique part and secret sauce. Uh, great talk, Lynn. Raj from Vercel. I'm curious about I think you mentioned what are the pitfalls of starting too early, but based on what you're seeing
1472.2s dimensions of rubrics. And then
1474.4s different company have a different way
1475.8s to blend those. And that's that's your
1479.1s um, unique part and secret sauce.
1485.3s Uh, great talk, Lynn. Raj from Vercel.
1488.3s I'm curious about I think you mentioned
1491.0s what are the pitfalls of starting too
1493.2s early, but based on what you're seeing
1495.0s from Vercel, Vercel cognition, or other companies, when do they actually start thinking about post training? Is it when they feel ready? Is it when the cost is now too much? Because the benchmarks where they're beating it, that's a great signal that yes, you can uh signal that yes, you can uh um
1497.5s companies, when do they actually start
1499.7s thinking about post training? Is it when
1502.4s they feel ready? Is it when the cost is
1504.6s now too much? Because the benchmarks
1508.0s where they're beating it, that's a great
1510.1s signal that yes, you can uh
1512.5s signal that yes, you can uh um
1513.4s you know, beat the frontier models, but is that the primary motivator? Like that is that 2 bump worth the cost or investment into this whole framework? And and how do they think about ROI? You know, close models are more expensive, but then post training has some intercept of, you know, money and investment and maintenance going
1515.5s is that the primary motivator? Like that
1517.3s is that 2 bump worth the cost or
1520.0s investment into this whole framework?
1522.2s And and how do they think about ROI? You
1524.6s know, close models are more expensive,
1526.8s but then post training has some
1529.6s intercept of, you know, money and
1531.2s investment and maintenance going
1532.6s forward. So, I'm curious of like, are they approaching when cost becomes a concern or is it when the usage spiking up and they really have to now plan for a year or two in advance? Yeah, this is this is excellent Yeah, this is this is excellent question. Uh so, usually so, think about uh in the AI building
1535.4s they approaching when cost becomes a
1536.8s concern or is it when the usage spiking
1539.2s up and they really have to now plan for
1541.1s a year or two in advance?
1542.6s Yeah, this is this is excellent
1544.0s Yeah, this is this is excellent question.
1545.1s Uh so, usually so, think about uh in the
1548.4s AI building
1549.8s product market fit and the scaling the business as actually two phases. And then in the SaaS time, it's one concept. You hit a product market fit, just scale. I know you got to scale as much as I know you got to scale as much as possible. Um but now we see a bifurcation.
1552.0s business as actually two phases.
1555.1s And then in the SaaS time, it's one
1556.4s concept. You hit a product market fit,
1558.1s just scale.
1559.3s I know you got to scale as much as
1560.6s I know you got to scale as much as possible.
1561.8s Um but now we see a bifurcation.
1565.3s Product market fit doesn't really mean you have a durable scalable business. And typically, we see uh companies deploy the strategy of focus on product market fit first by building on top of um you know, uh Frontier Labs model because you don't need to worry about anything. So, just kind of spend your
1566.8s you have a durable scalable business.
1568.7s And typically, we see uh companies
1572.0s deploy the strategy of focus on product
1573.8s market fit first by building on top of
1577.0s um you know, uh Frontier Labs model
1580.3s because you don't need to worry about
1581.3s anything. So, just kind of spend your
1583.0s money and hit a product market fit. Uh and the other very important thing is only after you hit the product market fit, the data you collect from product surface area are really meaningful. A- and you will get uh also the volume of high quality data start to collect from um product surface area. That
1585.2s Uh and the other very important thing is
1587.2s only after you hit the product market
1588.4s fit, the data you collect from product
1590.2s surface area are really meaningful.
1593.2s A- and you will get uh also the volume
1595.8s of high quality data start to collect
1597.8s from um product surface area. That
1600.2s became the fuel of uh you start to own your intelligence. So, we um we see like that as kind of the very strong that as kind of the very strong indication um because once you hit upon market fit, you really think about start start to scale a business, you think about two
1602.6s your intelligence. So, we um we see like
1606.0s that as kind of the very strong
1608.1s that as kind of the very strong indication
1609.4s um because once you hit upon market fit,
1611.4s you really think about start start to
1613.1s scale a business, you think about two
1614.5s things. One is continue to keep your competitive edge. Two is um build a build durable business, so your um your revenue and your costs are, you know, in a healthy state. Um so, then um post-training becomes a very appealing solution because post-training allow you solution because post-training allow you to to um
1617.5s competitive edge.
1619.2s Two is
1620.9s um build a build durable business, so
1623.0s your um your revenue and your costs are,
1626.9s you know, in a healthy state.
1628.8s Um so, then um
1632.1s post-training becomes a very appealing
1634.1s solution because post-training allow you
1636.9s solution because post-training allow you to
1637.9s to um
1638.5s basic encode codify um your uh unique taste into into a model that no one can uh steal from. Because it's very easy to clone and copy application as is, as I as you all know, right? Uh from coding agent it's very easy. Uh so, from screenshot, boom, generate the same or
1642.4s taste into into a model that no one can
1646.5s uh steal from. Because it's very easy to
1649.0s clone and copy application as is, as I
1652.3s as you all know, right? Uh from coding
1654.4s agent it's very easy. Uh so, from
1656.4s screenshot, boom, generate the same or
1658.7s even better app. Uh so, that's a very scary that application itself is kind of the a moat is is being um reduced. Um and you bake your data uh where reflecting the product engagement and all the deep knowledge into your model is a way to preserve that. And second is you can post-train a model that bring
1660.8s scary that application itself is kind of
1663.6s the a moat is is being um reduced.
1667.8s Um and you bake your data uh where
1670.8s reflecting the product engagement and
1673.6s all the deep knowledge into your model
1675.8s is a way to preserve that. And second is
1678.3s you can post-train a model that bring
1680.6s down the cost five to time 10 times. And and then that means you can support five to 10 times much higher traffic with the same budget. And the your unit of economics of scaling is so much better than you avoid scaling into bankruptcy. So, we see kind of those as as two compelling story behind the reason and
1683.7s and then that means you can support five
1686.1s to 10 times much higher traffic with the
1688.6s same budget. And the your unit of
1691.2s economics of scaling is so much better
1694.0s than you avoid scaling into bankruptcy.
1696.3s So, we see kind of those as as two
1698.3s compelling story behind the reason and
1701.0s compelling story behind the reason and timing. Thank you, Lin. Thanks a lot. Thanks a lot. applause
1702.5s Thank you, Lin.
1703.4s Thanks a lot.
1704.3s Thanks a lot. applause