All Videos
RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor

RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor

Read the full transcript of "RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor" by Sequoia Capital. Practice English lis...

Channel: Sequoia Capital Duration: 26 min Sentences: 82
RL environments have become the hottest topic in AI training data. But what are they, exactly? At Sequoia Capital’s Own Your Intelligence event, Mercor CEO Brendan Foody breaks down the three components: worlds (the messages, docs, and files of a real project), apps (high-fidelity clones of tools like Salesforce and Google Workspace), and tasks (prompts paired with rubric-based verifiers). He walks through a real legal environment built with lawyers from top firms, and shares post-training results showing dramatic gains on domain-specific tasks from modest compute. Brendan also covers why humans remain essential for measuring the frontier, how Mercor prices data, the shift from crowdsourced labeling to expert-built environments, and what's next: ultra-long-horizon tasks and virtual coworkers. He argues that technology once limited to frontier labs is now reaching application companies, and the datasets they build are becoming their moat. 00:00 Introduction 00:47 A short history of the data market: crowdsourcing to agentic data 02:29 What an RL environment is: worlds, apps, tasks 03:57 Why only humans can measure the frontier 05:35 Building verifiers is the hard part 06:44 Walkthrough: a real legal RL environment 08:18 Leaderboards — and what open weights change 09:45 Post-training results on Apex Agents 11:17 Three ways companies buy data 12:49 Q&A: How do you price data? 14:17 Q&A: What "data quality" actually means 16:42 Q&A: The misunderstanding about synthetic data 18:17 Q&A: Why RL environments now — and what comes after 21:20 Q&A: Can you scale rubric generation with models? 23:00 Q&A: RL environments for cyber defense 25:33 Q&A: Build data in-house or partner?
Watch original video on YouTube →
Start Learning with Interactive Transcript

Full Transcript

1.8s 1 to a 2 billion dollar revenue run rate in the last 4 months or so. Um, so this company's off to the races and I think you were just so front and center to how companies are thinking about uh post training their own models, uh building their own intelligence. So, thank you for joining us for this
3.9s in the last 4 months or so.
6.0s Um, so this company's off to the races
8.0s and I think you were just so front and
9.5s center to how companies are thinking
11.1s about uh post training their own models,
13.8s uh building their own intelligence. So,
15.7s thank you for joining us for this
16.8s conversation. Um, format-wise what we're going to do is we've 15 minutes or so of content from Brendan. He's going to talk about uh RL environments in particular, which I think is a, you know, new frontier topic. It'll be fun to fun to explore. And then we're going to leave 15 minutes or so at the end for QA
19.2s going to do is we've 15 minutes or so of
22.1s content from Brendan. He's going to talk
23.8s about uh RL environments in particular,
26.5s which I think is a, you know, new
28.2s frontier topic. It'll be fun to fun to
30.2s explore. And then we're going to leave
32.3s 15 minutes or so at the end for QA
34.8s again. So, uh please keep please keep questions back pocket. I will turn it over to you, Brendan. Sweet. So, I'll be talking about RL environments. Starting out, I figured it's helpful to give a little bit of the background on the history of the data market and how that history ties into Record's origin story. Where things
37.2s questions back pocket. I will turn it
38.8s over to you, Brendan.
40.0s Sweet. So, I'll be talking about RL
41.5s environments. Starting out, I figured
43.8s it's helpful to give a little bit of the
45.8s background on the history of the data
48.4s market and how that history ties into
51.7s Record's origin story. Where things
53.8s really started in 2020 in the era of crowdsourcing data for behavior cloning. So, this was mainly supervised fine-tuning data, inputs and outputs, and RLHF data where you would have a annotator select from a couple of model responses which they preferred. And we were able to make all this progress in fine-tuning GPT-3, making progress
56.3s crowdsourcing data for behavior cloning.
58.8s So, this was mainly supervised
60.2s fine-tuning data, inputs and outputs,
62.6s and RLHF data where you would have a
66.0s annotator select from a couple of model
68.1s responses which they preferred. And we
70.4s were able to make all this progress in
73.3s fine-tuning GPT-3, making progress
76.0s towards ChatGPT and GPT-4 in the crowdsourcing era of agentic data. But, what we saw changing in the market, especially as we uh headed into 2024, was this giant transition away from the low-skilled crowdsourcing era of behavior cloning data and moving towards the agentic era of data. Of how do we find the highest-skilled experts in the
78.5s crowdsourcing era of agentic data. But,
81.8s what we saw changing in the market,
84.5s especially as we uh headed into 2024,
87.7s was this giant transition away from the
90.8s low-skilled crowdsourcing era of
93.6s behavior cloning data and moving towards
96.3s the agentic era of data. Of how do we
98.4s find the highest-skilled experts in the
100.0s world that can work collaboratively in teams to build frontier evals and RL environments for the next generation of models models. All the software engineers, lawyers, doctors, bankers, et cetera that could measure the frontier of intelligence and help to use that to improve model capabilities. And so, Mercor grew up with our first big
102.1s teams to build frontier evals and RL
104.8s environments for the next generation of
106.5s models models. All the software
108.6s engineers, lawyers, doctors, bankers, et
111.0s cetera that could measure the frontier
113.9s of intelligence and help to use that to
118.0s improve model capabilities. And so,
120.0s Mercor grew up with our first big
122.0s project being deep research. I guess the first prominent RL agent scaling up dramatically with all of the frontier labs to become the primary agentic data vendor to all of the leading labs and also all of the leading application layer companies ranging from Harvey, Cera, Cognition to Ramp. And what's been really exciting over the last 12 months
124.4s first prominent RL agent
127.8s scaling up dramatically with all of the
129.6s frontier labs to become the primary
131.8s agentic data vendor to
134.1s all of the leading labs
136.0s and also all of the leading application
137.6s layer companies ranging from Harvey,
139.6s Cera, Cognition to Ramp. And what's been
142.9s really exciting over the last 12 months
146.3s especially is how RLVR within the agentic data paradigm has evolved to also include RL environments with these rich apps and worlds that teach agents how to use all of the tools on our laptops that we use every day. So, I'll be talking about that and of course how this technology that started in the frontier labs is now
149.5s agentic data paradigm has evolved to
152.8s also include RL environments with these
155.3s rich apps and worlds that teach agents
158.8s how to use all of the tools on
161.3s our laptops that we use every day. So,
163.4s I'll be talking about that
165.8s and of course how this technology that
168.8s started in the frontier labs is now
171.9s getting disseminated to the application layer and all of the products that all of you are building as you work on your company. So, high level on what an RL environment is is that it includes three parts. The first part is the worlds. So, this includes all of the messages, slides, docs, sheets, etc.
173.7s layer and all of the products that all
175.6s of you are building
177.1s as you
178.3s work on your company. So, high level on
180.3s what an RL environment is is that it
182.0s includes three parts. The first part is
184.3s the worlds. So, this includes all of the
186.8s messages, slides, docs, sheets, etc.
189.5s that correspond to everything you would have in a real project or company that you're working on. The second part is the apps which is high fidelity clones of popular applications, Salesforce, ServiceNow, Microsoft 365, etc. that agents can interact with via MCP, CLI, or Kua. And then the third part is the tasks where we have prompts and
191.7s have in a real project or company that
194.5s you're working on. The second part is
196.6s the apps which is high fidelity clones
198.6s of popular applications, Salesforce,
201.1s ServiceNow, Microsoft 365, etc. that
204.1s agents can interact with via MCP, CLI,
206.7s or Kua. And then the third part is the
208.8s tasks where we have prompts and
210.6s verifiers. Verifiers could be rubrics or unit tests that can be used either for eval or training. And the barrier for frontier labs to automate everything that you can do on your laptop using Claude is how do they cover the full distribution of all of the worlds, all of the apps, and all of the tasks in the
213.4s unit tests that can be used either for
215.4s eval or training. And the barrier for
217.9s frontier labs to automate everything
220.5s that you can do on your laptop using
222.1s Claude is how do they cover the full
224.1s distribution of all of the worlds, all
227.1s of the apps, and all of the tasks in the
229.0s economy. And so, there's been this enormous scale out in order to do that. Um, where humans have been really central to how we build these environments, obviously with models in the loop meaningfully. And so I put a graph here of the amount of expert hours that um, of throughput in from our talent network
231.5s enormous scale out in order to do that.
234.9s Um, where humans have been really
237.0s central to how we build these
238.8s environments, obviously with models in
240.7s the loop meaningfully. And so I put a
242.4s graph here of the amount of expert hours
245.7s that um,
247.1s of throughput in from our talent network
249.2s over the last 24 months. Um, and it's a a pretty crazy trajectory with respect to um, 2.5 million hours um, in Q2 alone with growth sort of accelerating on the amount of expert time uh, used to build out all of these environments. The reason being of course as I mentioned we need to scale out the
252.4s and it's a a pretty crazy trajectory
255.3s with respect to um, 2.5 million hours
259.0s um, in Q2 alone with growth sort of
261.6s accelerating on the amount of expert
263.8s time uh, used to build out all of these
267.0s environments. The reason being of course
269.7s as I mentioned we need to scale out the
271.5s environment distribution across every category in the economy. Many of you might know GDP val where there's 205 domains in the Bureau of Labor Statistics across all the different jobs, but then you have to think through how do we have all of the apps corresponding to all of those jobs, all the different scenarios, all the tasks.
274.2s category in the economy. Many of you
275.8s might know GDP val where there's 205
279.0s domains in the Bureau of Labor
280.4s Statistics across all the different
282.4s jobs, but then you have to think through
284.2s how do we have all of the apps
285.4s corresponding to all of those jobs, all
287.2s the different scenarios, all the tasks.
289.4s Is this enormous build out. Only humans can measure the frontier in most domains, not every domain. There are rare exceptions like math where you have a really clean simulation environment and so uh, the model's able to learn from whether it got the right answer, but in most domains like building a slide deck uh, the model has an
292.6s can measure the frontier in most
294.0s domains, not every domain. There are
295.6s rare exceptions like math where you have
298.0s a really clean simulation environment
299.7s and so uh, the model's able to learn
302.0s from whether it got the right answer,
303.5s but in most domains like building a
305.7s slide deck uh, the model has an
307.9s incredibly hard time identifying reliably where it made its own mistake. It's as if you would be asking a human to grade their own homework. And so that's why it's really valuable to have a human create a rubric similar to how a professor would create a rubric to grade an essay or a TA would grade that slide
309.8s reliably where it made its own mistake.
311.9s It's as if you would be asking a human
314.0s to grade their own homework. And so
316.5s that's why it's really valuable to have
319.0s a human create a rubric similar to how a
322.0s professor would create a rubric to grade
323.7s an essay or a TA would grade that slide
326.1s deck. Similar to the way that a lot of us learn it's in large part from the feedback we got from those around us rather than uh, purely plugging things into a calculator or clean simulation. Um, and then building these verifiers is hard cuz anytime you're building the slide deck you need to understand the
327.9s us learn it's in large part from the
330.0s feedback we got from those around us
331.8s rather than
333.2s uh, purely plugging things into a
335.1s calculator or clean simulation. Um,
338.1s and then building these verifiers is
339.8s hard cuz anytime you're building the
341.4s slide deck you need to understand the
344.4s full problem space of what are the 10 different slide decks that, you know, could be a good path to go down? What are the dozens of mistakes you could possibly make? And how do you build a comprehensive verifier that captures this full solution area of what's possible? And so, what I'll walk through is a sample RL environment.
346.4s different slide decks that, you know,
348.2s could be a good path to go down? What
350.5s are the dozens of mistakes you could
352.0s possibly make? And how do you build a
353.8s comprehensive verifier that captures
356.3s this full solution area of what's
357.7s possible? And so, what I'll walk through
359.8s is a sample RL environment.
362.0s Excuse me, to also break this down for all of you. Um and part of the reason that this is so cool, which I'll get to in a moment, is that we developed a lot of this technology in collaboration with the labs. These are, of course, ones that we have open-sourced and published
363.8s all of you. Um and part of the reason
366.3s that this is so cool, which I'll get to
367.9s in a moment, is that we developed a lot
370.2s of this technology in collaboration with
373.0s the labs. These are, of course, ones
375.0s that we have open-sourced and published
377.1s to the world, but now that's all starting to get disseminated to the application layer companies that are building and owning their own intelligence. As they realize that the three core pillars of their AI strategy are their compute, their algorithms or researchers, and the data sets they build. And data's often the most differentiating factor. And so, this is
379.1s starting to get disseminated to the
381.9s application layer companies that are
384.1s building and owning their own
385.3s intelligence. As they realize that the
387.6s three core pillars of their AI strategy
390.2s are their compute, their algorithms or
392.9s researchers, and the data sets they
394.7s build. And data's often the most
396.2s differentiating factor. And so, this is
398.2s one that we published, um as a legal environment, where we have lawyers from top law firms like Latham and Watkins write out a scenario of a real project that they worked on in their big law job. And then they create a full outline for a data room that corresponds to all of the different uh
400.6s as a legal environment, where we have
402.4s lawyers from top law firms like Latham
405.0s and Watkins write out a scenario of a
408.9s real project that they worked on in
410.5s their big law job. And then they create
412.4s a full outline for a data room that
415.1s corresponds to all of the different uh
417.6s messages, emails, files, size of files. I cut off the full data room cuz it's it's very extensive. Um and of course, there's a lot of model in the loop with how they effectively populate this. Similar to how a software engineer now should not be coding by hand entirely themselves, they should probably be
421.7s I cut off the full data room cuz it's
423.6s it's very extensive. Um and of course,
426.0s there's a lot of model in the loop with
427.6s how they effectively populate this.
429.4s Similar to how a software engineer now
431.5s should not be coding by hand entirely
433.8s themselves, they should probably be
435.6s orchestrating agents in how to do this very productively. Um and then we render that data room into the apps, the clones of Google Workspace you can see in this scenario, and have prompts uh to roll out model trajectories against this. And so, in this one, it's evaluating the maximum total liability for Star Tanker uh
437.7s very productively. Um
439.6s and then we render that data room into
441.4s the apps, the clones of Google Workspace
443.6s you can see in this scenario, and have
445.8s prompts uh to roll out model
447.9s trajectories against this. And so, in
449.6s this one, it's evaluating the maximum
452.2s total liability for Star Tanker uh
454.6s Tankers International Limited compared to Cooper Jefferies Energy Corporation under the Oil and Petroleum Act, uh considering all of the context uh from this real scenario in the data room. And then as I mentioned, similar to how a professor would create a rubric to grade an essay, they have these key rubric criteria that correspond to what are the
457.0s to Cooper Jefferies Energy Corporation
459.0s under the Oil and Petroleum Act, uh
461.1s considering all of the context uh from
463.6s this real scenario in the data room. And
464.9s then as I mentioned, similar to how a
466.7s professor would create a rubric to grade
468.3s an essay, they have these key rubric
470.4s criteria that correspond to what are the
473.4s characteristics of a of a accurate model response. And making sure that these rubric criteria avoid reward hacking and effectively align with the uh when you roll out 100 trajectories, making sure all of those scores are accurate is incredibly technically challenging. And so there's an enormous amount of research, agenda quality control, training on the data, etc. that goes
476.1s response. And making sure that these
478.1s rubric criteria avoid reward hacking and
481.4s effectively align with the uh when you
485.1s roll out 100 trajectories, making sure
486.9s all of those scores are accurate is
489.0s incredibly technically challenging. And
491.4s so there's an enormous amount of
493.5s research, agenda quality control,
495.3s training on the data, etc. that goes
497.4s into how you solve that problem and then ultimately produce these high-quality verifiers and leaderboards um that give you an aggregate model score across how well all of the different models are um doing on a particular domain. And as we can see, one of the big changes over the last few months is that GLM 52 and
498.9s ultimately produce these high-quality
500.6s verifiers and leaderboards um that give
503.3s you an aggregate model score across how
505.7s well all of the different models are um
507.8s doing on a particular domain. And as we
510.7s can see, one of the big changes over the
512.6s last few months is that GLM 52 and
515.4s Chimera K3 are on the leaderboard. And so that is a huge opportunity for all of you because that gives us the foundation to actually achieve frontier intelligence and all of the specific applications uh and verticals that you're focusing on um that's not too far away. And so to give a little bit of
517.6s so that is a huge opportunity for all of
520.5s you because that gives us the foundation
524.2s to actually achieve frontier
526.4s intelligence and all of the specific
528.6s applications uh and verticals that
531.2s you're focusing on um that's not too far
534.0s away. And so to give a little bit of
536.4s context on what that looks like, um I'll share an example of post-training on Apex Agents, which is the data set that I uh or the sample I just showed before, where we have 1,800 tasks in this example. This was post-training run of GLM 47, but we're redoing a lot of them for Chimera K3, so we'll have updated
538.6s share an example of post-training on
540.6s Apex Agents, which is the data set that
543.3s I uh or the sample I just showed before,
546.2s where we have 1,800 tasks in this
548.7s example. This was post-training run of
551.0s GLM 47, but we're redoing a lot of them
553.9s for Chimera K3, so we'll have updated
555.4s results for you all soon. Where you can see the jumps just on 1,800 tasks with about 500k in compute are pretty dramatic. Um corporate law going from 4.7 to 26.6, um but notice that this is just Apex Agents data set we gave it, and it actually generalized incredibly well to GDP valve and uh Apex V1, which doesn't
557.5s see the jumps just on 1,800 tasks with
560.8s about 500k in compute are pretty
563.3s dramatic. Um corporate law going from
566.2s 4.7 to 26.6,
569.6s um but notice that this is just Apex
572.6s Agents data set we gave it, and it
574.6s actually generalized incredibly well to
577.3s GDP valve and uh Apex V1, which doesn't
580.4s have these data rooms. Even just seeing nominal gains and some other benchmarks as well. And so, we're doing a lot of this work of working with customers like Harvey, who I know will present on stuff later to help build out the environments corresponding to their specific domain so that they can build frontier intelligence within that. And I
582.4s Even just seeing nominal gains and some
585.3s other benchmarks
587.2s as well. And so, we're doing a lot of
590.0s this work of working with customers like
592.7s Harvey, who I know will present on stuff
594.9s later to help build out the
598.4s environments corresponding to their
600.6s specific domain so that they can build
602.9s frontier intelligence within that. And I
605.3s think Andrew talked about how Cursor was a great first example of how an application layer company could build a industry-leading model that, you know, built an enormous amount of value for their customers. And I believe that over the next 12 months, there is going to be dozens of examples just like that where
607.6s a great first example of how an
610.2s application layer company could build a
613.3s industry-leading model that, you know,
615.4s built an enormous amount of value for
616.9s their customers. And I believe that over
619.0s the next 12 months, there is going to be
621.2s dozens of examples just like that where
623.7s companies own their own intelligence and that is the key source of the modes that they're building. And actually, Josh and I talked about this the other day as well. Um a couple of examples of ways to curate high-quality data sets. The general three that we see most that I'm happy to
625.9s that is the key source of the modes that
628.1s they're building. And actually, Josh and
629.9s I talked about this the other day as
631.2s well. Um
633.4s a couple of examples of ways to curate
635.8s high-quality data sets.
637.8s The general three that we see most that
640.5s I'm happy to
642.1s talk about and send people links to is first by task is the most common. Where people would say, I really like this data shape of environments in law and we will pay 2,000 per task to scale this up. And as an example, certain frontier labs might buy 50,000 tasks a month from us.
644.4s first by task is the most common.
647.2s Where people would say, I really like
648.8s this data shape of environments in law
653.2s and we will pay 2,000 per task to scale
656.6s this up. And
658.6s as an example, certain frontier labs
660.7s might buy 50,000 tasks a month from us.
663.8s And so, it tends to be pretty dramatic And so, it tends to be pretty dramatic scale. And these tasks would generally be very complex. Some would even take humans up to a month to complete that given task. Sometimes it would take just a few hours. And these would be sort of custom per task pricing. Second is
666.4s And so, it tends to be pretty dramatic scale.
667.6s And these tasks would generally be very
669.3s complex. Some would even take humans up
671.9s to a month to complete that given task.
674.0s Sometimes it would take just a few
676.1s hours. And these would be sort of custom
678.9s per task pricing. Second is
680.4s off-the-shelf data where we have we've invested hundreds of millions of dollars in building our own data sets that we sell to multiple customers. All these new neo labs are generally airing more towards off-the-shelf data because it doesn't make sense for 10 different labs to all be building their own uh data sets. Um
682.1s where we have we've invested hundreds of
684.2s millions of dollars in building our own
686.0s data sets that we sell to multiple
688.2s customers. All these new neo labs are
690.4s generally airing more towards
692.0s off-the-shelf data because it doesn't
694.1s make sense for 10 different labs to all
696.2s be building their own uh data sets. Um
698.9s there's a lot of value to building something once that can uh then be applied to everyone. And then the final which um we see a little bit of, but is uh less of our focus anymore, is just providing um the experts so that customers are able to uh or organize the experts on their own
700.1s something once that can uh then be
702.2s applied to everyone. And then the final
704.0s which um we see a little bit of, but is
706.9s uh less of our focus anymore, is just
709.8s providing um the experts so that
712.7s customers are able to
714.9s uh or organize the experts on their own
717.6s uh or organize the experts on their own um in just an hourly model. Um so we we do a little bit of that when people like that was how Harvey got started with us that was how Harvey got started with us hiring uh some lawyers, um but it generally moves towards more of these uh scaled
718.2s in just an hourly model. Um so we we do
720.8s a little bit of that when people like
722.8s that was how Harvey got started with us
724.3s that was how Harvey got started with us hiring
725.4s uh some lawyers, um but it generally
727.4s moves towards more of these uh scaled
730.0s offerings of data over time. So, that's a little bit of the background of how to build RL environments, what they are, and I'm really excited about all this technology that we have uh that has previously been limited to the frontier labs all making its way to all of you. And so, happy to answer any questions um
732.4s a little bit of the background of how to
734.4s build RL environments, what they are,
736.9s and I'm really excited about all this
738.3s technology that we have uh
741.4s that has previously been limited to the
743.1s frontier labs all making its way to all
745.4s of you. And so, happy to answer any
747.2s questions um
748.8s about that. Go ahead. Hey, uh um I'm Ali from Astro Guide. My question is it's kind of open-ended, but simple. How do you price data? Like how do you value data? So, there's a so many different ways. I So, there's a so many different ways. I mean, the most natural would be our customers
755.8s Go ahead.
757.7s Hey, uh
759.3s um I'm Ali from Astro Guide. My question
761.4s is it's kind of open-ended, but simple.
763.6s How do you price data?
765.5s Like how do you value data?
767.8s So, there's a so many different ways. I
770.1s So, there's a so many different ways. I mean,
771.2s the most natural would be our customers
773.7s care about model improvement, right? And so, our customers have a given goal of they want to you know, be at the frontier on a given leaderboard. And so, we're able to work backwards from how much is that worth to them and how much should we charge uh per task, how many
776.0s so, our customers have a given goal of
778.6s they want to you know, be at the
780.5s frontier on a given leaderboard. And so,
783.5s we're able to work backwards from how
785.6s much is that worth to them and how much
787.2s should we charge uh per task, how many
789.5s tasks do we think would get them to that goal. And so, when we think about uh a company like Nvidia, they're probably willing to pay, you know, a billion dollars to have a frontier open-source model. And so, there's a lot of complexity of like how do we you know, price all the different ingredients that
791.1s goal. And so, when we think about uh a
793.1s company like Nvidia, they're probably
795.4s willing to pay, you know, a billion
797.0s dollars to have a frontier open-source
799.8s model. And so, there's a lot of
801.6s complexity of like how do we you know,
803.2s price all the different ingredients that
805.0s go into um making that happen. The other way that we price when it lens we look at it through is also our cost structure to make them where of course when we have a task that takes 10 hours of human time and we're paying the human 150 an hour there might be a 1500 cost basis and so
808.4s The other way that we price
811.0s when it lens we look at it through is
813.6s also our cost structure to make them
816.6s where of course
819.1s when we have a task that takes 10 hours
822.3s of human time and
824.8s we're paying the human 150 an hour
827.2s there might be a 1500 cost basis and so
830.2s then it becomes a question of what margin do we want to run on top of that based on how differentiated and frontier that specific task is. But it's it's super wide range. Like we have tasks that range from 50 to 10,000. So. How do you think about data quality? You mentioned like you know utilize human
832.1s margin do we want to run on top of that
834.4s based on how differentiated and frontier
836.8s that specific task is.
839.3s But it's it's super wide range. Like we
840.9s have tasks that range from 50 to
843.2s 10,000. So.
845.8s How do you think about data quality? You
848.3s mentioned like you know utilize human
850.1s expert to label data and how do you how do you compare the human label data and you know the frontier lab you know you know the frontier lab you know frontier alliance the judgment data? How do you compare them to from your opinion? So first question was how do we think about quality? Second one was sort of
852.1s do you compare the human label data and
854.9s you know the frontier lab you know
856.6s you know the frontier lab you know frontier
857.9s alliance the judgment data? How do you
860.5s compare them to from your opinion?
863.6s So first question was how do we think
865.6s about quality? Second one was sort of
867.2s how do we compare the like judgment of preference labels to the auto graders? preference labels to the auto graders? Yeah. So the core spot which is generally when people say data quality they're referring to two things. First is realism and secondly is accuracy of realism and secondly is accuracy of verifiers.
870.7s preference labels to the auto graders?
873.1s preference labels to the auto graders? Yeah.
873.8s So the core spot which is generally when
876.4s people say data quality they're
878.5s referring to two things. First is
881.3s realism and secondly is accuracy of
884.5s realism and secondly is accuracy of verifiers.
886.2s On realism it's just like they want to automate everything in the economy that corresponds to corporate law in this case right? And so it's like how do we make sure that this actually reflects the real distribution of what we would see in a real lawyer's environment. And that's one of the reasons that experts
888.4s automate everything in the economy that
890.9s corresponds to corporate law in this
893.1s case right? And so it's like how do we
894.8s make sure that this actually reflects
897.3s the real distribution of what we would
899.6s see in a real lawyer's environment. And
902.9s that's one of the reasons that experts
904.7s create outlines and help to guide all the processes of the data curation. Realism of the environment the apps the tasks everything is incredibly important and and also granularly understanding the taxonomy that drives that realism across the entire distribution that you're looking for. The second part of it relates to the other way people think
906.5s the processes of the data curation.
908.3s Realism of the environment the apps the
911.0s tasks everything is incredibly important
914.1s and and also granularly understanding
916.2s the taxonomy that drives that realism
918.8s across the entire distribution that
920.1s you're looking for. The second part of
923.2s it relates to the other way people think
925.4s about quality, which is the accuracy of the verifiers. Cuz the way that you would train one of these models is you might roll out 100 trajectories of Kimi K3 and then use this rubric to score all of those trajectories. And as you can imagine, there's like so many different paths that a model can go down. And so
927.1s the verifiers. Cuz the way that you
929.0s would train one of these models is you
930.5s might roll out 100 trajectories of Kimi
933.1s K3 and then use this rubric to score all
936.5s of those trajectories. And as you can
938.5s imagine, there's like so many different
941.2s paths that a model can go down. And so
943.5s you want to make sure that the way this rubric is doing the scoring is the same as if we were to just have human stack rank those 100 trajectories. And so what we do for that is a process called trajectory analysis, where we roll out 10 trajectories of the model that we're focused on improving,
945.6s rubric is doing the scoring is the same
948.1s as if we were to just have human stack
949.6s rank those 100 trajectories. And so what
951.9s we do for that is a process called
953.7s trajectory analysis, where we roll out
956.6s 10 trajectories of the model that we're
958.5s focused on improving,
960.2s um, and then score all of those, and have some combination of agentic quality control systems and some human review go through to make sure that, um, all of the scores align with, uh, the goals. the scores align with, uh, the goals. Um, and sometimes you can also use, uh, human feedback evals or preference
962.7s have some combination of agentic quality
965.6s control systems and some human review go
968.6s through to make sure that, um, all of
972.1s the scores align with, uh, the goals.
975.4s the scores align with, uh, the goals. Um,
976.0s and sometimes you can also use, uh,
978.3s human feedback evals or preference
980.1s labels to as a eval for your auto grader, um, is the other way related to that that you're able to solve for it. Um, how much do you think like synthetic data generation plays into all of this, especially like creating these large data rooms? So the fascinating thing is I think that there's been a lot of misinterpretation
982.6s grader, um, is the other way related to
985.1s that that you're able to solve for it.
990.8s Um, how much do you think like synthetic
992.7s data generation plays into all of this,
994.9s especially like creating these large
996.1s data rooms?
998.0s So the fascinating thing is I think that
999.9s there's been a lot of misinterpretation
1002.8s of what people mean when they say synthetic data because, like, RLVR is a bet on synthetic data. It's basically let's roll out a bunch of synthetic model trajectories rather than having the humans write the SFT, and then let's score all of them, and let the models learn from all of these like, uh,
1004.0s synthetic data because, like, RLVR is a
1006.8s bet on synthetic data. It's basically
1009.0s let's roll out a bunch of synthetic
1011.0s model trajectories rather than having
1012.7s the humans write the SFT, and then let's
1014.8s score all of them, and let the models
1016.8s learn from all of these like, uh,
1019.2s synthetic model trajectories. So I think that's the first way that synthetics get used. The second way is that models play giant role in the way that we populate environments and create tasks in the same way that, um, a lawyer that is writing a legal memo should definitely be using Claude or ChatGPT to do that. The
1021.4s that's the first way that synthetics get
1023.9s used. The second way is that models play
1027.2s giant role in the way that we populate
1031.1s environments and create tasks in the
1033.1s same way that, um, a lawyer that is
1036.6s writing a legal memo should definitely
1039.0s be using
1040.6s Claude or ChatGPT to do that. The
1043.2s experts that are building out these data rooms should definitely be using Claude, ChatGPT, or whatever model to help them do that. And there's a lot of ways that the model can make them more efficient. But the reason that humans are still an essential component of the process that's incredibly differentiated is that you need humans
1045.0s rooms should definitely be using Claude,
1046.7s ChatGPT, or whatever model
1049.3s to help them do that. And there's a lot
1051.5s of ways that the model can make them
1052.7s more efficient. But the reason that
1054.5s humans are still an essential component
1056.8s of the process that's incredibly
1058.2s differentiated is that you need humans
1061.5s almost definitionally to measure what is beyond the frontier of the model beyond the frontier of the model capabilities. Like the models you you can't just tell the model like come up with the legal environment and then like tell me which of your legal memos are like good and bad. It's super noisy and there's not
1063.4s beyond the frontier of the model
1065.0s beyond the frontier of the model capabilities.
1066.3s Like the models you you can't just tell
1068.2s the model like come up with the legal
1069.9s environment and then like tell me which
1071.6s of your legal memos are like good and
1073.1s bad. It's super noisy and there's not
1075.6s like clear signals associated with that. You need something that has capabilities beyond the frontier of that model to do so the frontier of that model to do so reliably. is RL environments seem like they're all the rage now and maybe have been for about a year. I had I hadn't really been
1077.4s You need
1078.8s something that has capabilities beyond
1080.7s the frontier of that model to do so
1082.1s the frontier of that model to do so reliably.
1088.1s is RL environments seem like they're all
1090.2s the rage now and maybe have been for
1091.8s about a year. I had I hadn't really been
1094.3s hearing about them prior to that and it was all human expert labeling. And so kind of why why why is it all about RL environments now? Is that the most relevant thing for application companies to be thinking about? And is there any something after RL environments? Um so I'll I'll start with why it's
1095.7s was all
1097.1s human expert labeling. And so
1099.6s kind of why why why is it all about RL
1101.7s environments now? Is that the most
1103.4s relevant thing for application companies
1105.2s to be thinking about? And is there any
1106.6s something after RL environments?
1109.3s Um so I'll I'll start with why it's
1112.6s become the rage and maybe some of the differences also between the deep research paradigm and the like envi- the sort of environments paradigm as we saw it in 2025. And then I'll talk about looking forward what we see evolving in the data landscape. Specifically, I think that the reason the deep research environments were the first was because
1115.4s differences also between the deep
1116.9s research paradigm and the like envi- the
1119.6s sort of environments paradigm as we saw
1121.8s it in 2025. And then I'll talk about
1123.6s looking forward what we see evolving in
1125.2s the data landscape. Specifically, I
1127.4s think that the reason the deep research
1130.0s environments were the first was because
1131.6s like deep research had tool use with search. So search was the tool in the environment that the model would work in. But the experts would not necessarily be populating apps. So it was sort of a lighter version of an RL environment where they would just like create rubrics corresponding to this.
1133.5s search. So search was the tool in the
1135.9s environment that the model
1138.1s would work in. But the experts would not
1140.5s necessarily be populating apps. So it
1142.9s was sort of a lighter version of an RL
1144.8s environment where they would just like
1146.2s create rubrics corresponding to this.
1148.2s And again, like I only talk about this stuff cuz it's a couple of years old at this point, so it's no longer uh super confidential. And then for the like trend of apps in 2025, I think that really became giant because people realized that the primary bottleneck to making the models useful was how they
1150.0s stuff cuz it's a couple of years old at
1151.4s this point, so it's no longer
1153.5s uh super confidential. And then for the
1156.4s like trend of apps in 2025, I think that
1161.0s really became giant because people
1163.2s realized that the primary bottleneck to
1166.0s making the models useful was how they
1167.7s started to use both all the context in the code base and all of the tools on our laptops, right? And so if we want this in the user distribution of usage, then we need to get it in the data distribution that the models are learning from. Uh and so there's going to continue be
1170.6s the code base and all of the tools on
1172.3s our laptops, right? And so if we want
1174.7s this in the user distribution of usage,
1177.4s then we need to get it in the data
1179.3s distribution that the models are
1181.2s learning from.
1182.5s Uh and so there's going to continue be
1185.1s to be this giant scale up of diversity across all three of these categories on a going forward basis, but there's going to be some changes um to maybe name two of those changes that we're thinking about the most. The first one is ultra long horizon. Like right now agents mostly aren't trained
1188.4s across all three of these categories on
1190.7s a going forward basis, but there's going
1192.2s to be some changes
1193.8s um to maybe name two of those changes
1196.4s that we're thinking about the most. The
1198.2s first one is ultra long horizon. Like
1201.2s right now agents mostly aren't trained
1203.7s to do things that are over 10 hours, and we need to start building tasks for things that might take a human 100 hours or even 1,000 hours to do. And so that's going to be a giant shift. And then the other large shift that we're seeing is introducing virtual co-workers, which corresponds to that. Um like one my
1206.2s we need to start building tasks for
1207.8s things that might take a human 100 hours
1210.2s or even 1,000 hours to do. And so that's
1212.4s going to be a giant shift. And then the
1214.6s other large shift that we're seeing is
1217.6s introducing virtual co-workers, which
1219.5s corresponds to that. Um like one my
1222.4s favorite questions to ask people when they're thinking about their data distribution is what percentage of tasks that they do in their job require interacting with other people. Uh and most people would say like 60 or 70. Some people say a lot more, some people say a little bit less. Um but
1224.0s they're thinking about their data
1225.0s distribution is what percentage of tasks
1227.8s that they do in their job require
1230.2s interacting with other people.
1232.2s Uh and most people would say like 60 or
1234.4s 70. Some people say a lot more, some
1236.2s people say a little bit less. Um but
1238.6s then if you map that on to what percentage of evals measure how well the models can interact with other people, it's like 1, maybe Tau bench has a little bit of this. little bit of this. Um and so there's this giant realism gap associated with how you actually measure how well agents engage in social
1241.0s percentage of evals measure how well the
1244.2s models can interact with other people,
1246.4s it's like 1, maybe Tau bench has a
1248.9s little bit of this.
1250.4s little bit of this. Um
1251.0s and so there's this giant realism gap
1254.0s associated with how you actually measure
1257.1s how well agents engage in social
1259.1s interaction um throughout uh all of the different people and other agents that they need to work with uh in their jobs. they need to work with uh in their jobs. Good. You talked about rubric generation, which is like is it like a bespoke access of verify that you put task? And from what I understood, that's like
1262.3s different people and other agents that
1264.2s they need to work with uh in their jobs.
1267.7s they need to work with uh in their jobs. Good.
1268.4s You talked about rubric generation,
1270.2s which is like is it like a bespoke
1272.6s access of verify that you put task? And
1274.9s from what I understood, that's like
1276.0s bottleneck by experts. Have you found any success of like being able to scale that up with your models giving you like some heuristic or even like post any models for We have found that you can make it a lot more efficient if you have a like AI copilot that's able to work with the
1278.0s any success of like being able to scale
1280.2s that up with your models giving you like
1282.2s some heuristic or even like post any
1284.4s models for
1286.4s We have found that you can make it a lot
1289.3s more efficient if you have a like AI
1292.0s copilot that's able to work with the
1294.2s expert in creating the task and the expert in creating the task and the verifier. Um so, the expert can talk to the trajectory and understand exactly what's happening and where it's going wrong. The challenge is just that if you're trying to improve FABLE, FABLE cannot reliably write out the like rubric criteria for where it's making
1296.2s expert in creating the task and the verifier.
1297.6s Um so, the expert can talk to the
1300.8s trajectory and understand exactly what's
1302.7s happening and where it's going wrong.
1304.9s The challenge is just that if you're
1307.0s trying to improve FABLE,
1308.9s FABLE cannot reliably write out the like
1313.1s rubric criteria for where it's making
1314.6s mistakes. It might get like half of them right and half of them wrong, and that amount of noise is unworkable from a training standpoint. Um and so, that's the reason that the the process that requires humans the most is the task creation. Uh like a lot of the environments, we can get
1317.4s right and half of them wrong, and that
1320.0s amount of noise is unworkable from a
1323.2s training standpoint. Um and so, that's
1326.0s the reason that the the process that
1328.6s requires humans the most is the task
1330.6s creation. Uh like a lot of the
1332.2s environments, we can get
1334.9s like use a lot of synthetic. Uh it's helpful to have humans write the outlines they're familiar with the environment and and uh and they're grounded in reality, a realistic distribution, but with the task that's those tend to really require humans. Um with with rare exceptions, um in code or if you're sort of distilling
1337.0s it's helpful to have humans write the
1338.3s outlines they're familiar with the
1339.5s environment and and uh
1342.2s and they're grounded in reality, a
1343.8s realistic distribution, but with the
1346.6s task that's those tend to really require
1349.0s humans. Um
1350.6s with with rare exceptions, um
1353.6s in code or if you're sort of distilling
1357.3s from like if you have a model that's worse than Kimikaze 3, then it can definitely learn from tasks Kimikaze 3 is creating. So, there are there are some exceptions if you're doing it that some exceptions if you're doing it that way. Hey, this is uh Nikhil from uh Cribl. So, when we think about RL environments
1358.8s worse than Kimikaze 3, then it can
1361.0s definitely learn from tasks Kimikaze 3
1362.5s is creating. So, there are there are
1363.7s some exceptions if you're doing it that
1365.6s some exceptions if you're doing it that way.
1367.9s Hey, this is uh Nikhil from uh Cribl.
1370.2s So, when we think about RL environments
1372.3s for certain provable domains like cyber defense or incident response, where there's the model is or the agent is trying to find a flaw in an existing trying to find a flaw in an existing system, Do you use uh humans for just authoring that environment or setting it up or do you
1375.0s defense or incident response, where
1377.2s there's the model is or the agent is
1379.1s trying to find a flaw in an existing
1381.0s trying to find a flaw in an existing system,
1382.2s Do you use uh
1384.0s humans for just authoring that
1386.0s environment or setting it up or do you
1387.3s also use that for grading? Is there Is there a way to scale that up? I I actually think cyber is one where you don't necessarily always need humans for the verifiers cuz you can have an attacker and a defender agent. attacker and a defender agent. Um and uh I think you're right in saying
1389.1s there a way to scale that up?
1391.4s I I actually think cyber is one where
1394.5s you don't necessarily always need humans
1396.9s for the verifiers cuz you can have an
1398.1s attacker and a defender agent.
1400.3s attacker and a defender agent. Um
1400.8s and uh I think you're right in saying
1403.6s that for cyber you can have humans more so or or architect what is a realistic like environment um and sort of set up the environment um cuz you do need a lot of diversity. Uh but then it's less human intensive with respect to uh building verifiers. with respect to uh building verifiers. Thanks. the base model?
1406.9s so or or architect what is a realistic
1410.0s like environment um and sort of set up
1412.9s the environment um cuz you do need a lot
1415.5s of diversity.
1416.8s Uh but then it's less human intensive
1419.7s with respect to uh building verifiers.
1423.3s with respect to uh building verifiers. Thanks.
1427.9s the base model?
1430.2s What do you mean by that? Well, ostensibly you're using the same data set for all these models here and they somewhat land around the same final performance on this list here. But maybe if you try to smaller model, which is maybe a good place to start, it would land much lower. Um what is the cause
1431.2s Well, ostensibly you're using the same
1433.3s data set for all these models here and
1435.5s they somewhat land around the same final
1437.6s performance on this list here. But maybe
1440.1s if you try to smaller model, which is
1441.3s maybe a good place to start, it would
1443.1s land much lower. Um what is the cause
1445.6s there? Is it just parameter count or there? Is it just parameter count or So the parameter count will definitely play a role in how effectively the model does insofar as how trainable it is. I think that um the main thing to look at is generally the gap between the like
1448.3s there? Is it just parameter count or So
1449.9s the parameter count will definitely play
1452.2s a role in
1454.2s how effectively the model does insofar
1458.2s as how trainable it is. I think that um
1462.5s the main thing to look at is generally
1464.5s the gap between the like
1468.4s pass at 16 and the pass at one. If you have a model where you roll out 16 trajectories and it gets all of them totally wrong, the it's sort of like hopeless that the model is going to learn from that um for the most part. Maybe you roll out another 100 trajectories and it gets one of them
1471.6s have a model where you roll out 16
1473.9s trajectories and it gets all of them
1475.7s totally wrong, the it's sort of like
1478.4s hopeless that the model is going to
1480.0s learn from that um for the most part.
1482.5s Maybe you roll out another 100
1484.4s trajectories and it gets one of them
1485.7s right. Um versus if you have um the ideal case is that you have the ideal case is that you have um pass at one it fails, but then pass at 16 when you roll out 16 trajectories it gets it right once or twice, and then the model is able to learn very effectively from that. So, that's
1487.3s versus if you have um
1489.7s the ideal case is that you have
1492.0s the ideal case is that you have um
1492.8s pass at one it fails, but then pass at
1495.0s 16 when you roll out 16 trajectories it
1497.1s gets it right once or twice, and then
1499.6s the model is able to learn very
1501.0s effectively from that. So, that's
1502.4s generally the heuristic we would use for how strong the base model needs to be to effectively learn from a given data set. Cool. Maybe Oh, final question really quick. What advice do you have for companies as they partner with you on data that they should rely on themselves as part of their RL post-training versus
1506.0s how strong the base model needs to be to
1509.1s effectively learn from a given data set.
1515.9s Cool. Maybe Oh,
1516.8s final question really quick.
1519.1s What advice do you have for
1521.5s companies as they partner with you
1523.8s on data that they should rely on
1525.8s themselves as
1527.6s part of their RL post-training versus
1529.8s part of their RL post-training versus relying What's the complementary Well, I think this is why a giant portion of our business is custom data, where it's like we have teams that are siloed and fully exclusive to critical customers to make sure that we build out the best data sets in the world that they own.
1531.0s What's the complementary
1532.8s Well, I think this is why a giant
1534.6s portion of our business is custom data,
1536.8s where it's like we have teams that are
1539.1s siloed and fully exclusive to
1542.5s critical customers to make sure that we
1545.0s build out the best data sets in the
1546.4s world that they own.
1548.1s And that allows them to maintain their competitive advantage associated with this, while also benefiting from all of the infrastructure that we've built. And I think there are some companies that try to like build out all of the talent network and infrastructure in-house, but I think if you look at what the frontier
1551.1s competitive advantage associated with
1552.8s this, while also benefiting from all of
1554.8s the infrastructure that we've built. And
1556.2s I think there are some companies that
1557.6s try to like build out all of the talent
1560.2s network and infrastructure in-house, but
1562.4s I think if you look at what the frontier
1564.0s labs do and the best models do, it's a pretty good indication that there's so many economies of scale from working with a partner that has all of these economies of scale, the platform, the talent network, etc. So, anyways, thanks for all having me. for all having me. applause
1565.7s pretty good indication that there's so
1567.8s many economies of scale from working
1569.5s with a partner that has all of these
1571.9s economies of scale, the platform, the
1573.8s talent network, etc. So, anyways, thanks
1576.3s for all having me.
1578.6s for all having me. applause