@LangChain: On the latest Max Agency, @cognition president @russelljkaplan shared why devs stopped chasing the best model and start…
Summary
LangChain's Max Agency podcast features Cognition president Russell Kaplan discussing why developers now prioritize efficiency over the best model, and how agents like Devin are being deployed at scale.
View Cached Full Text
Cached at: 08/11/26, 05:42 AM
On the latest Max Agency, @cognition president @russelljkaplan shared why devs stopped chasing the best model and started chasing efficiency. YouTube: https://youtube.com/watch?si=Ff-wq0tBPTx-SC8V&v=bBUotstDLdk&feature=youtu.be… Apple: https://podcasts.apple.com/us/podcast/the-misaligned-incentives-behind-ai-coding-agents/id1891551672?i=1000779120914… Spotify: https://open.spotify.com/episode/7JNB5YXA6NVgUDAhy8eA3V…
@LangChain: On the latest Max Agency, @cognition president @russelljkaplan shared why devs stopped chasing the best model and start…
Channel: @LangChain Source: https://www.youtube.com/watch?si=Ff-wq0tBPTx-SC8V&v=bBUotstDLdk&feature=youtu.be
Transcript
I started my own machine learning career at Tesla on the autopilot team. I talked to my friends at Tesla today and the bottleneck is no longer just training bigger and bigger models. It’s actually running the evals.
Today I’m talking to Russell Kaplan, president at Cognition, the company behind Devon, an agent that went from viral demo to deploying code inside some of the most complex orgs in the world. We essentially have an evaluator agent that can take a session and say was it productive or not? It gave us the confidence to actually go to our customers and say we are actually going to make a $10 million productivity [music] guarantee. He explains why and how proactive agents are giving human engineers outsized leverage. You have like all these great suggestions of fixes that need to be applied. Oh yeah, that looks good. I want you to change that here. Individual developers have essentially realized I can be the CTO of an army of 10,000 agents. We get into what running agents at scale actually costs and how cognition drives it down. There’s organizations where the per person token spend is starting to eclipse the human salary spend. By [music] being a little bit more clever about some of the routing, we can get about 35% better [music] price performance with actually a slight increase in quality. And he argues there’s no such thing as the best model anymore. The fable class models are really good for [music] high precision, but we actually find that GP 5.5 and 5.5 Cyber are better on recall. You have to use both. Welcome to Max Agency, [music] the podcast that goes deep into how the best agents are being built by builders like you. [music] You guys had a massive launch about two and a half years ago, three years ago at this point and I think you pioneered a lot of really interesting concepts especially around UX and interacting with Devon in Slack. I think that was the first time I saw that so prominently. So I remember a lot from that launch. I remember a lot about the UX. I’m sure others do as well. What should people know about what you guys have been up to over the past 2 and a half years? Yeah, we launched Devon in March of
- uh was the original demo video that went super viral and at the time I would say coding agents were just at the edge of possible. you kind of squint and see, okay, this is going to work at some point. And we sort of put together a first pass example of what could that look like? You know,
SweetBench was like 13% I think you guys hit and it like tripled the previous best. Totally. It was like Yeah, we were like 13% Sweetbench. Um this is a you know, really exciting. And it actually that by the way that that number reminds me a lot of u some of our more recent evals and where we are now in Aenta Coding with a with a new eval um called Frontier Code which we can talk about. But uh when we you know we kind of launched in March of 24 we had this point of view that eventually you’re just going to have uh agents as teammates that you can delegate complete units of work to. And what we learned is that took us a few months from okay this is a prototype that we can kind of see the future of to this is something that’s actually useful internally. I think it was June of 24 that Devon became the number one committer to Devon which was like our first big milestone. That’s still a while ago. It was a while ago, but it was like a lot of manual dog fooding and like really kind of grinding to make the internal workflows nice and and then it took us a few more months to actually get this deployed in production and useful at a customer. And in in the sort of 2024 era async cloud coding agents, they really couldn’t do most of the tasks in software engineering, but there were some, you know, niches already. And I think like one of the best early ones was migrations and refactors. If you ran essentially a you think of like a reax++ style workflow across a large codebase with sort of bits of intelligence sprinkled in that actually already worked reasonably well in late 2024. And so that’s where we got our you know early product market fit. It was with bigger companies you know like more like enterprise companies who they just had lots of code that needed all this transformation. And why why was that such a good fit? There’s a few like technical reasons this made sense. one is um for a very large refactor or migration or you know ETL transformation whatever it’s sort of worth it to put in the effort to carefully prompt engineer your uh your cloud agents to be accurate and so you know if you if you kind of iterate on your prompt and your setup and your uh you know the context you’re feeding in and you you get it like just right and you can apply it across you know 10,000 modules in your codebase. This is actually really high ROI and you could it’s obviously much better than just sort of a find and replace you know string change but um it didn’t have to have this like full general software intelligence that we’re now delegating all sorts of coding task to so that worked well even the early days but it was it was still kind of niche and then in yeah December of 24 we launched Devon so anyone could sign up and we had this big debate internally of what should be you know what should we really emphasize or what should be the focus and I remember we did like that we just we flew the whole company to Utah to just like lock in in December and get this out uh you know kind of before the end of the year. And what we settled on was actually Slack as the primary interface. And so the entire launch video for Devon available self-service uh we called it internally at Devon. It was just at Devon at Devon at Devon and really trying to drive home this sort of user experience change which is collaborating with agents more like teammates. And so from there we got a lot more users, we got a lot more traction. you know, cloud agents have just gotten better and better since then. Both as the infrastructure has gotten better, as we’ve matured on that, as the models have gotten better, as the harnesses have gotten better, and now I think most of the work uh is being done by just delegating to to async agents. Yeah, maybe diving into that a little bit. What are some of the things that have gotten better that have allowed them to really take off? So, a a few things. So I think people always talk about the models getting better and that’s that’s super important but I think the the kind of infrastructure maturity is what was what was the first model that you felt was like really actually good enough. Oh yeah it’s you know we kept saying oh this is the model this is the model you know this is the new model. There were a few like phase changes I would say. I think sort of like the the the November 2024 uh like model series was uh one kind of that was like a big step function where like okay this actually can work a lot better now we we did a self-serve launch you know in part based on that um from open AI and then and then you know the anthropic model started getting really good um and we saw I think probably like you know July of of 25 another big step change I think now we see you know with like fable 5 this like new generation of models continuous step changes and now the step changes are so high that it’s actually no longer about the capabilities increases. I think we’re kind of seeing the opposite trend in some sense which um more and more of the tasks in software engineering are getting intelligent saturated like and so the sort of the sensitivity of a lot of developers has totally shifted from wait I need to be using the best possible model to holy cow I’m spending so much money on my coding agents how can I how can I be a little bit more efficient you know about this and I think I think this is like an underappreciated aspect of continuous gains in frontier intelligence which is that you For some workloads, you always want the most frontier intelligence. Um, especially if it’s like an adversarial game between parties and, you know, your intelligence has to be higher than the the counterparties intelligence, like if you’re trading in competition or something. But for a lot of software we want to build, you know, there’s kind of this um this saturation threshold where once it’s good enough, what you what you what you care about is speed and cost. And I think that’s actually where we see a huge amount of demand from developers and the companies we work with is okay, this already works. How can I now optimize the speed and cost? On the model side, maybe one question. What are tasks where you think fable type models are like where where you see that jump today? Like where where would you recommend people use them? Where do you see people using them? So I think we’re seeing um maybe in like this most recent generation of models and fable being one example um particularly big gains is in first of all there’s like detection remediation of cyber security vulnerabilities and we have there’s a lot of guardrails now in the public versions of these models and so it’s actually can be tricky to to get them to help. We see for example that if you’re trying to go just like clean up my security backlog, which by the way is a huge workload right now for a lot of the a lot of the companies we work with, you know, the the the fable class models are really good um for high precision. Um but we actually find that um you know, GP 5.5 uh and 5.5 cyber are better on recall. Um and so the best possible harness for finding and remediating security v vulnerabilities right now you have to use you have to use both and you sort of filter down you know the the ones that are maybe better found by the GP series with with the Fable series. I think that’s one big example. Um another one is that processing data from integrations for example from like data dog or um or other systems. It’s we do see like a step change in our emails in in Fable in particular. Is that because that data is like so large and messy that it needs more intelligence? I think a lot. Yeah, it’s just like there’s like a volume consideration. There’s sort of a multi-step reasoning consideration. And I think there’s just you can tell that there’s a lot of RL going on on like these very realistic data sources that have been scaled up a lot. Maybe the third category that’s maybe the most noticeable is a category that we’ve only even gotten better at detecting and understanding recently ourselves with the work we’ve been doing on new evals for coding. Because to your point when we launched Devon uh and it was you know uh in the teens on Sweetbench now Sweetbench is totally saturated and so how do we find how do we find the next um level of difficulty for evaluating models? We really we we looked for all of the ebells and we we couldn’t find any that could actually kind of match our sort of hazy internal intuitions of ah this model feels a lot better and when we sort of dug into it and really asked ourselves why is that I think the core gap for evals that we found was around mergeability which is okay this this code is technically correct but like would you actually merge it like would you feel happy would this improve the quality of your codebase and there’s lots of different subtle examples um of what that can mean it could be stylistically following the expectations that you already have. It could mean that it’s done in a way that it’s easy to modify in the future or that if other modifications happen, it gracefully handles those modifications. All these little stylistic things which we felt like were not being captured in uh in the evals. And so we we set out to basically make a new eval to to measure this better and also just have a new higher watermark of difficulty uh for Frontier coding capabilities. And then we released it recently. Um it’s called Frontier Code. And it’s once again uh it’s it’s it’s the new hardest coding eval. Uh the Frontier Co Diamond subset is like in the teens uh you know pre Fable 5 and now I think Fable 5 has has pushed to the 30s. So we have I wonder how many more months we have before that eval. I was going to say yeah in like a year we’ll be saying oh remember when it first got in the teens for this one. Yeah. I’m wondering I do wonder like it’s going to get harder you know it gets harder and harder to make these evals. I mean I started my own machine learning career at Tesla on the autopilot team and this was in like 2017 and the bottleneck was very clearly you know GPUs and compute and was like okay how do we get more compute we just need more compute and you know I talked to my friends at at Tesla today and the bottleneck is actually no longer just training bigger and bigger models it’s actually running the evals because for a self-driving system the interventions are so rare that you have to do enormous volume of driving to even find any problem at all in the stack And you know, I think for for software engineering, we’re not there yet. Like we still find bugs, but but we might we might hit that threshold sooner than we think. We talk with a bunch of teams around creating evals for their own systems using linksmith and other things like that. How did you guys create the frontier code? Like what did that process look like? So it started with okay we need to go find the sort of peak high taste developers who are going to have really strong both kind of stylistic quality but also code correctness standards to help us answer the question would you actually merge this code not just is this code passing tests is it is it functionally correct so that’s kind of part one and so we actually did like a you know a deep collaboration and recruiting campaign with a lot of um the sort of leading open-source authors who were you know really grateful for their collaboration on this it wouldn’t have been possible without them to go find, you know, what are the best known libraries in open source with with really high coding standards and then collaborating closely with them to help encode their their human intuition of what makes a PR something I would accept uh into these very carefully designed tests. That was part one. What what what do those tests look like? Are they LM as a judge? Are they programmatic assertions? Yeah, so multiple so so part of them are programmatic assertions. Of course, there’s like a correctness test, programmatic. There can be programmatic assertions on um kind of stylistic elements too like uh okay if you make this change in this poll request uh you know you have to actually use this module even though using this other module would be equivalent right now over time if they drift you know if you didn’t use the correct module then like you’re going to introduce a bug in the future so it can be asserted deterministically but it requires the judgment of you know the human author to to settle and then the other big thing which I think is still underdone by people practicing uh machine learning and working with agents is every single researcher on the cognition team, you know, hand contributed and reviewed and eval directly. And so this is not this is not something you can sort of throw over the wall and say, okay, the data quality piece that’s that’s such a slo um someone else can do that. You you sort of have to work on it directly, too. I’m assuming they worked with the open source authors to put together these assertions and and they would be running on these open source code bases. I guess like that was kind of the setup. it would be on on the open source codebase and you know we could run it in our infrastructure but then um we test and and you to avoid contamination we haven’t published the full set of question we publish some example questions but we’re trying to preserve this eval for the community for as many months as possible until it gets you know completely saturated um but running um yeah essentially running the code that already exists with the the patches applied by these models how do you guys score these evals is it binary like 01 is there some numeric component like if if you have 10 assertions on one of the test and it passes nine out of 10. How how is that scored? Yeah. So uh we broke it out into two components uh to have first there’s a binary element which is essentially yeah did this pass all of the blocking constraints of how we would evaluate this PR you know the simple things of like do the test pass are there hard are there hard deterministic criteria that the um maintainer of the codebase has inputed must be met by the uh by the agent who’s writing this code do you know off top like is that like five assertions or like 50 assertions like it’s highly varied per task but each task is you know like hundreds of hours of work of people putting into so it’s it is like this is not like a quick thing it’s the the process of constructing a single test uh you know eval line item in a data set like this it’s it feels like a it’s like a full project and labor of love per question you know where you’re really trying to think holistically what are all the things that I as the expert maintainer want to sort of imbue in my expectations so that’s how you get some of the binary you know pass past fail metrics but then we also to your point that’s often not enough to to really know okay would I merge this code and then also like how would this code rank to someone else’s code because there might be you know two different pieces of code you would merge but one is like a little bit more preferable than the other and so uh we also have the con concept of a score which is essentially like a linearly weighted aggregation of all of the non-blocking evaluation criteria so if it’s blocking evaluation criteria we’d say okay you know if it doesn’t pass you fail but like uh or this is just you know we’re not we’re not merging them if it’s non-blocking criteria. Uh, for example, there stylistic elements. A common one for us is is scope. So, I think one of the code smells of LM still is like they’ll make the right change, but then they’ll kind of mess with other files too that you really wish they didn’t mess with. And so, that’s not going to impact your correctness score, but it’s certainly going to impact the sort of stylistic elements. Um, and that’s going to downweight um, you know, if you had unnecessary touches in other files, that’s going to downweight your aggra metric. Same there for you know LM as a judge having heristics that you impose in LM as a judge can get added to this um to the to the linear score. We also found that um reverse classical evaluation is is really helpful. So commonly people say okay we need to make sure that after you accept this code change the test pass right but we also care just as much that um you know without this code change the test should fail right and that if you make this other code change the test should fail. So, how do you how do you impose kind of blocking constraints on both sides? But yeah, eval are really hard. You know, we spent a lot of time on this. I think I think we got it uh you know, we I think we did it right enough that all of the kernelms uh are pretty bad are pretty bad at this val. And it also more importantly to us, you know, as new models come out and we test them and we see their scores on Frontier Code, it’s it’s roughly matching the vibes like like if if a model is scoring really well on Frontier Code, then we get really excited. I like this idea of kind of like binary plus fail and then some numerical after that for the or more explicit I like the like blocking and non-blocking. So so we’re building a benchmark of our own we call like issue bench for linksmith engine which goes through and finds issues and there’s a bunch of stuff that like I would classify under like the non-blocking stuff like it’s yeah it would be nice if it did this it should probably do this and then there’s some other stuff that’s kind of like more blocking does it just find this like really bad issue should be more of a blocking thing. So I like that kind of like dual justosition. Do you guys use Harbor as a format for running evals or do you have your own internal kind of like eval? We do use Harbor. We we think in general a lot of the kind of standards are quite early and immature and I think one of the funny things if you sort of look in the whole Devon codebase is because it was the first coding agent there’s actually a whole bunch of stuff that like there might be a standard now or a correct way of doing things and we just we just rolled our own our our old implementation you know as funny as example funny examples even a basic um agent things like the concept of you know skills.mmd file like uh we we had implemented a concept in devon of of knowledge that uh was like before any open source standard existed for skills and now we’ve basically grafted on how can you also work with the open source standards but it is like a recurring theme in our codebase that we’ve gone off and sort of invented something weird and then it becomes some some permutation of it becomes an open source standard that we incorporated back in. We went down this rabbit hole talking about uh models and talking about the best models but you also mentioned kind of like cost and and presumably kind of like speed and latency are becoming other issues. You guys have done a few things here if I’m correct. You have Devon Fusion. You also have your own series of models that you guys have post-trained. How do you guys think about this uh this section of the model universe? We’re kind of in this unique time in history where anyone can hire as many AI agents as they want usually the way the tools are work. So, you know, uh me as a developer, yeah, I’m not going to I’m not going to use the cheap the cheap model. I want us I want to use the best one for everything. Um, but everyone is kind of collectively making that decision and then you sort of roll it all up and you realize, oh wow, like we are spending a lot of money. You know, there’s organizations where the the per person token spend is is very rapidly approaching or even starting to eclipse the, you know, the the human salary spend. And so it would be kind of crazy, you know, if if like the way you ran Langchain, for example, anyone could just hire a thousand people tomorrow without, you know, without talking to you. But that’s kind of how we run our teams today. And this is starting to become like a real issue for us. And like I like to think we’re like we’re pretty like you know we’re still startup we’re a lax like AI native kind of like forward organization but we are absolutely caring about kind of like token spend. I’m like do are you guys internally kind of like worried about token spend for your own kind of like engineering schemes as well? We have very high uh you know compute investment in general. So I would think our internal spend while being very high per person to us it’s like really valuable dog fooding investment in everything we do. But our customers are definitely thinking about if you just extrapolate the trend line, you know, it’s going to like eclipse the whole economy in that long with with exponential growth. And so so so people are asking the question, okay, well, how how should we approach this? How should we even think about this? And I think we now there’s sort of two different trends that are happening at the same time that sometimes get conflated. So so one is obviously the models are getting a lot better and more and more capable. What’s happening is you know if you look at the most frontier capabilities and then maybe the set of models that are just behind them or a little bit behind them every time we have these new generations of models we we both move the high water mark on the frontier but then the set of tasks that all the models can do is growing a lot and the distribution of tasks that people are like trying to accomplish their any day in their everyday lives are changing a bit actually not that much you know and so it’s like okay I’m I’m trying to build a front end you know for my marketing registration page it’s like that’s not that hard of a t right and so So I think the way I think this plays out is is as more and more the models do more of these things, the marginal returns to investing in just like having a reasonably intelligent harness that can make sure you’re using the right model for the right job in the right moment uh goes up a lot. And both for individual developers who you know they want a fast answer, they want a correct answer and they don’t want you know they don’t want any necessary waste or even really like the cognitive overhead of having to decide every time you interact with a coding agent, oh like what model should I use for this one? I I was going to ask, do you let users of Devon choose what models they use? For Devon desktop, which is uh what we rebranded Windsurf relatively recently, and for our CLI, we do. People developers like the individual control when they’re working with local agents. Um for our for our cloud agent, we don’t. It’s a you know, we’ll have like Devon infusion, which we talked about, which is the sort of um frontier performance with cost optimization option. We’ll have like a more affordable agent. We have, you know, the like the max or ultra agent. So we have this like sense of kind of tearing but I actually think it’s like a UX bug in the fullness of time for people to have to think about this for their cloud agents. And you know, we saw this early days. We ran some tests where we let people pick the models and then two things happened where one is sometimes people would then give a test to Devon and it wouldn’t work and they would complain Devon was so dumb here. And we looked at well why did you pick this model? But actually that’s kind of more our fault I think than the user’s fault. And then the same thing would happen where they would run tests like oh this is like really expensive. Uh well why did you use this very expensive uh thing? And so so I think um this sort of natural equilibrium is people want the best performance uh for the best price without compromising on correctness, right? And so folks who are building agents I think have in some sense like a user experience responsibility as well as you know just building good price performant products to try to do as much of that optimization as possible. And that’s what we recently put out with Dev Infusion, which is a sort of next generation of our own harness designed for frontier level capabilities, but with maximum price performance. And so what we found is like by being a little bit more clever about some of the routing and the decisions of like what you’re doing with model and and thinking about this in a very cache aware way, uh, we can get about 35%, you know, better price performance with actually a slight a slight increase in quality. Um, and I think that’s sort of where a lot of this needs to go for the technology to be continue to adopt it at the at the rates being adopted. So, talking about that like a little bit more because there’s there’s there’s routing and then there’s also open router launched open router fusion which is not routing, it’s running on multiple models in parallel and then combining things back together. So, when you guys have Devon infusion, is it is it routing? Is it like running multiple things? Yeah. So, it’s it’s doing both. So, so we we we put out a technical blog post that shares a little bit more detail on our implementation. Um but one core component of it is this idea of a sidekick agent. And so you know what we’ll do is we’ll have the sort of frontier quality model executing on the task and then in parallel we’ll have a more price performant model execute on the task and then there will be some decision- making that the the frontier quality model has to do of okay well when do I delegate to my sidekick but having them both work on the same task in parallel it lets you make sure that there’s still context for both of these agents and we try to share a lot of context too and we do you know writes to the file system frequently if there’s if there’s things that are exploding out of context But having basically both work in parallel then frontier model knowing ah okay I can pass this off and actually as the models get smarter the most frontier models get smarter and smarter they’re one of the key skills we see is they’re like way better at delegating to like fable is very good at knowing ah this task I can delegate to a dumber model and it’s going to be okay which kind of makes sense if you think of human career progression also like one of the aspects of of growing in your own career is learning how to delegate tasks and how to like do the highest leverage task yourself. So we kind of see the same technical pattern in in in the models. You guys have your own set of models as well. SWE 1.6 I think is the most recent one. Why did you guys train that? What do you see people using it for? Yeah, this is a great question. So sometimes people ask us, you know, why even RL or post-train your own models at R? Like aren’t the next models from the Frontier Labs just going to be better and uh and better and better and definitely and we get super excited every time new Frontier Lab models come out that are better at um their task. But I think there’s two important reasons for us to be spending a lot of energy on our own RL and post training for our own models. So the first is that you actually can deliver frontier capabilities at any given moment in time through greater specialization. And I’ll give you a recent example which is we shipped a product uh we shipped a product relatively recent called Devon review which you guys are uh are using in really interesting ways inside inside lang chain. Yes. To deliver Dev and Review, we wanted to not just have sort of a human interface for understanding diffs and kind of groing large amounts of AI generated code quickly, but also be able to run quick static analysis and lightweight um kind of model driven analysis on are there bugs in this code? Are there security vulnerabilities in this code? Are there things that we should be sort of linting and automatically checking? We can use um frontier models for that. Uh we can use cheaper models, but then uh you get really tough trade-offs on price performance curve. you know, you don’t want to have to spend tons of money just by virtue of having created a new PR. And that’s the exact type of problem where very specialized RL and post training can lead you to producing an extremely price performant model for a very specialized task. And we know that that model has a halflife that’s going to expire. And that’s fine because by the time it expires, we’ll be very happy and we’ll be working on the next specialized model for the next workload. And so I think a lot of kind of building an AI startup right now is is being very willing to think in these like 3 to six month increments of okay given this state of frontier capabilities today what is the most differentiated set of new product experiences I can build through my own specialization that’s part one the second reason we work on it’s really helpful is it does go back to price performance which is as more and more models are capable of doing more and more things how can we continue to deliver you know the best possible price performance for our customers um Well, it starts with if any model can do it well, um, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, Su 1.6 is actually the most popular model in Dev and Desktop by number of tokens consumed. It’s about Opus 4.6 level um, and we’ll have more to say on on new models uh, on new models coming out soon. there’s going to continue to be this like frontier of capabilities that use the most uh frontier models for and startups not just us I think more startups should be considering how do I get how do I make my own models that are specialized for my domain because more and more of the tasks in my domain are going to be doable by any model how many of these specialized models do you guys have at any given point in time is it one for review one for coding and those are separate ones or yeah or order of magnitude it ranges yeah I would say it’s like single digits number of specialized models because because part of our own focus as a company is on software engineering and so you know having like the sweet 1.6 series that model series is something we can we can drop in and use at a lot of different places. So we don’t need to manage like 50 or 100 different of these super specialized models but for the areas of our product that we think is going to make the biggest impact uh not just on performance by or like kind of quality it could be literally on like latency and speed and and and you know affordability we want to we want to make sure we continue to invest there and you guys optimized a lot for latency with suite 1.6 right? Yeah, one thing that was really fun um uh working on SW we were the first kind of western firm to deploy Cerebrris at scale you know when when we tested their their their trips we were really confused because we were like this is really good you know like it’s it’s it operates a unique point on the paro curve of price and throughput and you know we were getting about 950 tokens per second on our own models which was many x faster than what we could get on on GPUs for the same size of model and it was a little more expensive for us to serve. They were great partners and they enabled us to ship a product experience that basically wouldn’t have been possible otherwise and now serious is more popular you know now you know opens us for for spark I think more people are considering them and but we want to constantly be basically pushing the frontier of well what’s the next most possible thing that wasn’t possible previously um and yeah doing our own RL and post training helps with that what do you think the next most possible thing that isn’t possible right now will be as far as like model serving I mean I’m really excited about you know the sort of uh continue generation of very very specialized inference as uh the architectures that we use for transformers have gotten a little bit more standardized and more predictable you know the people designing chips uh you can make a lot more specific assumptions uh that enable big really big meaningful technical uh gains I’ll give you one funny example from when I was at Tesla um so Tesla we built our own chips in house for inference on autopilot and it was a big part of why autopilot could work uh as performantly as it did and I remember we had uh I was like a deep learning researcher One of my jobs was like interfacing with the silicon design team to make sure that the next generations of silicon supported our needs. And you know a lot of the bottleneck in in chip design it’s actually thermals. It’s like how hot can you run without melting the chip. Uh that’s the amount of power you can put in. And then even within the amount of power you you put in a lot of that power is consumed by memory movement as opposed to as opposed to flops. And so we’re trying to be really efficient of okay like what are the things that we don’t actually need on this trip that we could take away so we could really run at max throughput for the stuff we did care about. And we had a big debate one day on whether the chip actually needed to support division because they were like well you know if we don’t have to support division we can make this way faster. And if you look at at that time it was convolutional neural networks. The layers in a convolutional neural network at inference time none of them actually needed division. You know for for normalization we were using batch normalization which has a division step but uh you know you could actually fuse the batchorm uh weights and biases into the uh into the conf weights. So I think we’re going to have crazy levels of specialization because there’s going to be so much inference demand and that’s going to unlock just like super exciting uh levels of of performance. And so when we you know we started off talking about all these things that kind of like made coding agents and Devon like way better over the past two and a half years all this infrastructure stuff at the chip layer. I’m presuming that’s part of it and and there’s probably a lot more stuff to go in in that capacity. Every layer of the stack we try to think about every single layer of the stack. And one of the layers I’m most excited about right now it’s how deeply can we integrate with the rest of the team and company you know organizational context and way of doing things. So I’ll tell you like the biggest internal difference I feel using uh you know deving and cognition versus a few months ago is we’ve gone really hard on this concept of automations. So how can you wire up your agents so that they’re proactive not reactive. And you know in steady state it’s probably going to be the case that most of the engineering work at a company it’s done proactively autonomously by default by your agents and humans are still going to be making the decisions but they’re really going to be driving the new bets and well what are the sort of change my company trajectory you know technical directions to take on versus you got some user feedback there’s a bug report uh you know there’s some crash needs to be investigated all that should be handled proactively by agents and so you know we have Devon wired up to a bunch of our Slack channels triaging every message saying, “Is that something I should chime in on? Is it something I should investigate?” And the the sort of role of being a human engineer, it it’s just getting like crazy leverage. Um, and so you have like all these great suggestions of fixes that need to be applied. Oh yeah, that looks good. I want you to change that here. And in particular, having the having the agent have the context of who’s responsible for what in, you know, in the codebase in the company, so it knows, oh, you know, Harrison, you should review this change or I should review that change. Um, it’s really it’s really quite exciting. I’d love to hear how you guys think about UX because you started off, I think, and made kind of like the Slack coding experience. Like I I think I think of you guys when I think of that. You mentioned uh you you guys teamed up with Windsurf. You also have Devon CLI. You’re now talking about kind of like a different type of UX where you’re not kicking things off, but it’s coming to you. How do you see like the UX of coding agents over time and in the future evolving? Like where will humans be in the loop and what will working with these things look like? Yeah. So I think individual developers working with coding agents are going to ask themselves the question of how do I create self-driving software? How can I do sort of you know fixed cost upfront work for continuous productivity benefits and gains? Is this what the term software factory means? I think some people you said that to me it’s sort of the wrong term because when I think of a a factory I think of like the same thing being produced you know again and again versus the whole magic of of coding agents and this like proactive shift of engineering is that it’s the opposite of that it’s bespoke it’s fully custom and with all the right context of exactly what’s going on. So, I personally have never been a fan of the term, but but generally this idea that the baseline engineering and technical function of your company is going to just be getting better and better by default because it knows what objectives it’s optimizing for. It knows the input context coming in. It knows how to react to that and it can learn from your feedback over time, I think is super powerful. And we see this within the companies we work with too where individual developers have essentially realized I can be the CTO of an army of 10,000 agents. Um, and maybe the most acute use case that’s just growing like wildfire right now for us in this regard is vulnerability remediation because you know there’s sort of this generational shift in the threat surface to a lot of the especially large firms uh given how good these models are getting at um cyber security both offense and defense. So there’s kind of a lot of urgency right now on okay let’s go patch all the vulnerabilities that we have. But how do you actually do that? Well, you know, a developer probably has to sit down and think, okay, how can I really leverage Devon to go find all the different issues and then remediate them and mass? So, you can have one person setting off, you know, like a enterprisewide API call that’s doing the equivalent of thousands of engineers worth of work. How do you do that? Right. So, for security specifically, there’s enough nuance that we’ve actually shipped pretty like vertical specific features and and product services in security. We announced something recently called Devon security swarm. One way to think about it is agentic map produce for finding and then fixing security vulnerabilities. One of the original advantages of Devon is that it was the first cloud agent. So every every session is running on its own microVM and because of that you can actually safely reproduce potential security vulnerabilities and and replicate them and then validate that they’re fixed. And so there’s a sort of a scanning phase where you have to figure out, okay, of my massive codebase that doesn’t fit in the context window of any one LLM, what are the subp parts that are actually vulnerable and you have to sort of shard it out. And then once you find those parts, how do you fix them, combine them, and aggregate them in a way that um is is technically correct and actually patching the issues. But that like that paradigm of you know, one engineer is now empowered to take on the whole company uh make something better, I think is really I think is really exciting. It’s like, you know, each person there’s nothing stopping you from just having way more output than, you know, than you could have had only recently. And so that’s like a really big change and it’s it’s like changing the skills of like the skills that are most valuable I think of being an engineer. Uh thinking much more big picture of okay like what are the most important things we have to get done? How can I do this like massive set of changes to get them done and then leave behind a system that is self-improving? Who who has those skills? And kind of what I’m getting at is like do you have to have a lot of years of expertise to understand how the whole system fits together? And if so like how will people who are just graduating now get that expertise in order to do this new skill which is now maybe the only skill that matters. Yeah. Sometimes uh people ask like oh what’s the future of being like a new grad engineer because uh how are you even going to get the experience if like the coding agents are so good at new grad engineer level tasks. Uh but I think the flip side is everyone has barely any years of experience when it comes to working with agents. So in some sense we’re kind of all starting at the same level and there’s a new skill ladder to climb which is which is how can you work with agents really well and the best the best way to do that is to just try and just learn and you know do stuff. um you know consume the the resources and the training and the reading materials but but actually just go build stuff and then see what’s working and not working and in particular keep your pul your sort of finger on the pulse of the current limits of the frontier of capabilities and constantly reassess that. I think one of the biggest mistakes people make when um you know trying to learn like a new way of working is you you sort of try something say okay that that didn’t work I guess that doesn’t work but as we know in AI every every two months every 3 months you have to continually re retest that because these systems are getting so much better I actually say that what I found both internally at cognition and just working with Devon is you actually have more relative advantage as a new grad or as a recent engineer because onboarding has never been easier. You can ask as many questions as you want, including the silly ones, to an agent that will not judge you. It will help you understand exactly what’s going on. It will point you to the right parts of the codebase to say, “Oh, yeah, like this is actually where this is implemented. This is where it’s done.” And so, we’re seeing the opposite where people who would never have called themselves engineers not that long ago are steering the creation of software in ways that are totally crazy. you know, we started focused on professional software engineers, and that’s still, I think, the main users of Davin, but so many people are now using it to to make things um that we didn’t think possible, even though, you know, that wasn’t our original focus. You mentioned a few of these like specialized versions of Devon. So, Devon review, Devon security swarm. I think there’s also this idea that coding agents are general purpose and can be used for everything. How do you guys think about what should be like a special version of Devon versus just hey just go use normal Devon for that? Yeah, so we debate this internally a lot. Um I think one of the original kind of like founding thesis of the company was that yeah if you solve code you can kind of solve a lot of other things. We are seeing this I think play out in in real time both in terms of the popularity of coding agents and in you know the revenue numbers of people working on code and really the whole adoption profile of of of code. We don’t want to overengineer something that then the next version of Devon is going to make you know completely outdated and useless. At the same time part of the role we serve in the ecosystem is to be the independent agent lab. So how do we take the best of every underlying model and then go really solve our users problems end to end like in the very specific details of what they’re working on. So we’re still super focused on software engineering as the core of what we do, but more and more things touch software engineering. And so, you know, one of the reasons we did this like specialization push on security is that we just looked at the distribution of how people were using Devon and realize, hey, that’s like one of the most popular uses of Devon. So, how do we double down on that? And I think in general, like my sort of personal philosophy on product development is you want to be allocating sort of some portion of your time to fixing the the frictions and the paper cuts that your user experience. And then you also want to be allocating some portion of your time on doubling down on positive surprises. What are the ways that people are using Devon that we didn’t even think of? And then how can we double down on that to make their experience even better? And a lot of times doubling down is hey let’s like try to just improve general agent capabilities. But other times it’s okay what are the more specific interfaces um that we should build to to make it work well. You mentioned being an agent lab. There’s also these model labs which you are competing with at times but also using a lot of their models at other times. How do you think about that and what does that landscape look like over the past few years? Yeah, so we both like have really deep collaborations with the model labs and you know kind of use each other’s stuff and Frontier models are a really important part of of Devon’s own delivery and then there’s definitely I think model labs not just in coding but in all of the application domains are starting to get into the application layer too. I think for us um we’ve learned a few things you know working in that space. One is that there is actually a pretty interesting structural differentiation of where we sit in the ecosystem versus anyone model lab which is whether you’re an individual developer or the CIO of a large enterprise. Um the only thing you can be really sure about is that the underlying models are going to keep changing and it’s pretty hard to predict in advance who’s going to have the best model, who’s going to have the best price performance model and so on. And I think feedback we pretty consistently get from you know the individual users and the leaders and um you know the CTO’s that we work with is that it’s it’s really nice to have some sort of decoupling between the agents and the models so that as the models keep getting better uh you’re not stuck on you know the wrong tooling because now your model is is no longer the best. you you want to be in a situation where every day when you wake up you’re very happy to hear that a new model has come out that’s even better than before regardless of of who came out. I think the second big lesson is just uh focus on the last mile of complexity and problem solving that you know is most important for the customers. So we were pretty early in building you know forward deployed engineering team for example and today our for deployed engineering team it’s we have more for deployed engineers than um non-forward deployed engineers at cognition and so you know by the numbers we’re actually mostly working with customers because the core product is is devon working on itself most of the time. What does forward deployed engineering mean for you guys because I feel like everyone’s got forward deployed engineers and sometimes they mean different things. So for us it means folks who are very technical but also good at understanding customer and business problems and can go on-site with customers to help solve big outcomes. The way we work with the large enterprise, it’s very different from how someone signs up on Devon.ai and uses the product. It’s often um pointed at a specific objective to go solve. uh it’s not just about here’s the tooling you’re using but it’s hey you know we have got to uh move on to this new system by the end of the year or we’re going to have a lot of problems and our current schedule has this taking four years how do we do this you know much faster and better and part of I think the appeal for individual engineers to be for deployed right now is that you’re kind of flexing a lot more muscles as as an individual engineer where you’re you’re still in the technical weeds of what’s going on but you can be that that CTO of an army of a thousand agents and like work on your own skills to just sort of master this new way of working. Um, and that’s like really valuable for us but also really valuable for for our customers. What’s the right profile of someone who wants to be a for deployed engineer at cognition? We have had a lot of success with uh ex-founders in particular. So people who are very comfortable in ambiguous kind of problem areas and you know can do the technical parts but also can understand customer pain points talk to them and and really have very high internal agency to recognize oh yeah that’s that’s the problem we should go solve which is another big thing we’ve learned in the process of deploying Devon more and more widely you know you can use coding agents for almost anything and so one of the big questions is well what’s the most impactful thing or set of things to get started on and helping steer that is I think one the things that makes like a great forward deployed engineer so great. We talked about this a little bit earlier, but I think costs for coding agents are becoming something that a lot of companies are paying attention to. And the other part of that is kind of like the return that you get for everything that you’re spending on coding agents. And quantifying ROI is really hard. I think you guys have done maybe one of the best jobs at it or at least maybe even just one of the only people I’ve seen actually try to take this on in a general way. you wrote a blog post on like estimating productivity and you have some productivity guarantee. Could you talk about how you guys think about that? Yeah, we think about this all the time. So, first on the productivity side, um what’s the problem? The the problem right now is that a lot of people in the industry have a strong incentive to get customers to token max. You know, let’s build leaderboards of usage and kind of have things go up. And the incentives are actually quite fundamentally misaligned, I think. And at some point, you know, the bill comes due. And I think we’re starting to see that now. uh at a lot of enter customers where they are they have been token maxing they’ve been using a lot and now suddenly they’re spending in the tens or hundreds of millions of dollars or even more in some cases a year and and then the the leadership is asking hang what what do we get for this so there’s kind of this reset happening going back also to your earlier question of like where do we sit in the ecosystem as the independent agent lab I think one of the side effects of being independent is that structurally we’re the most incentive aligned with customers to not just focus on oh you got to drive usage of this model but really focus on well what’s going to be the value you’re going to get and let’s not even work on the stuff that’s going to be low value. Um so that was kind of one thing we focused on early on and we’ve always tried to work closely directly with customers to just help them estimate and then quantify their own ROI uh of using of using Devon and and AI generally. Um but part of that’s a manual process. It’s literally sitting down and saying okay well what are the most important things for the company right now and how do we you know how could we move the needle together if we worked on them closely. What we were looking to do is answer a sort of a more scalable question which is how could we make this happen in an automated way so everyone can benefit uh whether you know you’re a big enterprise or an individual developer using Devon how can everyone benefit from that and so that immediately imposes one design constraint which is okay how can you automatically estimate the productivity or ROI or value and I think our our conclusion was at least today in 2026 it’s very hard to automatically estimate ROI because you you don’t have the full business context of oh you know this is going to drive your, you know, your revenue or profits up this much or this new product is worth worth this much to you. But we could do one level below that which is productive engineering output. So there’s a difference between I’m using an agent, I’m getting a lot of tokens out or a lot of code out and oh this was actually valuable time savings for for me. And so to measure that, we actually worked with a bunch of our of our customers to collect a data set of a bunch of Devon sessions where we scored them together with the customer to say, okay, first of all, was this actually productive or not? And some rules you can do automatically. So for example, if you create a Devon session and that creates a PR and it never gets merged, we call that an unproductive session. Now in reality, like maybe it was still actually helpful for you, but we want to earn conservative. So we say, okay, that’s just blanket unproductive session. If you merge a PR production uh then we would say okay that’s productive. If it didn’t result in any PR because sometimes people are using Devon for data analysis or other workloads then we do an ML driven classification based on kind of a manually tagged data set in collaboration with customers. So the result of all that is we essentially have an evaluator agent that can take a session and say was it productive or not part one. Part two is if it was productive, how many hours of work did that actually save you? And again, manual data collection grind to figure out that question where we surveyed a bunch of folks. We worked closely with them to see, hey, you’re still using AI maybe even if not dev how so like what’s the delta there? And so we collected this data set of essentially engineering hours estimates for their equivalent Devon session. And by combining these things together, we can now automatically apply for everything you run from Devon um a score of was it productive or not? And if it was productive, how many hours did this save you? And that’s really interesting visibility for individuals, for teams, and for companies. And when we ran the numbers, we realized that as we were hoping, people are saving a lot of time uh wi-i with Devon. Uh but they were saving so much time that it gave us the confidence to actually financially underwrite it and go to our customers and say we are actually going to make a $10 million productivity guarantee which is if you’re paying for Devon and people are in fact just wasting usage and they’re not getting value. You know if you end up paying us more than the like engineering hours of value that you would expect in dollar terms we’re just going to refund the difference up to $10 million. And when we debated this internally, I got a lot of push back actually because, you know, some people were saying, “Wait, this is like a big risk. We’re signing up for a potential liability.” Like, what if someone just like makes an API call and they just run debit in a loop and it totally burns, you know, millions of dollars. And the reality is, yeah, like if you did that, then we would we would be on the hook and we would lose money. So, let’s go add some guardrails in the product to prevent you from being able to do that. It’s like, how do we align our incentives with the customer’s incentives? So, we now have really good cost controls. We have the ability for admins to say like this is the overall budget I want to allocate. Help me be smart about like who should get you know more compute, who’s still learning and needs kind of a tighter a tighter leash. And I think it’s it’s all about like aligning the incentives with customers so that they’re actually getting value for what they’re using. I was going to ask about that last thing you mentioned which basically like how do people use this info and do they do they use it at a team level? Do they use it at an individual level? I think intuitively inside lang chain if you asked our VP of engineering I think he’d probably say that more senior engineers he trusts to use these things more and more junior engineers it’s easy to mess up how you use these things it’s easy to accidentally spend stuff but I’m curious yeah like how how do people use these insights I think that my favorite way and one of the most popular ways of using these insights it’s really for learning and like professional development because you can look at you know two teams you can say oh wow this team is actually you know really economically efficient with their use they’re getting tons of of productive output for what they’re putting in. And this team is kind of blowing up their cost a little bit. Well, it’s not like they’re maliciously intending to do that. They’re just maybe using the tools differently and maybe they don’t realize that they’re doing some things that are really inefficient. So, we’ve tried to then now that we have these dashboards in where you can see, oh yeah, how is my usage compared to, you know, other folks and where do I sit and like what can I be doing better? We’re trying to push more of that coaching and that like professional development sort of into the product itself where you can like learn from Devon. demo will actually tell you, Harrison, that was a bad prompt, you know, like like like this is like way under specified, you know, you should fix this. Like here’s here’s some tips on how to be better. And so we’re trying to put more of that training into the product so people can kind of continue to level up with the capabilities because it’s actually really hard for anyone to keep up with how fast everything is moving and things that you didn’t expect would be possible a few months ago are possible now or things that you’re doing today as like a workaround for limitations actually you shouldn’t do tomorrow. So we we’ve tried to move as much of it in product as possible and I think people are actually using the dashboards the most is like literally seeing okay team by team individual by individual how how can I learn how to be more productive more efficient and just really master this new way of working. Thanks for listening to Max Agency. If you liked this episode leave a review and subscribe. Send feedback or questions to max agency langchain.dev. We want to hear from you.
Similar Articles
@LangChain: Brand new Max Agency with @cognition President @russelljkaplan + @hwchase17. YouTube: https://youtu.be/bBUotstDLdk?si=F…
In the interview, Cognition President Russell Kaplan reviews the development of coding agent Devon, discusses model intelligence saturation and cost optimization, and introduces how the new evaluation Frontier Code measures code mergeability.
@LangChain: "A lot of the bottleneck in chip design, it's actually thermals." @cognition president @russelljkaplan on the Tesla deb…
Cognition 总裁 Russell Kaplan 在播客中回顾了编码智能体 Devon 的发展,介绍新评估 Frontier Code 衡量“可合并性”,并讨论模型选择、成本与速度取代能力成为关键关注点。
@LangChain: In a real conversation, deciding when to speak takes about as much brainpower as deciding what to say. Voice agents hav…
Sierra Platform's approach to voice agents parallelizes thinking, listening, and talking to mimic human conversation, as discussed on the Max Agency podcast.
@rajeevchhajer: Key ideas I picked up at the @LangChain conference this week: "AI engineering is data science. Look at your data." — @s…
Key takeaways from the LangChain conference, including insights on AI engineering as data science, the priority of data strategy before agents, coding agents as the new substrate, and context plus memory as the new moat.
@LangChain: Our data agent now handles roughly 40x the request volume our 3 person data team could manage directly. Now, our data t…
LangChain rebuilt its data stack around an AI agent that handles ~40x the request volume of its 3-person data team, enabling self-serve analysis and shifting the team's focus to models, context, and guardrails.