2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents

(upbeat music) - Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO at Decibel Partners, and I'm joined by my co-host Zwickz for the 100th time today. - Yay, and we're so glad that everyone has followed us in this journey. How do you feel about it, 100 episodes?

- Yeah, almost two years that we've been doing this. We've had four different studios. We've had a lot of changes. You know, we used to do this lightning round when we first started that we didn't like, and we tried to change the questions. - Because every answer was cursor and perplexity.

- Yeah, exactly. I love mid-journey. It's like, do you really not like anything else? Like, what's the unique thing? And I think, yeah, we've also had a lot more research-driven content. You know, we had like Three Dow, we had Jeremy Howard, we had more folks like that. I think we want to do more of that too in the new year, like having some of the Gemini folks, both on the research and the applied side.

Yeah, but it's been a ton of fun. I think we both started, I wouldn't say as a joke, we were kind of like, oh, we should do a podcast. And I think we kind of caught the right wave, obviously. And I think your rise of the AI engineer posts just kind of give people somewhere to congregate and then the AI engineer summit.

And that's why when I look at our growth chart, it's kind of like a proxy for like the AI engineering industry as a whole, which is almost like, like, even if we don't do that much, we keep growing just because there's so many more AI engineers. Did you expect that growth or did you expect it would take longer for like the AI engineer thing to kind of like become, you know, everybody talks about it today.

- Yeah, my sign of that that we have won is that Gartner puts it at the top of the hype curve right now. So Gartner has called the peak in AI engineering. I did not expect to what level, I knew that I was correct when I called it because I did like two months of work going into that.

But I didn't know how quickly it could happen. And obviously there's a chance that I could be wrong. But I think like most people have come around to that concept. Hacker News hates it, which is a good sign, but there's enough people that have defined it, you know, GitHub, when he launched GitHub models, which is the hugging face clone, they put AI engineers in the banner and like above the fold, like in big letters.

So I think it's like kind of arrived as a meaningful and useful definition. I think people are trying to figure out where the boundaries are. I think that was a lot of the quote unquote drama that happens behind the scenes at the World's Fair in June, because I think there's a lot of doubt or questions about where ML engineering stops and AI engineering starts.

That's a useful debate to be had. In some sense, I actually anticipated that as well. So I intentionally did not put a firm definition there because most of the successful definitions are necessarily underspecified and it's actually useful to have different perspectives. And then you don't have to specify everything from the outset.

- Yeah, I was at AWS reInvent and the line to get into like the AI engineering talk, so to speak, which is, you know, applied AI and whatnot was like, there are like hundreds of people just in line to go in. I think that's kind of what enabled people, right?

Which is what you kind of talked about is like, hey, look, you don't actually need a PhD, just use the model. And then maybe we'll talk about some of the blind spots that you get as an engineer with the earlier posts that we also had on the sub stack.

But yeah, it's been a heck of a two years. - Yeah, you know, I was trying to view the conference as like NeurIPS is a thing like 16, 17,000 people and the latent space live event that we held there was 950 signups. I think the AI world, the ML world is still very much research heavy.

And that's as it should be because ML is very much in a research phase. But as we move this entire field into production, I think that ratio inverts into becoming more engineering heavy. So at least I think engineering should be on the same level, even if it's never as prestigious, like it'll always be low status because at the end of the day you're manipulating APIs or whatever, but wrapping GPTs, but there's gonna be an increasing stack and an art to doing these things well.

And I think that's what we're focusing on for the podcast, the conference and basically everything I do seems to make sense. And I think we'll talk about the trends here that apply. It's this very strange mix of like keeping on top of research while not being a researcher and then putting that research into production.

So like people always ask me like, why are you covering NeurIPS? Like, this is a ML research conference. And I'm like, well, yeah, I mean, we're not going to like understand everything or reproduce every single paper, but the stuff that is being found here is going to make its way into production at some point, you hope.

And then actually, like when I talk to the researchers, they actually get very excited because they're like, oh, you guys are actually caring about how this goes into production. And that's what they really, really want. The measure of success is previously just peer review, right, like getting sevens and eights on their academic review conferences and stuff.

Like citations is one metric, but money is a better metric. - Right. Yeah, and there were about 2,200 people on the live stream or something like that. - Yeah, 2,200 on the live stream. - So I tried my best to moderate, but it was a lot spicier in person with Jonathan and Dylan than it was in the chat on YouTube.

- I would say that I actually also created Latent Space Live in order to address flaws that are perceived in academic conferences. It's not NeurIPS specific, it's ICML, it's ICLR, it's NeurIPS. Basically, it's very sort of oriented towards the sort of PhD students market, job market, right? Like literally, basically everyone's there to advertise their research and skills and get jobs.

And then obviously all the companies go there to hire them. And I think that's great for the individual researchers, but for people going there to get info is not great because you have to read between the lines, bring a ton of context in order to understand every single paper.

So what is missing is effectively what I ended up doing, which is domain by domain, go through and recap the best of the year, survey the field. And there are, like NeurIPS had a, I think ICML had a like a position paper track, NeurIPS added a benchmarks and datasets track.

These are ways in which to address that issue. And there's always workshops as well. Every conference has a last day of workshops and stuff that provide more of an overview, but they're not specifically prompted to do so. And I think really organizing a conference is just about getting good speakers and giving them the correct prompts.

And then they will just go and do that thing and they do a very good job of it. So I think Sarah did a fantastic job with the startups prompt. I can't list everybody, but we did best of 2024 in startups, vision, open models, post-transformers, synthetic data, small models, and agents.

And then the last one was the, and then we also did a quick one on reasoning with Nathan Lambert. And then the last one, obviously, was the debate that people were very hyped about. It was very awkward. And I'm really thankful for John Frankel, basically, who stepped up to challenge Dylan, 'cause Dylan was like, "Yeah, I'll do it." But he was pro-scaling.

And I think everyone who is in AI is pro-scaling. So you need somebody who's ready to publicly say, "No, we've hit a wall." So that means you're saying Sam Altman's wrong, you're saying everyone else is wrong. It helps that this was the day before Ilya went on, went up on stage and then said, "Pre-training has hit a wall, data has hit a wall." So actually, Jonathan ended up winning, and then Ilya supported that statement.

And then Noam Brown, on the last day, further supported that statement as well. So it's kind of interesting that I think the consensus kind of going in was that we're not done scaling, like you should believe in a better lesson. And then four straight days in a row, you had Sepp Hochreiter, who is the creator of the LSTM, along with everyone's favorite OG in AI, which is Jurgen Schmidhuber.

He said that we're pre-training has hit a wall, or we've run into a different kind of wall. And then we have, you know, John Frankel, Ilya, and then Noam Brown, all saying variations of the same thing, that we have hit some kind of wall in the status quo of what pre-trained, scaling large pre-trained models has looked like, and we need a new thing.

And obviously the new thing for people is some make, either people are calling it inference time compute or test time compute. I think the collective terminology has been inference time. And I think that makes sense because test time, calling it test, meaning, has a very pre-trained bias, meaning that the only reason for running inference at all is to test your model.

That is not true. - Right. - So I quite agree that OpenAI seems to have adopted, or the community seems to have adopted this terminology of ITC instead of TTC. And that makes a lot of sense because like now we care about inference, even right down to compute optimality.

Like I actually interviewed this author who we've recovered or reviewed the Chinchilla paper. Chinchilla paper is compute optimal training, but what is not stated in there is it's pre-trained compute optimal training. And once you start caring about inference, compute optimal training, you have a different scaling law and in a way that we did not know last year.

- I wonder, because John is, he's also on the side of attention is all you need. Like he had the bet with Sasha. So I'm curious, like he doesn't believe in scaling, but he thinks the transformer. I wonder if he's still of that camp. - So he, obviously everything is nuanced and you know, I told him to play a character for this debate, right?

So he actually does. Yeah, he still believes that we can scale more. He just assumed the character to be very game for playing this debate. So even more kudos to him that he assumed a position that he didn't believe in and still won the debate. - Get wrecked, Dylan.

Do you just want to quickly run through some of these things like Sarah's presentation, just the highlights? - Yeah, we can't go through everyone's slides, but I pulled out some things as a factor of like stuff that we were going to talk about. - Yeah, and we'll publish the rest.

- Yeah, we'll publish on this feed, the best of 2024 in those domains. And hopefully people can benefit from the work that our speakers have done. But I think it's, these are just good slides. And I've been looking for sort of end of year recaps from people. The field has progressed a lot.

You know, I think the max ELO in 2023 on LMSYS used to be 1200 for LMSYS ELOs. And now everyone is at least at 1275 in their ELOs. And this is across Gemini, HGBT, Grok, O1 AI, which with their E-Large model and Enthopic, of course. It's a very, very competitive race.

There are multiple frontier labs all racing, but there is a clear tier zero frontier. And then there's like a tier one, it's like everything else. And tier zero is extremely competitive. It's effectively now three horse race between Gemini, Enthopic and OpenAI. I would say that people are still holding out a candle for XAI.

XAI, I think for some reason, because their API was very slow to roll out, it's not included in these like metrics. So it's actually quite hard to put on there. Like as someone who also does charts, XAI is continually snubbed because they don't work well with the benchmarking people.

So it's a little trivia for why XAI always gets ignored. The other thing is market share. So these are slides from Sarah and we have it up on the screen. It has gone from very heavily OpenAI. So we have some numbers and estimates. These are from RAMP, estimates of OpenAI market share in December, 2023.

And this is basically, what is it? GPT 3.5 and GT for being 95% of production traffic. And I think if you correlate that with stuff that we asked Harrison Chase on the LinkedIn episode, it was true. And then CLAWD 3 launched middle of this year. I think CLAWD 3 launched in March, CLAWD 3.5 Sonnet was in June-ish.

And you can start seeing the market share shift towards Enthopic very, very aggressively. And the more recent one is Gemini. So if I scroll down a little bit, this is an even more recent dataset. So RAMP's dataset ends in September, 2024. Gemini has basically launched a price war at the low end with Gemini Flash being basically free for personal use.

I think people don't understand the free tier. It's something like a billion tokens per day. Unless you're trying to abuse it, you cannot really exhaust your free tier on Gemini. They're really trying to get you to use it. They know they're in like third place, fourth place, depending how you count.

And so they're going after the lower tier first and then maybe the upper tier later. But yeah, Gemini Flash, according to OpenRouter, is now 50% of their OpenRouter requests. Obviously, these are the small requests. These are small, cheap requests that are mathematically going to be more. The smart ones obviously are still going to OpenAI.

But it's a very, very big shift in the market. Like basically over the course of 2023 going into 2024, OpenAI has gone from 95 market share to reasonably somewhere between 50 to 75 market share. - Yeah, I'm really curious how RAMP does the attribution to the model, if it's API, because I think it's all- - Credit card spend.

- Well, but the credit card doesn't say. Maybe when they do expenses, they upload the PDF. But yeah, the Gemini, I think, makes sense. I think that was one of my main 2024 takeaways that like the best small model companies are the large labs, which is not something I would have thought that the open source kind of like long tail would be, like the small model.

- Yeah, different sizes of small models we're talking about here, right? Like so small model here for Gemini is 8B, right? Mini, we don't know what the small model size is, but yeah, it's probably in the double digits or maybe single digits, but probably double digits. The open source community has kind of focused on the one to 3B size, maybe 0.5B, that's Moon Dream.

And that is small for you, then that's great. It makes sense that we have a range for small now, which is like maybe one to 5B. I'll even put that at the high end. And so this includes Gemma from Gemini as well, but also includes the Apple Foundation models, which are also like, I think Apple Foundation is 3B.

- Yeah, no, that's great. I mean, I think in the start, small just meant cheap. I think today small is actually a more nuanced discussion, you know, that people weren't really having before. - Yeah, we can keep going. This is a slide that I smiley disagree with Sarah. She's pointing to the scale seal leaderboard.

I think the researchers that I talked to in Europe were kind of positive on this because basically you need private test sets to prevent contamination. And scale is one of maybe three or four people this year that has really made an effort in doing a credible private test set leaderboard.

Lama 4 or 5B does well compared to Gemini and GPT-40. And I think that's good. I would say that, you know, it's good to have an open model that is that big that does well on those metrics. But anyone putting 4 or 5B in production will tell you, if you scroll down a little bit to the artificial analysis numbers, that it is very slow and very expensive to infer.

It doesn't even fit on like one node of H100s. Cerebris will be happy to tell you they can serve 4 or 5B on their super large chips. But, you know, if you need to do anything custom to it, you're still kind of constrained. So is 4 or 5B really that relevant?

Like, I think most people are basically saying that they only use 4 or 5B as a teacher model to distill down to something. Even Meta is doing it. So when Lama 3.3 launched, they only launched the 70B because they use 4 or 5B to distill the 70B. So I don't know if like open source is keeping up.

I think the open source industrial complex is very invested in telling you that the gap is narrowing. I kind of disagree. I think that the gap is widening with O1. I think there are very, very smart people trying to narrow that gap and they should. I really wish them success, but you cannot use a chart that is nearing 100 in your saturation chart.

And look, the distance between open source and closed source is narrowing. Of course it's going to narrow it because you're near 100. - Yeah, yeah, yeah. - This is stupid. But in metrics that matter, is open source narrowing? Probably not for O1 for a while. And it's really up to the open source guys to figure out if they can match O1 or not.

- I think inference time compute is bad for open source just because Doc can donate the flops at training time, but he cannot donate the flops at inference time. So it's really hard to like actually keep up on that axis. - Big business model shift. - So I don't know what that means for the GPU clouds.

I don't know what that means for the hyperscalers, but obviously the big labs have a lot of advantage because it's not a static artifact that you're putting the compute in. You're kind of doing that still, but then you're putting a lot of compute at inference too. - Yeah, yeah, yeah.

I mean, Llama4 will be reasoning oriented. We talked with Thomas Shalom. Kudos for getting that episode together. That was really nice. Well-timed. Actually, I connected with the AI meta guy at NeurIPS and we're gonna coordinate something for Llama4. - And our friend Clara Shi just joined to lead the business agent side.

So I'm sure we'll have her on in the new year. - Yeah, so my comment on the business model shift, this is super interesting. Apparently it is wide knowledge that OpenAI wanted more than $6.6 billion for their fund raise. They wanted to raise higher and they did not. What that means is basically like, it's very convenient that we're not getting GPT-5, which would have been a larger pre-train.

We should have a lot of upfront money. Instead, we're converting fixed costs into variable costs, right? And passing it on effectively to the customer. And it's so much easier to take margin there because you can directly attribute it to like, oh, you're using this more, therefore you pay more of the cost and I'll just slap a margin on there.

So like that lets you control your gross margin and like tie your spend or your sort of inference spend accordingly. And it's just really interesting to, that this change in the sort of inference paradigm has arrived exactly at the same time that the funding environment for pre-training is effectively drying up, kind of.

I feel like maybe the VCs are very in tune with research anyway. So like they would have noticed this, but it's just interesting. - Yeah, and I was looking back at our yearly recap of last year and the big thing was like the mixed trial price fights, you know?

And I think now it's almost like there's nowhere to go. Like, you know, Gemini Flash is like basically giving it away for free. So I think this is a good way for the labs to generate more revenue and pass down - It's great. - Some of the compute to the customer.

- Yeah, I think we're going to keep going. I think that $2,000 chat GPT will come. - Yeah, I know, totally. I mean, next year, the first thing I'm doing is signing up for Devon, signing up for the pro chat GPT. Just to try, I just want to see what does it look like to spend $1,000 a month on AI?

- Yes, I think if your job is at least AI content creator or VC or, you know, someone whose job it is to stay on top of things, you should already be spending like $1,000 a month on stuff. And then obviously easy to spend, hard to use. You have to actually use.

The good thing is that actually Google lets you do a lot of stuff for free now. So like deep research that they just launched uses a ton of inference and it's free while it's in preview. So you should use it. - They need to put that in Lindy. I've been using Lindy lately.

I've built a bunch of things once we had flow because I like the new thing. It's pretty good. I even did a phone call assistant. - Yeah, they just launched a new voice. - Yeah, I think once they get advanced voice mode, like capability, today it's still like speech to text, you can kind of tell.

But it's good for like reservations and things like that. So I have a meeting prepper thing. And so it's good. - Okay, I feel like we've covered a lot of stuff. Yeah, I think we will go over the individual talks in a separate episode. I don't want to take too much time with this stuff, but suffice to say that there was a lot of progress in each field.

We covered vision. Basically this is all like the audience voting for what they wanted. And then I just invited the best speaker I could find in each audience, especially agents. Graham, who I talked to at ICFL in Vienna, he is currently still number one. It's very hard to stay on top of SweetBench.

OpenHand is currently still number one on SweetBench full, which is the hardest one. He had very good thoughts on agents, which I'll highlight for people. Everyone is saying 2025 is the year of agents, just like they said last year. (laughing) But he had thoughts on like eight parts of what are the frontier problems to solve in agents.

And so I'll highlight that talk as well. - Yeah, the number six, which is the how can agents learn more about the environment has been super interesting to us as well, just to think through. Because yeah, how do you put an agent in an enterprise where most things in an enterprise have never been public?

You know, a lot of the tooling, like the code bases and things like that. So yeah, there's not-- - So just indexing and rag? - Well, yeah, but it's more like, you can't really rag things that are not documented, but people know them based on how they've been doing it.

You know? So I think there's almost this like-- - Oh, institutional knowledge. - So yeah, the boring word is kind of like a business process extraction. It's like, how do you actually understand how these things are done? - I see. - And I think today the agents that most people are building are good at following instruction, but are not as good as like extracting them from you.

So I think that will be a big unlock. - Cool. - Just to touch quickly on the Jeff Dean thing. I thought it was pretty, I mean, we'll link it in the things, but I think the main focus was like, how do you use ML to optimize the systems instead of just focusing on ML to do something else?

- Yeah, I think speculative decoding, we had, you know, Eugene from RWKB on the podcast before, like he's doing a lot of that with Federalist AI. - Everyone is, I would say it's the norm. I'm a little bit uncomfortable with how much it costs, because it does use more of the GPU per call, but because everyone is so keen on fast inference, then yeah, it makes sense.

- Exactly. Yeah, but we'll link that. Obviously Jeff is great. - Yeah, so that Jeff is, Jeff's talk was more, it wasn't focused on Gemini. I think people got the wrong impression from my tweet, is more about how Google approaches ML and uses ML to design systems and then systems feedback into the ML.

And I think this ties in with Lubna's talk on synthetic data, where it's basically the story of bootstrapping of humans and AI in AI research or AI in production. So her talk was on synthetic data, where like how much synthetic data has grown in 2024 in the pre-training side, the post-training side and the eval side.

And I think Jeff then also extended it basically to chips, to chip design. So he'd spent a lot of time talking about AlphaChip. And most of us in the audience are like, we're not working on hardware, man. Like you guys are great. TPU is great. Okay, we'll buy TPUs.

- And then there was the Ilya talk. - Yeah. - But, and then we have a essay tied to it. What Ilya saw. I don't know if we're calling them essays. What are we calling these? But post. - For me, it's just like bonus for the InSpace supporters, because I feel like they haven't been getting anything.

And then I wanted a more high-frequency way to write stuff. Like that one I wrote in an afternoon. I think basically we now have an answer to what Ilya saw. It's one year since the blip. And we know what he saw in 2014. We know what he saw in 2024.

We think we know what he sees in 2024. He gave some hints. And then we have vague indications of what he saw in 2023. So that was the. - Yeah. - Oh, and then 2016 as well, because of this lawsuit with Elon, OpenAI is publishing emails from Sam's, like his personal text messages to Siobhan, Zelis, or whatever.

So like, we have emails from Ilya saying, this is what we're seeing in OpenAI, and this is why we need to scale up GPUs. And I think it's very prescient in 2016 to write that. And so like, it is exactly like, basically his insights is him and Greg, basically kind of driving the scaling up of OpenAI, while they're still playing Dota.

They're like, no, like. - Yeah, yeah, yeah. - Like, we see the path here. - Yeah, and it's funny. Yeah, they even mentioned, you know, we can only train on 1v1 Dota. We need to train on 5v5, and that takes too many GPUs. - Yeah, and at least for me, I can speak for myself, like I didn't see the path from Dota to where we are today.

I think even maybe if you ask them, like they wouldn't necessarily draw a straight line, but. - Yeah, yeah, no, I definitely. But I think like that was like the whole idea of almost like the RL. And we talked about this with Naden on his podcast. It's like with RL, you can get very good at specific things, but then you can't really like generalize as much.

And I think the language models are like the opposite, which is like, you're gonna throw all this data at them and scale them up, but then you really need to drive them home on a specific task later on. And we'll talk about the OpenAI reinforcement, fine tuning, announcement too, and all of that.

But yeah, I think like skill is all you need. That's kind of what LEI will be remembered for. - Will be remembered for, yeah. - And I think just maybe to clarify on like the pre-training is over thing, that people love to tweet. I think the point of the talk was like, everybody we're scaling these chips, we're scaling the compute, but like the second ingredient, which is data, it's not scaling at the same rate.

So it's not necessarily pre-training is over. It's kind of like what got us here, won't get us there. In his email, he predicted like 10X growth every two years or something like that. And I think maybe now it's like, you can 10X the chips again, but- - I think it's 10X per year, was it?

I don't know. - And Moore's law is like 2X. So it's like much faster than that. And yeah, I like the fossil fuel of AI analogy. It's kind of like the little background tokens thing. And so the OpenAI reinforcement fine tuning is basically like, instead of fine tuning on data, you fine tune on a reward model.

So it's basically like, instead of being data driven, it's like task driven. And I think people have tasks to do. They don't really have a lot of data. So I'm curious to see how that changes, how many people fine tune. Because I think this is what people run into.

It's like, oh, you can fine tune Lama. And it's like, okay, where do I get the data to fine tune it on? So it's great that we're moving the thing. And then I really like, he had this chart where the brain mass and the body mass thing is basically like mammals that scale linearly by brain and body size.

And then humans kind of like broke off the slope. So it's almost like maybe the mammal slope is like the pre-training slope. And then the post-training slope is like the human one. - Yeah, I wonder what the, I mean, we'll know in 10 years, but I wonder what the y-axis is for Ilya's SSI.

We'll try to get them on. - Ilya, if you're listening, you're welcome here. Yeah, and then he had what comes next, like agent, synthetic data, inference compute. I thought all of that was like- - I don't think he was dropping any alpha there. - Yeah, yeah, yeah. Any other new reps, highlights, or?

- I think that there was comparatively a lot more work. Oh, by the way, I need to plug that my friend Yi made this like little nice- - Yeah, that was really nice. - Of like all the, she called it must read papers of 2024. So I laid out some of these at NeurIPS and it was just gone.

Like everyone just picked it up 'cause people are dying for like little guidance and visualizations of each paper. And so I thought it was really super nice that we got that. - Should we do a late in space book for each year? - I thought about it. - For each year, we should.

- Coffee table book. - Yeah. - Yeah. - Okay, put it in there, Will. - Hi, Will, by the way, we haven't introduced you. These are new, you know, you're organized, Jamie, we got Will. - You need to pull up more things. One thing I saw that, okay, there- - What this book see.

- Okay, one fun one, and then one more general one. So the fun one is this paper on agent collusion. This is a paper on steganography. This is secret collusion among AI agents, multi-agent deception via steganography. I tried to go to NeurIPS in order to find these kinds of papers because the real reason, like NeurIPS this year has a lottery system.

A lot of people actually even go and don't buy tickets because they just go and attend the side events. And then also the people who go and end up crowding around the most popular papers, which you already know and already read them before you showed up to NeurIPS. So the only reason you go there is to talk to the paper authors.

But there's like something like 10,000 other papers out there that, you know, are just people's work that they did on the air. They failed to get attention for one reason or another. And this was one of them. It was like all the way at the back. And this is a DeepMind paper that actually focuses on collusion between AI agents by hiding messages in the text that they generate.

So that's what steganography is. So a very simple example would be the first letter of every word. If you pick that out, you know, it decodes a different message than that. But something I've always emphasized is to LLMs, we read left to right. LLMs can read up, down, sideways, you know, in random character order.

And it's the same to them as it is to us. So if we were ever to get, you know, self-motivated, underlined LLMs that we're trying to collaborate to take over the planet, this would be how they do it. They spread messages among us in the messages that we generate.

And he developed a scaling law for that. So he marked, I'm showing it on screen right now, the emergence of this phenomenon. Basically, for example, for Cypher encoding, GPT-2, Lama2, mixed trial, GPT-3.5, zero capabilities, and it's sudden emergence of GPT-4. And this is the kind of Jason Wei type emergence properties that people kind of look for.

I think what made this paper stand out as well, so he developed a benchmark for steganography collusion, and he also focused on Schelling point collusion, which is very low coordination. Like for agreeing on a decoding encoding format, you kind of need to have some agreement on that. But Schelling point means like very, very low or almost no coordination.

So for example, if I ask someone, if the only message I give you is meet me in New York, and you have no idea where or when, you would probably meet me at Grand Central Station. That is, the Grand Central Station is a Schelling point, and it's probably somewhere during the day.

That is, the Schelling point of New York is Grand Central. To that extent, Schelling points for steganography are things like the common decoding methods that we talked about. It will be interesting at some point in the future when we are worried about alignment. It is not interesting today, but it's interesting that DeepMind is already thinking about this.

- Interesting. I think that's like one of the hardest things about NeurIPS, is like the long tail. - Very long tail. I found a pricing guy. I'm going to feature him on the podcast. Basically, this guy from NVIDIA worked out the optimal pricing for language models. It's basically an econometrics paper at NeurIPS, where like everyone else is talking about GPUs.

- And the guy with the GPUs is talking about pricing. - That was the sort of fun one. The broader focus I saw is that model papers at NeurIPS are kind of dead. No one really presents models anymore. It's just data sets. This is all the grad students are working on.

So like there was a data sets track, and then I was looking around, I was like, you don't need a data sets track because every paper is data sets paper. (laughing) And so data sets and benchmarks, they're kind of flip sides of the same thing. So yeah, if you're a grad student, your GPU board, you kind of work on that.

And then the sort of big model that people walk around and pick the ones that they like, and then they use it in their models. And that's kind of how it develops, I feel like. Last year, you had people like Hao Tian, who worked on Lava, which is Ticlama and AdVision.

And then obviously XAI hired him, and he added Vision to Grok. He's the Vision Grok guy. This year, I don't think there was any of those. - Yeah, what were the most popular like orals? Last year, it was like the Mixed Monarch, I think was like the most attended.

- Yeah, I need to look it up. - Yeah, I mean, if nothing comes to mind, that's also kind of like an answer in a way. But I think last year, there was a lot of interest in like furthering models and like different architectures and all of that. - I will say that I felt the oral picks this year were not very good.

Either that, or maybe it's just a highlight of how I have changed in terms of how I view papers. So like in my estimation, two of the best papers in this year for data sets or Data Comp and RefineWeb, or FineWeb. These are two actually industrially used papers, not highlighted for oral.

I think DCLM got the spotlight, FineWeb didn't even get the spotlight. So like, it's just that the picks were different. But one thing that does get a lot of play that a lot of people are debating is the role that's scheduled. This is the Schedule Free Optimizer paper from Meta, from Aaron DeFazio.

And this year in the ML community, there's been a lot of chat about shampoo, soap, all the bathroom amenities for optimizing your learning rates. And most people at the big labs who I asked about this say that it's cute, but it's not something that matters. I don't know, but it's something that was discussed and very, very popular.

- Four Wars of AI recap, maybe just quickly. Where do you want to start, Data? - Yeah, so to remind people, this is the Four Wars piece that we did as one of our earlier recaps of this year. And the belligerents are on the left, journalists, writers, artists, anyone who owns IP, basically, New York Times, Stack Overflow, Reddit, Getty, Sarah Silverman, George R.R.

Martin. Yeah, and I think this year we can add Scarlett Johansson to that side of the fence. So anyone suing open the eye, basically. I actually wanted to get a snapshot of all the lawsuits. I'm sure some lawyer can do it. That's the Data Quality War. On the right-hand side, we have the synthetic data people.

And I think we talked about Lumna's talk, really showing how much synthetic data has come along this year. I think there was a bit of a fight between scale AI and the synthetic data community, 'cause scale published a paper saying that synthetic data doesn't work. Surprise, surprise, scale is the leading vendor of non-synthetic data.

- Only cage-free annotated data is useful. - So I think there's some debate going on there, but I don't think it's much debate anymore that at least synthetic data, for the reasons that are blessed in Lumna's talk, makes sense. I don't know if you have any perspectives there. - I think, again, going back to the reinforcement, fine-tuning, I think that will change a little bit how people think about it.

I think today, people mostly use synthetic data for distillation and fine-tuning a smaller model from a larger model. I'm not super aware of how the frontier labs use it, outside of the rephrase, the web thing that Apple also did. But yeah, I think it'll be useful. I think whether or not that gets us the big next step, I think that's maybe like TBD.

I think people love talking about data because it's like a GPU-poor thing. I think synthetic data is something that people can do, so they feel more opinionated about it compared to the optimizers stuff, which is like, they don't really work on. - I think that there is an angle to the reasoning synthetic data.

So this year, we covered in a paper club the STAR series of papers. So that's STAR, QSTAR, VSTAR. It basically helps you to synthesize reasoning steps or at least distill reasoning steps from a verifier. And if you look at the OpenAI RFT API that they released or that they announced, basically, they're asking you to submit graders or they choose from a preset list of graders.

Basically, it feels like a way to create valid synthetic data for them to fine-tune the reasoning paths on. So I think that is another angle where it starts to make sense. And so it's very funny that basically all the data quality wars between, let's say the music industry or the newspaper publishing industry or the textbooks industry on the big labs, it's all of the pre-training era.

And then the new era, the reasoning era, nobody has any problem with all the reasoning. Fine-tuning, especially because it's all sort of math and science-oriented with very reasonable graders. I think the more interesting next step is how does it generalize beyond STEM? We've been using O1 for AI News for a while.

And I would say like for summarization and creative writing and instruction following, I think it's underrated. I started using O1 in our intro songs before we killed the intro songs, but it's very good at writing lyrics. You know, it can actually say like, I think one of the O1 Pro demos that Noam was showing was that, you know, you can write an entire paragraph or three paragraphs without using the letter A, right?

So like literally just anything in the sort of token, like not even token level, character level manipulation and counting and instruction following, it's very, very strong. And so no surprises when I ask it to rhyme and to create song lyrics, it's going to do that much better than in previous models.

So I think it's underrated for creative writing. - Yeah. What do you think is the rationale that they're going to have in court when they don't show you the thinking choices of O1, but then they want us to like, they're getting sued for using other publishers' data, you know, but then on their end, they're like, well, you shouldn't be using my data to then train your model.

So I'm curious to see how that kind of comes. - Yeah, I mean, O1 has many ways to punish people without taking them to court, already banned by Dance for distilling their info. And so anyone caught distilling the chain of thought will be just disallowed to continue on the API and it's fine.

It's no big deal. Like, I don't even think that's an issue at all, just because the chain of thoughts are pretty well hidden. Like, you have to work very, very hard to get it to leak. And then even when it leaks the chain of thought, you don't know if it's the real one.

So there's much less concern here. - Yeah, yeah, yeah. - The bigger concern is actually that there's not that much IP hiding behind it. That Cosine, which we talked about, we talked to him on Dev Day, can just fine tune 4.0 to beat O1. That Cloud Sonnet so far is beating O1 on coding tasks without, at least O1 preview, without being a reasoning model and same for Gemini Pro or Gemini 2.0.

So like, how much is reasoning important? How much of a moat is there in this proprietary sort of training data that they've presumably accomplished? Because like, even DeepSeek was able to do it and they had two months notice to do this, to do R1. So it's actually unclear how much moat there is.

Obviously, if you talk to the Strawberry team, they'll be like, yeah, I mean, we spent the last two years doing this. So we don't know. And it's going to be interesting because there'll be a lot of noise from people who say they have inference time compute and actually don't because they just have fancy chain of thought.

And then there's other people who actually do have very good chain of thought and you will not see them on the same level as OpenAI because OpenAI has invested a lot in building up the mythology of their team, which makes sense. Like the real answer is somewhere in between.

- Yeah, I think that's kind of like the main data more story developing. GPU poor versus GPU rich. - Yeah. - Where do you think we are? I think there was, again, going back to like the small model thing, there was like a time in which the GPU poor were kind of like the rebel faction working on like these models that were like open and small and cheap.

And I think today people don't really care as much about GPUs anymore. You also see it in the price of the GPUs. Like, you know, that market is kind of like plummeted because there's, people don't want to be, they want to be GPU free. They don't even want to be poor.

They just want to be, you know, completely without them. Yeah, how do you think about this war developing? - You can tell me about this, but like, I feel like the appetite for GPU rich startups, like the, you know, the funding plan is, we will raise 60 million and we'll give 50 of that to Nvidia.

That is gone, right? Like no one's pitching that. This was literally the plan, the exact plan of like, I can name like four or five startups, but you know, this time last year. So yeah, GPU rich startups gone. But I think like the GPU ultra rich, the GPU ultra high net worth is still going.

So now we're, you know, we had Leopold's essay on the trillion dollar cluster. We're not quite there yet. We have multiple labs, you know, XAI, very famously, you know, Jensen Huang praising them for being best boy number one in spinning up a hundred thousand GPU cluster in like 12 days or something.

So likewise at Meta, likewise at OpenAI, likewise at the other labs as well. So like the GPU ultra rich are going to keep doing that because I think partially it's an article of faith now that you just need it. Like you don't even know what we're going to use it for.

You just need it. And it makes sense that if, especially if we're going into more researchy territory than we are, so let's say 2020 to 2023 was let's scale big models territory because we had GPT-3 in 2020. And we were like, okay, we'll go from 1.75B to 1.8B, 1.8T.

And that was GPT-3 to GPT-4. Okay, that's done. Like, and as far as everyone is concerned, Claude, you know, Opus 3.5 is not coming out. GPT-4.5 is not coming out. And Gemini 2, like we don't have pro whatever. We've hit that wall, whatever that wall is. Maybe I'll call it like the 2 trillion parameter wall.

Like we're not going to 10 trillion. Like it's just like, no one thinks it's a good idea, at least from training costs, from amount of data, or at least the inference, like, would you pay 10X the price of GPT-4? Probably not. Like you want something else that is at least more useful.

So it makes sense that people are pivoting in terms of their inference paradigm. And so when it's more researchy, then you actually need more just general purpose compute to mess around with at the exact same time that production deployments of the previous paradigm are still ramping up pretty aggressively.

So it makes sense that the GPU rich are growing. We have now interviewed both together in Fireworks and Replicates. We haven't done any scale yet, but I think Amazon may be kind of a sleeper one, Amazon, in a sense of like they, at reInvent, I wasn't expecting them to do so well, but they are now a foundation model lab.

It's kind of interesting. I think, you know, David went over there and started just training models. - Yeah, I mean, that's the power of prepaid contracts. I think like a lot of AWS customers, you know, they do this big reserve instance contracts and now they got to use their money.

That's why so many startups get bought through the AWS marketplace. So they can kind of bundle them together and prefer pricing. - Okay, so maybe GPU super rich, doing very well. GPU middle-class, dead. And then GPU poor. - I mean, my thing is like, everybody should just be GPU rich.

There shouldn't really be, even the GPU poorest is like, does it really make sense to be GPU poor? Like if you're GPU poor, you should just use the cloud. - Yes. - You know, and I think there might be a future once we kind of like figure out what the size and shape of these models is, where like the tiny box and these things come to fruition where like you can be GPU poor at home.

But I think today it's like, why are you working so hard till I get these models to run on like very small clusters where it's like, it's so cheap to run the- - Yeah, yeah. - You know? - Yeah, yeah. I think mostly people think it's cool. People think it's a stepping stone to scaling up.

So they aspire to be GPU rich one day and they are working on new methods. Like news research, like probably the most deep tech thing they've done this year is distro or whatever the new name is. There's a lot of interest in heterogeneous computing, distributed computing. I tend generally to de-emphasize that historically, but it may be coming to a time where it is starting to be relevant.

I don't know. You know, SF Compute launched their compute marketplace this year. And like, who's really using that? Like it's a bunch of small clusters, disparate types of compute. And if you can make that useful, then that will be very beneficial to the broader community, but maybe still not the source of frontier models.

- Yeah. - It's just going to be a second tier of compute that is unlocked for people and that's fine. But yeah, I mean, I think this year, I would say a lot more on device. We are, I now have Apple intelligence on my phone, doesn't do anything apart from summarize my notifications, but still not bad.

Like it's multi-modal. - Yeah. The notification summaries are so-and-so in my experience. - Yeah, but they add juice to life. And then Chrome Nano, Gemini Nano is coming out in Chrome. They're still feature flagged, but you can try it now if you use the alpha. And so like, I think like, we're getting the sort of GPU poor version of a lot of these things coming out.

And I think it's like quite useful. Like Windows as well, rolling out RWKB in sort of every Windows department. It's super cool. And I think the last thing that I never put in this GPU poor war that I think I should now, is the number of startups that are GPU poor, but still scaling very well, as sort of wrappers on top of either a foundation model lab or a GPU cloud.

A GPU cloud, it would be Suno. Suno, RAMP has rated as one of the top ranked, fastest growing startups of the year. I think the last public number is like zero to 20 million this year in ARR. And Suno runs on Modo. So Suno itself is not GPU rich, but they're just doing the training on Modo, who we've also talked to on the podcast.

The other one would be Bolt. Straight cloud wrapper. (laughs) And again, now they've announced 20 million ARR, which is another step up from our 8 million that we put on the title. So yeah, I mean, it's crazy that all these GPU poors are finding a way, while the GPU riches are also finding a way.

And then the only failures, I kind of call this the GPU smiling curve, where the edges do well, 'cause you're either close to the machines and you're like number one on the machines, or you're like close to the customers and you're number one on the customer side. And the people who are in the middle inflection, character, didn't do that great.

I think character did the best of all of them. Like you have a note in here that we apparently said that character's price tag was 1B. - Yeah. - Did I say that? - Yeah. You said Google should just buy them for 1B. I thought it was a crazy number.

Then they paid 2.7. - I mean, for like, yeah. What do you pay for now? Like, I don't know what the beginning was like. Maybe the starting price was 1B. (laughs) I mean, whatever it was, it worked out for everybody involved. Multi-modality were, in this one, we never had text-to-video in the first version, which now is the hottest.

- Yeah, I would say it's a subset of image, but yes. - Yeah. Well, but I think at the time, it wasn't really something people were doing. And now we had VO2 just came out yesterday. Sora was released last week. - Have you tried Sora? - I've not tried Sora because the day that I tried- - Yeah, it wasn't.

I think it's generally available now. You can go to Sora.com and try it. - Yeah, and then they had the outage, which I think also played a part into it. - Small things, just very important. - What's the other model that you posted today that was on Replicate Video or OneLive?

- Yeah, very, very nondescript name, but it is from Minimax, which I think is a Chinese lab. The Chinese labs do surprisingly well at the video models. I'm not sure it's actually Chinese. I don't know, hold me up to that. Yep, China. (laughs) - No, it's good. - So, Heiluo.

Yeah, the Chinese love video. What can I say? They have a lot of training data for video, or a more relaxed regulatory environment. - Well, sure, in some way. Yeah, I don't think there's much else there. I think like, you know, on the image side, I think it's still open.

- Yeah, I mean, 11 Labs, now Unicorn. So basically, what is multi-modality war? Multi-modality war is, do you specialize in a single modality, right? Or do you have God model that does all the modalities? Right, so this is definitely still going in a sense of 11 Labs, you know, now Unicorn.

Pico Labs doing well. They launched Pico 2.0 recently. HN, I think has reached a hundred million ARR. Assembly, I don't know, but they have billboards all over the place. So I assume they're doing very, very well. So these are all specialist models and specialist startups. - And product, especially.

- Yep, and then there's the big labs who are doing the sort of all-in-one play. And here I would highlight Gemini 2 for having native image output. Have you seen the demos? - No. - Yeah, it's hard to keep up. Literally, they launched this last week. And shout out to Paige Bailey, who came to the Latent Space event to demo on the day of launch, and she wasn't prepared.

She was just like, "I'm just gonna show you." So they have voice. They have, you know, obviously image input, and then they obviously can code gen and all that. But the new one that OpenAI and Meta both have, but they haven't launched yet, is image output. So you can literally, I think their demo video was that you put in an image of a car and you ask for minor modifications to that car, they can generate you that modification exactly as you asked.

So there's no need for the stable diffusion or comfy wireflow of like mask here, and then like infill there, inpaint there, and all that stuff. This is small model nonsense. Big model people are like, "Huh, we got you everything in the transformer." This is the multimodality war, which is, do you bet on the God model?

Or do you string together a whole bunch of small models like a chump? - Yeah, I don't know, man. Yeah, that would be interesting. I mean, obviously I use Midjourney for all of our thumbnails. - Yes, still Soda. - They've been doing a ton on the product, I would say.

They launched a new Midjourney editor thing. They've been doing a ton. Because I think, yeah, the model is kind of like, maybe, you know, people say Black Forest, the Black Forest models are better than Midjourney on a pixel by pixel basis. But I think when you put it together.

- Have you tried the same problems on Black Forest? - Yes, but the problem is just like, you know, on Black Forest, it generates one image. And then it's like, you got to regenerate. You don't have all these UI things. Like what I do- - Skill issue, bro. - No, but it's like time issue, you know?

It's like on Midjourney- - Call the API four times. - No, but then there's no like variate. Like the good thing about Midjourney is like, you just go in there and you're cooking. There's a lot of stuff that just makes it really easy. And I think people underestimate that.

Like, it's not really a skill issue because I'm paying Midjourney. So it's a Black Forest skill issue because I'm not paying them, you know? - So, okay. So this is a UX thing, right? Like you understand that at least we think that Black Forest should be able to do all that stuff.

I will also shout out ReCraft. Has come out on top of the image arena that artificial analysis has done. Has apparently taken Flux's place. Is this still true? So artificial analysis is now a company. I highlighted them, I think, in one of the early AI news of the year.

And they have launched a whole bunch of arenas. So they're trying to take on LM Arena, Anastasios and crew. And they have an image arena. Oh yeah, ReCraft V3 has now beaten Flux 1.1, which is very surprising 'cause Flux and Black Forest Labs are the old stable diffusion crew who left stability after the management issues.

So ReCraft has come from nowhere to be the top image model. Very, very strange. I would also highlight that Grok has now launched Aurora, which is, it's very interesting dynamics between Grok and Black Forest Labs because Grok's images were originally launched in partnership with Black Forest Labs as a thin wrapper.

And then Grok was like, "No, we'll make our own." And so they've made their own. I don't know, there are no APIs or benchmarks about it. They just announced it. So yeah, that's the multi-modality war. I would say that so far the small model, the dedicated model people are winning because they are just focused on their tasks.

But the big model people are always catching up. And the moment I saw the Gemini 2 demo of image editing where I can put an image and just request it and it does, that's how AI should work. Not like a whole bunch of complicated steps. So it really is something.

And I think one frontier that we haven't seen this year, like obviously video has done very well and it will continue to grow. You know, when we have the release of Sora Turbo today, but at some point we'll get full Sora, or at least the Hollywood Labs will get full Sora.

We haven't seen video to audio or video synced with audio. And so the researchers that I talked to are already starting to talk about that as the next frontier. But there's still maybe like five more years of video left to actually be Sora. I would say that Gemini's approach compared to OpenAI, Gemini seems, or DeepMind's approach to video seems a lot more fully-fledged than OpenAI.

Because if you look at the ICML recap that I published that so far nobody has listened to, (laughs) that people have listened to it. It's just a different, definitely a different audience. - It's only seven hours long. - It's only seven hours. - Why are people not listening? - Basically everything in one video is good.

So DeepMind is working on Genie. They also launched Genie 2 and VideoPoet. So like they have maybe four years advantage on world modeling that OpenAI does not have. 'Cause OpenAI basically only started Diffusion Transformers last year when they hired Bill Peebles. So DeepMind has a bit of advantage here, I would say, in showing, like the reason of VO2, well, one, they cherry-picked their videos.

So obviously it looks better than Sora. But the reason I would believe that VO2, when it's fully launched, will do very well is because they have all this background work in video that they've done for years. Like last year's NeurIPS, I already was interviewing some of their video people.

I forget their model name, but for people who are dedicated fans, they can go to NeurIPS 2023 and see that paper. - And then last but not least, the LLMOS/RegOps, formerly known as RegOps War. I put the latest chart on the Brain Trust episode. I think I'm gonna separate these essays from the episode notes.

So the reason I used to do that, by the way, is 'cause I wanted to show up on Hacker News. I want the podcast to show up on Hacker News. So I always put an essay inside of there because at Hacker News, people like to read and not listen.

- So episode essays, I don't know about you, we're just doing them separately. - You say Liangchen Lama Index is still growing. - Yeah, so I looked at the PiPi stats. I don't care about stars. On PiPi, you see- - Do you wanna share your screen? - Yes. I prefer to look at actual downloads, not at stars on GitHub.

So if you look at, you know, Liangchen still growing. These are the last six months, Lama Index still growing. What I've basically seen is like things that, one, obviously these things have a commercial product. So there's like people buying this and sticking with it versus kind of hopping in between things versus, you know, for example, crew.ai, not really growing as much.

The stars are growing. If you look on GitHub, the stars are growing, but kind of like the usage is kind of like flat in the last six months. - Have they done some kind of a reorg where they did like a split of packages and now it's like a bundle of packages?

Sometimes that happens, you know. - I didn't see that. - I can see both. I can see both happening. The crew.ai is very loud, but not used. And then- - Yeah, but anyway, to me, it's just like- - Yeah, there's no split. I mean, similar with AutoGPT is like, there's still a wait list for AutoGPT to be used.

- Yeah, they're still kicking. They announced some stuff recently. - But I think that's another one. It's the fastest growing project in the history of GitHub. But I think, you know, when you maybe like run the numbers on like the value of the stars and like the value of the hype, I think in AI you see this a lot, which is like a lot of stars, a lot of interest at a rate that you didn't really see in the past in open source where nobody's running to start, you know, a NoSQL database.

It's kind of like just to be able to actually use it. - Yeah. I think one thing that's interesting here, one obviously is that in AI, you kind of get paid to promise things. - Yeah. - And then you, to deliver them, you know, people have a lot of patience.

I think that patience has come down over time. One example here is Devon, right, this year, where a lot of promise in March and then it took nine months to get to GA. But I think people are still coming around now on Devon. Devon's product has improved a little bit and even you're going to be a paying customer.

So I think something Devon-like will work. I don't know if it's Devon itself. The Auto-GPT has an interesting second layer in terms of what I think is the dynamics going on here, which is a very AI-specific layer. Over-promising, under-delivering applies to any startup. But for AI specifically, there's this promise of generality that I can do anything, right?

So Auto-GPT's initial problem was making money, like increase my net worth. And I think that means that there's a lot of broad interest from a lot of different people who are trying to do all different things on this one project. So that's why this concentrates a lot of stars.

And then obviously, because it does too much, maybe, or it's not focused enough, then it fails to deploy. So that would be my explanation for why the interest to usage ratio is so low. And the second one is obviously pure execution. Like the team needs to have a vision and execute like half the core team left right after AI Engineer Summit last year.

(laughs) That will be my explanation as to why, like this promise of generality works basically only for ChatGPT. - Right. - And maybe for this year, Notebook LM. Like sticking anything in there, it'll mostly be correct. And then for basically everyone else, it's like, you know, we will help you complete code.

We will help you with your PR reviews, like small things. - Yeah, yeah, yeah. All right, code interpreting, we talked about a bunch of times we soft announced the E2B fundraising on this podcast. Code Sandbox got acquired by Together AI last week, which they're now also going to offer as an API.

So more and more activity, which is great. Yeah, and then in the last step, two episodes ago with Bolt, we talked about the web container stuff that we're working on. I think like there's maybe the spectrum of code interpreting, which is like, you know, dedicated SDK. There's like, yeah, the models of the world, which is like, hey, we got a sandbox.

Now you just kind of run the commands and orchestrate all of that. I think this is one of the, I mean, E2B's growth has just been crazy, just because, I mean, everybody needs to run code, right? And I think now all the products and everybody's graduating to like, okay, it's not enough to just do chat.

So Perplexity, which is a E2B customers, they do all these nice charts for like finance and all these different things. It's like the products are maturing. And I think this is becoming more and more of kind of like a hair on fire problem, so to speak. So yeah, excited to see more.

And this was one that really wasn't on the radar when we first wrote "The Four Wars." - Yeah. I think mostly because I was trying to limit it to Ragnops, but I think now that the frontier has expanded in terms of the core set of tools, core set of tools would include code interpreting, like tools that every agent needs, right?

And Graham in his state of agents talk had this as well, which is kind of interesting for me, 'cause like everyone finds the same set of things. So it's basically like, everyone needs web browsing, everyone needs code interpreting, and then everyone needs some kind of memory or planning or whatever that is.

We'll discover this more over time, but I think this is what we've discovered so far. I will also call out Morphlabs for launching a time travel VM. I think that basically the statefulness of these things needs to be locked down a lot. Basically, you can't just spin up a VM, run code on it, and then kill it.

It's because sometimes you might need to time travel back, like unwind or fork to explore different paths for sort of like a tree search approach to your agent development. I would call out the newer ones, the new implementations as the emerging frontier in terms of like what people kind of are going to need for agents to do very fan out approaches to all these sort of code execution.

And then I'll also call out that I think ChatGPT Canvas with what they launched in the 12 days of shipments that they announced has surprisingly superseded Code Interpreter. Like Code Interpreter was last year's thing. And now Canvas can also write code and also run code and do more than Code Interpreter used to do.

So right now it has not killed it. So there's a toggle box for Canvas and for Code Interpreter when you create a new custom GPTs. My old thesis that custom GPTs is your roadmap for investing 'cause it's what everyone needs. So now there's a new box called Canvas that everyone has access to, but basically there's no reason why you should use Code Interpreter over Canvas.

Like Canvas has incorporated the diff mode that both Anthropic and OpenAI and Fireworks has now shipped that I think is going to be the norm for next year, that everyone needs some kind of diff mode Code Interpreter thing. Like Ader was also very early to this. Like the Ader benchmarks were also all based on diffs and Coursera as well.

- Yeah, you want to talk about memory? - Memory, you think it's not real. - Yeah, I just don't. I think most memory product today just like a summarization and extraction. - Yeah. - I don't think there is. - They're very immature. - Yeah, there's no implicit memory. It's not explicit memory of what you've written.

There's no implicit extraction of like, oh, use a node to this, use a node to this 10 times so you don't like going on hikes at 6 a.m. Like it doesn't, none of the memory products do that. They'll summarize what you say explicitly. - When you say memory products, you mean the startups that are more offering memory as a service?

- Yeah, or even like, you know, Lindy has like memories, you know, it's like based on what I say, it remembers it. So it's less about making an actual memory of my preference. It's more about what explicitly said. And I'm trying to figure out at what level that gets solved.

You know, like is it, do these memory products like the MMPTs of the world create a better way to implicitly extract preference or can that be done very well? You know, I think that's why I don't think, it's not that I don't think memory is real. I just don't think that like the approaches today are like actually memory or what you need a system to have.

- Yeah, I would actually agree with that. But I would just point it to it being immature rather than not needed. Like it's clearly something that we will want at some point. And so the people developing it now are, you know, not very good at it. And I would definitely predict that next year will be better and the year after that will be better than that.

I definitely think that last time we had the "Shouldn't You" pod with Harrison as guest host, I over-focused on LangMEM as a separate product. He has now rolled it into LangGraph as a memory service, the same API. And I think that everyone will need some kind of memory. And I think that this has distinguished itself now as a separate need from a normal rag vector database.

In fact, you will need a memory layer, whether it's on top of a vector database or not, it's up to you. A memory database and a vector database are kind of two different things. Like I've had to justify this so much, actually, that I have a draft post in the "Latent Space" dashboard that basically says like, what is the difference between memory and knowledge?

And to me, it's very clear. It's like, knowledge is about the world around you. And like, there's knowledge that you have, which is the rag corpus that your, maybe your company docs or whatever. And then there's external knowledge, which is the stuff that you Google. So you use something like Exa, whatever.

And then there's memory, which is my interactions with you over time. Both can be represented by vector databases or knowledge graphs, doesn't really matter. Time is a specifically important one in memory because you need a decay function. And then you also need like a review function. A lot of people are implementing this as sleep.

Like when you sleep, you like literally, you sort of process the day's memories and you come up with new insights that you then persist and bring into context in the future. So I feel like this is being developed. LandGraph has a version of this, ZEP. It's another one that's based on Neo4j's knowledge graph that has a version of this.

MemGPT used to have this, but I think, I feel like Letter, since it was funded by Quiet Capital, has broadened out into more of a sort of general LLMOS type startup, which I feel like there's a bunch of those now. There's All Hands and all this. - Do you think this is a LLMOS product or should it be a consumer product?

- I think it's a building block. I think every, I mean, there should, just like every consumer product is going to have a, going to eventually want a gateway, you know, for managing their requests. An ops tool, you know, that kind of stuff. Code interpreter for maybe not exposing the code, but executing code under the hood, for sure.

So it's going to want memory. It's going to want long-lived memory. So as a consumer, let's say you are a new.computer, who, you know, they've launched their own little agents, or if you're a friend.com, you're going to want to invest in memory at some point. Maybe it's not today.

Maybe you can push it off a lot further with like a million token context, but at some point you need to compress your memory and to selectively retrieve it. And then what are you going to do? You have to reinvent the whole memory stack. And these guys have been doing it for a year now.

- Yeah, to me, it's more like I want to bring the memories. It's almost like they're my memories, right? So why- - So you selectively choose the memory to bring in. - Why does every time that I go to a new product, it needs to relearn everything about me?

- Okay, you want portable memories. - Yeah, is it like a protocol? Like how does that work? - Speaking of protocols, Anthropic's model context protocol that they launched has a 300 line of code memory implementation. Very simple, very bad news for all the memory startups, but that's all you need.

And yeah, it would be nice to have a portable memory of you to ship to everyone else. Simple answer is there's no standardization for a while because everyone will experiment with their own stuff. And I think Anthropic's success with MCP suggests that basically no one else but the big labs can do it because no one else has the sway to do this.

Then that's how it's going to be. (laughing) Like, unless you have something silly, like, okay, one form of standardization basically came from Georgi Gaganov with Lama CPP, right? And that was completely open source, completely bottoms up. And that's because there's just a significant amount of work that needed to be done there.

And then people build off from there. Another form of standardization is Confit UI from Confit Anonymous. So like that kind of standardization can be done. So someone basically has to create that for the role-play community because those are the people with the longest memories. Right now, the role-play community, as far as I understand it, I've looked at Sully Tavern, I've looked at Cobalt, they only share character cards.

And there's like four or five different standardized standard versions of these character cards, but nobody has exportable memory yet. If there was anyone that developed memory first, that became a standard, it would be those guys. - Cool, I'm excited to see what people built. - Benchmarks. - Okay. - One of our favorite pet topics.

- Yeah, yeah. So basically, I just wanted to mention this briefly. Like, I think that in a year, end of year review, it's useful to remind everybody where we were. So we talked about how in LMS's ELO, everyone has gone up and it's a very close race. And I think benchmarks as well.

I was looking at the OpenAI live stream today when they introduced O1 API with structured output and everything. And the benchmarks they're talking about are like completely different than the benchmarks that we were talking about this time last year. This time last year, we were still talking about MMLU.

Little bit of, there's still like GSMAK. There's stuff that's basically in V1 of the Hugging Face Open Models leaderboard, right? We talked to Clementine about the decisions that she made to upgrade to V2. I will also say LMS's now LM Arena also has emerged this year as the leading like battlegrounds between the big frontier labs.

But also, we have also seen like the emergence of SweBench, LiveBench, MMLU Pro, and AIMI. AIMI specifically for O1. It will be interesting to see like top most cited benchmarks of the year from 2020 to 2021, two, three, four, and then going to five. And you can see what has been saturated and solves and what people care about now.

And so now people care a lot about frontier math coding, right? There's literally a benchmark called frontier math, which I spent a bit of time talking about at NeurIPS. There's AIMI, there's LiveBench, there's MMLU Pro, and there's SweBench. I feel like this is good. And then there was another one, this time last year, it was GPQA.

I will put math and GPQA here as sort of top benchmarks of last year. At NeurIPS, GPQA was declared dead, which is very sad. People were still talking about GPQA Diamond. So literally the name of GPQA is called Google Proof Question Answering. So it's supposed to be resistant to saturation for a while.

And Noam Brown said that GPQA was dead. So now we only care about SweBench, LiveBench, MMLU Pro, AIMI, and even SweBench. We don't care about SweBench proper. We care about SweBench verified. We care about the SweBench multimodal. And then we also care about the new Kowinski Prize from Andy Kowinski, which is the guy that we talked to yesterday, who has launched a similar sort of ArcAGI attempt on a SweBench type metric, which arguably is a bit more useful.

OpenAI also has MLEBench, which is more tracking sort of ML research and bootstrapping, which arguably like this is the key metric that is most relevant for the Frontier Labs, which is when the researchers can automate their own jobs. So that is a kink in the acceleration curve if we were ever to reach that.

- Yeah, that makes sense. I mean, I'm curious. I think Dylan at the debate, he said SweBench 80% was like a soap for end of next year. - As a kind of like, you know, watermark. - Yeah, yeah, yeah. And keeping when we started the year at 13%. - Yeah, exactly.

- And so now we're about 50. OpenHance is around there. And yeah, 80 sounds fine. Kowinski Prize is 90. - Yeah. And then as we get to 100 and the open source catches up. - Oh yeah, magically gonna close the gap between the closed source and open source. So basically I think my advice to people is keep track of the slow cooking of benchmark language because the labs that are not that frontier will keep measuring themselves on last year's benchmarks.

And then the labs that are actually frontier will tell you about benchmarks you've never heard of. And you'll be like, oh, like, okay, there's new territory to go on. That will be the quick tip there. Yeah, maybe I won't belabor this point too much. I was also saying maybe VL has introduced some new video benchmarks, right?

Like basically every new frontier capabilities and this is the next section that we're gonna go into introduces new benchmarks. We'll also briefly talk about Ruler as like the new sort of, last year we was like needle in a haystack and Ruler is basically a multi-dimensional needle in a haystack.

- Yeah, we'll link on the episodes. - Yeah, this is like a review of all the episodes that we've done, which I have in my head. This is one of the slides that I did on my Dev Day talk. So we're moving on from benchmarks to capabilities. And I think I have a useful categorization that I've been trying to sell.

I'd be curious on your feedback or edits. I think there's basically like, I kind of like the thought spot model of what's mature, what's emerging, what's frontier, what's niche. So mature is like stuff that you can just rely on in production, it's solved, everyone has it. So what's solved is general knowledge, MMLU.

And then what's solved is kind of long context. Everyone has 128K. Today, O1 announced 200K, which is very expensive. I don't know what the price is there. What's solved, kind of solved is RAG. There's like 18 different kinds of RAG, but it's mostly solved. Batch transcription, I would say Whisper is something that you should be using on as much as possible.

And then code generation, and kind of solved. There's different tiers of code generation and I really need to split out. - Yeah, yeah, yeah. - Single line autocomplete versus multi-file generation. I think that is definitely emerging. So on the emerging side, tool use, I would still consider emerging, maybe more mature already, but they only launched for short output this year.

- Yeah, yeah, yeah. I think emerging is fine. - Vision language models. Everyone has vision now, I think, including O1. So this is clear. A subset of vision is PDF parsing. And I think the community is very excited about the work being done with CodePoly and CodeQuin. - What's for you the break point for vision to go to mature?

- I think it's basically now. This is maybe two months old. - Yeah, yeah, yeah. - NVIDIA, most valuable company in the world. Also, I think this was in June. Then also they surprised a lot on the upside for their Q3 earnings. I think the quote that I highlighted in AI News was that it is the best, like Blackwell is the best selling series in the history of the company.

And they're sold. I mean, obviously they're always sold out, but for him to make that statement, I think it's another indication that the transition from the H to the B series is going to go very well. - Yeah. I mean, if you had just bought NVIDIA and tried to beat the game out, that would be insane.

- Yeah. Which one more, NVIDIA or Bitcoin? I think NVIDIA. - I think in gains, yeah. - Well, I think the question is like, people ask me like, what's the reason to not invest in NVIDIA? I think it's really just like, they have committed to this. They went for a two year cycle to one year cycle, right?

And so it takes one misstep to delay. You know, like there have been delays in the past and like when delays happen, they're typically very good buying opportunities. Anyway. - Hey, this is Fiks from the editing room. I actually just realized that we lost about 15 minutes of audio and video that was in the episode that we shipped and I'm just cutting it back in and rerecording.

We don't have time to rerecord before the end of the year. It's December 31st already. So I'm just going to do my best to recover what we have and then sort of segue you in nicely to the end. So our plan was basically to cover like what we felt was emerging capabilities, frontier capabilities, and niche capabilities.

So emerging would be tool use, vision language models, which you just heard, real-time transcription, which I have on one of our upcoming episodes to be, as well as you can try it in Whisper WebGPU, which is amazing. I think diarization capabilities are also maturing as well, but still way too hard to do properly.

Like we had to do a lot of stuff for the latent space transcripts to come out right. I think maybe, you know, Dwarkesh recently has been talking about how he's using Gemini 2.0 Flash to do it. And I think that might be a good effort, a good way to do it.

And especially if there's crosstalk involved, that might be really good, but there might be other reasons to use normal diarization models as well. Specifically PyAnode. Text and image, we talked about a lot, so I'm just going to skip. And then we go to frontier, which, you know, I think like basically I would say is on the horizon, but not quite ready for broad usage.

Like it's, you know, interesting to show off to people, but like we haven't really figured out how like the daily use, the large amount of money is going to be made on long inference, on real-time interruptive, sort of real-time API voice mode things, on on-device models, as well as all other modalities.

And then niche models, niche capabilities. I always say like base models are very underrated. People always love talking to base models as well. And we're increasingly getting less access to them. It's quite possible, I think, you know, Sam Altman for 2025 was like asking about what people want him to ship or what people want him to open source.

And people really want GPT-3 base. We may get it. We may get it. It's just for historical interest, but you know, at this point, but we may get it. Like it's definitely not a significant IP anymore for him. So we'll see, you know, I think OpenAI has a lot more things to worry about than shipping base models, but it will be a very, very nice things to do for the community.

State space models as well. I would say like the hype for state space models this year, even though, you know, the post-transformers talk at Lean Space Live was extremely hyped and very well attended and watched. I would say like, it feels like a step down this year. I don't know why.

It seems like things are scaling out in state space models and RWBKBs. So Cartesia, I think is doing extremely well. We use them for a bunch of stuff, especially for small talks and some of our sort of notebook Ellen podcasts clones. I think they're a real challenger to 11 Labs as well.

And RWBKB of course is rolling on on Windows. So I'll still say these are niche. We've been talking about them as the future for a long time. And I mean, we live technically in a year in the future from last year, and we're still saying the exact same things as we were saying last year.

So what's changed? I don't know. I do think the XLSTM paper, which we will cover when we cover the sort of NeurIPS papers is worth a look. I think they are very clear eyed as to how they wanna fix LSTM. Okay, so, and then we also wanna cover a little bit like the major themes of the year.

And then we wanted to go month by month. So I'll bridge you into it back to the recording, which we still have the audio of. So one of the major themes is sort of the inference race at the bottom. We started this time last year with the Mistral Price War of 2023, with Mistral going from $1.80 per token down to $1.27 in the span of like a couple of weeks.

And, you know, I think this, a lot of people also interested in the price war, sort of the price intelligence curve for this year as well. I started tracking it, I think roundabout in March of 2024 with Haiku's launch. And so this is, if you're watching the YouTube, this is what I initially charted out as like, here's the frontier.

Like, everyone's kind of like in a pretty tight range of LMS's ELO versus the model pricing. You can pay more for more intelligence and it'll be cheaper to get less intelligence, but roughly it correlates to aligned and a trend line. And then I could update it again in July and see that everything had kind of shifted right.

So for the same amount of ELO, let's say GPT-4, 2023 would be about sort of $11.75 as an ELO, you could, and you used to get that for like $40 per token, a per million tokens. And now you get Cloud3 Haiku, which is about the same ELO for 50 cents.

And so that's a two orders of magnitude improvement in about two years, sorry, in about a year. But more importantly, I think you can see the more recent launches like Cloud3 Opus, which launched in March this year, now basically superseded completely, completely dominated by Gemini 1.5 Pro, which is both cheaper, $5 a month, $5 per million, as well as smarter.

So it's about slightly higher than 12.50 in ELO. So the March frontier and shift to the July frontier is roughly one order of magnitude improvement per sort of ISO ELO. And I think what you're starting to see now in July is the emergence of 4.0 Mini and DeepSea V2 as outliers to the July frontier, where July frontier used to be maintained by 4.0, Lama 4.5, Gemini 1.5 Flash and Mistral Nemo.

These things kind of break the frontier. And then if you update it like a month later, I think if I go back a month here, you update it, you can see start, you can see more items start to appear here as well with the August frontier, with Gemini 1.5 Flash coming out with an August update as compared to the June update, being a lot cheaper and roughly the same ELO.

And then we update for September. And this is one of those things where we really started to understand the pricing curves being real instead of something that some random person on the internet drew on a chart, because Gemini 1.5 cut their prices and cut their prices exactly in line with where everyone else is in terms of their ELO price charts.

So if you plot by September, we had O1 preview and pricing and costs and ELOs. So the frontier was O1 preview, GPC 4.0, O1 mini, 4.0 mini, and then Gemini Flash at the low end. That was the frontier as of September. Gemini 1.5 Pro was not on that frontier.

Then they cut their prices, they halved their prices, and suddenly they were on the frontier. And so it's a very, very tight and predictive line, which I thought it was really interesting and entertaining as well. And I thought that was kind of cool. In November, we had 3.5 Haiku new.

And obviously we had Sonnet as well. Sonnet is not, I don't know where this Sonnet on this chart, but Haiku new basically was 4X the price of old Haiku. Oh, sorry, 3.5 Haiku was 4X the price of 3 Haiku. And people were kind of unhappy about that. There's a reasonable assumption, to be honest, that it's not a price hike, it's just a bigger model.

So it costs more, but we just don't know that. There was no transparency on that. So we are left to draw our own conclusions on what that means. That's just is what it is. So yeah, that would be the sort of price ELO chart. I would say that the main update for this one, if you go to my LLM pricing chart, which is public, you can ask me for it, or I've shared it online as well.

The most recent one is Amazon Nova, which we briefly, briefly talked about on the pod, where they've really sort of come in and basically offered Amazon Basics LLM, where Amazon Pro, Nova Pro, Nova Light, and Nova Micro are the efficient frontier for their intelligence levels of 1200 to 1300.

You wanna get beyond 1300, you have to pay up for the O1s of the world, and the 4Os of the world, and the Gemini 1.5 Pros of the world. But 2Flash is not on here, and is probably a good deal higher. Flash Thinking is not on here, as well as all the other QWQs, R1s, and all the other sort of thinking models.

So I'm gonna have to update this chart. It's always a struggle to keep up to date. But I wanna give you the idea that basically for, through 2024, for the same amount of ELO, what you used to pay at the start of 2024, let's say $40 to $50 per million tokens, now is available approximately, with Amazon Nova, approximately at, I don't know, $0.075 per token.

So like 7.5 cents. So that is a couple orders of magnitude at least. Actually, almost three orders of magnitude improvement in a year. And I used to say that intelligence, the cost of intelligence was coming down one order of magnitude per year, like 10x. That is already faster than Moore's Law.

But coming down three times this year is something that I think not enough people are talking about. And so even though people understand that intelligence has become cheaper, I don't think people are appreciating how much more accelerated this year has been. And obviously, I think a lot of people are speculating how much more next year will be with H200s becoming commodity, Blackwell's coming out.

It's very hard to predict, and obviously there are a lot of factors beyond just the GPUs. So that is the sort of thematic overview. And then we went into sort of the annual overview. This is basically us going through the AI News releases of the year, and just picking out favorites.

I had Will, our new research assistant, help out with the research, but you can go on to AI News and check out all the sort of top news of the day. But we had a little bit of an AI Rewind thing, which I'll briefly bridge you in back to the recording that we had.

So January, we had the first round of the year for Perplexity. And for me, it was notable that Jeff Bezos backed it. Jeff doesn't invest in a whole lot of companies, but when he does, he backed Google back in the day, and now he's backing the new Google, which is kind of cool.

Perplexity is now worth $9 billion. I think they have four rounds this year. Will also picked out that Sam was talking about GPT-5 soon. This was back when he was, I think, at one of the sort of global summit type things, Davos. And yeah, no GPT-5. It's actually, we got '01 and '03.

In February, you know, people were, we were still sort of thinking about last year's Dev Day. And this was three months on from Dev Day. People were kind of losing confidence in GPTs. And I feel like that hasn't super recovered yet. I hear from people that there are still stuff in the works and you should not give up on them.

And they're actually underrated now, which is good. So I think people are taking a stab at the problem. I think it's a thing that should exist. And we just need to keep iterating on them. Honestly, any marketplace is hard. It's very hard to judge given all the other stuff they've shipped.

ChachiGPT also released Memory in February, which we talked about a little bit. We also had Gemini's Diversity Drama, which we don't tend to talk a ton about in this podcast 'cause we try to keep it technical. But we also started seeing context window size blow out. So this year, I mean, it was Gemini with 1 million tokens.

But also I think there's 2 million tokens talked about. We had a podcast with Gradients talking about how to fine tune for 1 million tokens. It's not just like what you declare to be your token context, but you also have to use it well. And increasingly, I think people are looking at not just Ruler, which is sort of multi needle in a haystack we talked about, but also Muser and like reasoning over long context, not just being able to retrieve over long context.

And so that's what I would call out there. Specifically, I think magic.dev as well, made a lot of waves for the 100 million token model, which was kind of teased last year, but whatever it was, they made some noise about it. Still not released, so we don't know, but we'll try to get them on the podcast.

In March, Cloud 3 came out, which huge, huge, huge for a topic. This basically started to mark the shift of market share that we talked about earlier in the pod, where most production traffic was on OpenAI, and now Enthropic had a decent frontier model family that people could shift to.

And obviously, now we know that Sonnet is kind of the workhorse, just like 4.0 is the workhorse of OpenAI. Devon came out in March, and that was a very, very big launch. It was probably one of the most well-executed PR campaigns, maybe in tech, maybe this decade. And then I think there was a lot of backlash as to what specifically was real in the videos that they launched with, and then they took nine months to ship to GA.

And now you can buy it for $500 a month and form your own opinion. I think some people are happy, some people less so, but it's very hard to live up to the promises that they made and the fact that for some of them they do, which is interesting.

I think the main thing I would caution out for Devon, I think people call me a Devon show sometimes 'cause I say nice things, like one nice thing doesn't mean I'm a show, basically is that a lot of the ideas can be copied. And this is always the threat of "GPT wrappers" that you achieve product market fit with one feature, it's gonna be copied by 100 other people.

So of course, you gotta compete with branding and better products and better engineering and all that sort of stuff, which Devon has in spades, so we'll see. - April, we actually talked to Yurio and Suno. We talked to Suno specifically, but Yurio I also got a beta access to and like AI music generation.

We played with that on the podcast. I loved it. Some of our friends at the pod like play in their cars, like I rode in their cars while they played our Suno intro songs and I freaking loved using O1 to craft the lyrics and Suno and Yurio to make the songs.

But ultimately like a lot of people, some people were skipping them. I don't know what exact percentages, but those 10% of you that skipped it, you're the reason why we cut the intro songs. (laughs) We also had Lama 3 release. So I think people always wanna see like a good frontier open source model and Lama 3 obviously delivered on that with the 8B and 70B, the 400B came later.

Then May, GPC 4.0 released. And it was a kind of a model efficiency thing, but also I think just a really good demo of all the things that 4.0 was capable of. Like this is where the messaging of omni-model really started kicking in. Previously 4 and 4 Turbo were all text and not natively sort of vision.

I mean, they had vision, but not natively voice. And I think everyone fell in love immediately with the Sky Voice and Sky Voice got taken away before the public release. And I think it's probably self-inflicted. I think that the version of events that has Sam Altman basically putting a foot in his mouth with a three-letter tweet causing decent grounds for a lawsuit where there was no grounds to be had because they actually just use the voice actress that sounded like Scarlett Johansson is unfortunate because we could have had it and we don't.

So that's what it is. And that's what the consensus seems to be from the people I talked to. People be pining for the Scarlett Johansson voice. In June, Apple announced Apple Intelligence at WWDC. And we haven't, most of us, if you update your phones, have it now if you're on an iPhone.

And I would say it's like decent. I think it wasn't the game changer thing that caused the Apple stock to rise like 20%. And just because everyone was gonna upgrade their iPhones just to get Apple Intelligence, it did not become that. But it is probably the largest scale rollout of transformers yet after Google rolled out BERT for search.

And people are using it. And it's a 3B foundation model that's running locally on your phone with Loras that are hot swaps and we have papers for it. Honestly, Apple did a fantastic job of doing the best that they can. They're not the most transparent company in the world and nobody expects them to be.

But they gave us more than I think we normally get for Apple tech. And that's very nice for the research community as well. NVIDIA, I think we continue to talk about, I think I was at the Taiwanese trade show, Comtex, and saw him signing women body parts. And I think that was maybe a sign of the times, maybe a sign that things have peaked, but things are clearly not peaked 'cause they continued.

Going, ILIA, and then that bridges us back into the episode recording. I'm gonna stop now and stop yapping. But yeah, we recorded a whole bunch of stuff. We lost it and we're scrambling to re-record it for you, but also we're trying to close the chapter on 2024. So now I'm gonna cut back to the recording where we talk about the rest of June, July, August, September, and the second half of 2024's news.

And we'll end the episode there. ILIA came out from the woodwork. - Yeah, raised a billion. - Saw a term sheet, raised a billion dollars. Dan Gross seems to have now become full-time CEO of the company, which is interesting. I thought he was gonna be an investor for life, but now he's operating.

- He was an investor for a short amount. - For two years. What else can we say about ILIA? I mean, like, I think this idea that you only ship one product and it's a straight shot at super intelligence seems like a really good focusing mission, but then it runs counter to basically both Tesla and OpenAI in terms of the ship intermediate products that get you to that vision.

- Well, I think the question is like, OpenAI now needs then more money because they need to support those products. And I think maybe their bet is like with 1 billion, we can get to the thing. But we don't wanna have to have intermediate steps. Like we're just making it clear that like, this is what it's about.

- Yeah, but then like, where do you get your data? - Yeah, totally. - So I think that's the question. I think we can also use this as part of a general theme of the safety wing of OpenAI leaving. - Yeah. - It's fair to say that Jan Leica also left and like basically the entire super alignment team left.

- Yeah, then there was artifacts kind of like the Chajupiti Canvas equipment that came out. - I think more code-oriented. - Yeah. - No one has a Canvas clone yet apart from OpenAI. Interestingly, I think the same person responsible for artifacts and Canvas, Karina, officially left Anthropic after this to join OpenAI on the rare reverse moves.

- Yeah, and then we had AI Engineer World's Fair in June. I was over 2,000 people, not including us. I would love to attend the next one. - If only we can get tickets. - Yeah, but I think a really good demo. We now have it and deploy it for everybody.

And Gemini actually kind of beat them to the GA release, which is kind of interesting. I think that everyone should basically always have this on as long as you're comfortable with the privacy settings because then you have a second person kind of looking over your shoulder. - Yeah. - And like this time next year, I would be willing to bet that I would just have this running on my machine.

And I think that assistance always on that you can talk to with vision that sees what you're seeing. I think that is where at least one hour software experience to go. Then it will be another few years for that to happen in real life in outside of the screen.

But for screen experiences, I think it's basically here, but not evenly distributed. And we've just seen the GA of this capability that was demoed in June. - And then July was Lama 3.1, which we've done a whole podcast on, but that was great. - July and August is kind of quiet.

- Yeah, structure uploads. We also did a full podcast on that. And then September we got O1. - Yes. - Strawberry, AKA Q*, AKA we had a nice party with strawberry glasses. - Yes, I think very underrated. Like this is basically from the first internal demo of Q of strawberry was let's say November, 2023.

So between November to September, like the whole red teaming and everything, honestly, a very good ship rate. Like I don't know if like people are giving OpenAI enough credit for like this all being available. In ChargeGBT and then shortly after in API. I think maybe in the same day, I don't know.

I don't remember the exact sequence already, but like this is like the frontier model that was like rolled out very, very quickly to the whole world. And then we immediately got used to it, immediately said it was shit because we're still using Sonnet or whatever, but like still very good.

And then obviously now we have O1 Pro and O1 Full. I think like in terms of like biggest ships of the year, I think this is it, right? - Yeah, yeah, totally. Yeah, and I think it now opens a whole new Pandora's box for like the inference time compute and all that.

- Yeah, yeah, yeah. It's funny because like it could have been done by anyone else before. - Yeah. - Literally, this is an open secret. They were working on it ever since they hired GNOME, but no one else did. - Yeah. - Another discovery, I think Ilya actually worked on a previous version called GPT-0 in 2021.

Same exact idea and it failed, whatever that means. - Yeah, timing. - Voice mode also. - Voice mode, yeah. Yeah, I think most people have tried it by now because it's generally available. - Yeah, I think your wife also likes it. - Yeah, yeah, yeah. She talks to it all the time.

- Okay. - Canvas in October. - Okay. - Another big release. - Have you used it much? - Not really, honestly. - I use it a lot. - What do you use it for mostly? - Drafting anything. I think that people don't see where all this is heading. Like OpenAI is really competing with Google in everything.

Canvas is Google Docs. And like it's a full document editing environment with the auto assister thing at the side that is arguably better than Google Docs, at least for some editing use cases, right? 'Cause it has a much better AI integration than Google Docs with Gemini on the side.

And so OpenAI is taking on Google and Google Docs. It's also taking it on in search. They launched their little Chrome extension thing to be the default search. And I think like piece by piece, it's kind of really tackling on Google in a very smart way that I think is additive to workflow and people should start using it as intended because this is a peek into the future.

Maybe they're not successful, but at least they're trying. And I think Google has gone without competition for so long that anyone trying will at least receive some attention from me. - Yeah, and then yeah, computer use also came out. Yeah, that was a busy, it's been a busy couple of months.

- Busy couple of months. I would say that computer use was one of the most upvoted demos on Hacker News of the year, but then comparatively, I don't see people using it as much. This is how you feel the difference between a mature capability and an emerging capability. Maybe this is why vision is emerging because I launched computer use, you're not using it today, but you'll use everything else in a mature category.

And it's mostly because it's not precise enough or it's too slow or it's too expensive. And those will be the main criticisms. - Yeah, that makes sense. It's also just like overall uneasiness about just letting it go crazy. - I don't care. - Yeah, no, no, totally. But I think a lot of people do.

- November. - R1, so that was kind of like the open source of one competitor by DeepSeek. - Yeah, nobody knew it was coming. Everyone knew like F1, we had a preview at the Fireworks HQ. And then I think some other labs did it, but I think R1 and QWQ, Quill from the Quent team, both Alibaba affiliated, I think, are the leading contenders on that front.

And we'll see, we'll see. - What else to highlight? I think the Stripe agent toolkit, I mean, it's a small thing, but it's just like people are like agents are not real. It's like when you have companies like Stripe and like start to build things to support it, it might not be real today, but obviously they don't have to do it because they're not an AI company.

But the fact that they do it shows that there's one demand and so there's belief on their end. - This is a broader thing about, broader thesis for me that I'm exploring around, do we need special SDKs for agents? Why can't normal SDKs for humans do the same thing?

Stripe agent toolkits happens to be a wrapper on the Stripe SDK, it's fine. It's just like a nice little DX layer. But like, it's still unclear to me. I have been asked my opinion on this before. And I said, I think I said it on a podcast, which is like the main layer that you need is separate off roles so that you don't assume it's a human doing these things.

And you can lock things down much quicker or you can identify whether it is an agent acting on your behalf or actually you. And that is something that you need. I had my 11 labs key pwned because I lost my laptop. And I saw a whole bunch of API calls and I was like, "Oh, is that me?

"Or is that someone?" And it turned out to be a key that I committed onto GitHub and that I didn't scrape. And so sourcing of where API usage is coming from, I think you should attribute it to agents and build for that world. But other than that, I think SDKs, I would see it as a failure of tech and AI that we need every single thing needs to be reinvented for agents.

- I agree in some ways. I think in other ways, we've also like not always made things super explicit. There's kind of like a lot of defaults that people do when they design APIs. But like, I think if you were to redesign them in a world in which the person or the agent using them as like almost infinite memory and context, like you will maybe do things differently.

- Yeah. - But I don't know. I think to me that the most interesting is like REST and GraphQL is almost more interesting in the world of agents because agents could come up with so many different things to query versus like before I always thought GraphQL was kind of like not really necessary because like, you know what you need, just build the REST endpoint for it.

So yeah, I'm curious to see what else changes. And then yeah, the search wars. I think that was, you know, search GPT, perplexity. - Dropbox. - Dropbox dash. Yeah, we had Drew on the pod and then we added the Pioneer Summit. The fact that Dropbox is a Google Drive integration, it's just like, if you told somebody five years ago, it's like, Dropbox doesn't really care about hosting your files.

You know, it's like, that doesn't compute. So yeah, I'm curious to see where that goes. - Cool. - This whole space. - And that brings us up to December, still developing. I'm curious what the last day of opening as shipments will be. I think everyone's expecting something big there.

I think so far has been a very eventful year. Definitely has grown a lot. We were asked by Will actually, like whether we made predictions. I don't think we did, but. - Not really. - Maybe we should. - Well, I think we definitely talked about agents. - Yes. - And I don't know if we said it was the year of the agents, but we said.

- But next year is the year. - No, no, but well, you know, the anatomy of autonomy, that was April, 2023, you know? So obviously there's been belief for a while. - Yes. - But I think now the models are, I would say maybe the last, yeah, two months.

- Yeah. - I've made a big push in like capability for like 3.16, 4.1. - Yeah, I mean, Ilya saying the word agentic on stage in Europe, it's a big deal. Satya, I think also saying that a lot these days. I mean, Sam has been saying that for a while now.

So DeepMind, when you announced Gemini 2.0, they announced Deep Research, but also Project Mariner, which is a browser agent, which is their computer use type thing, as well as Jules, which is their code agent. And I think that basically complements with whatever OpenAI is shipping next year, which is Codename Operator, which is their agent thing.

It makes sense that if it actually replaces a junior employee, they will charge $2,000 for it. - Yeah, I think that's my whole, I did this post that it's pinned on my Twitter, so you can find it easily, but about skill floor and skill ceiling in jobs. And I think the skill floor more and more, I think 2025 will be the first year where the AI sets the skill floor of a role.

I don't think that has been true in the past, but yeah, I think now really like if Devon works, if all these customer support agents are working. So now to be a customer support person, you need to be better than an agent because the economics just don't work. I think the same is gonna happen to in software engineering, which I think the skill floor is very low.

There's a lot of people doing software engineering that are really not that good. So I'm curious to see in the next year of the recap, what other jobs are gonna have that change. - Yeah, every NeurIPS that I go, I have some chats with researchers and I'll just highlight the best prediction from that group.

And then we'll move on to end of year recap in terms of we'll just go down the list of top five podcasts and then we'll end it. So the best prediction was that there will be a foreign spy caught at one of the major labs. So this is part of the consciousness already that whenever you see someone who is like too attractive in a San Francisco party, where the ratio is like a hundred guys to one girl and suddenly the girl's like super interested in you, like it may not be your looks.

So there's a lot of like state level secrets that are kept in these labs. And not that much security. I think if anything, the situational awareness essay did to raise awareness of it, I think it was directionally correct, even if not precisely correct. We should start caring a lot about this.

OpenAI has hired a CISO this year. And I think like the security space in general, oh, I remember what I was gonna say about Apple foundation model before we cut for a break. They announced Apple secure cloud, cloud compute. And I think we're also interested in investing in areas that are basically secure cloud, LLM inference for everybody.

I think like what we have today is not secure enough. And because it's like normal security when like this is literally a state level interest. - Agreed. - Top episodes? - Yeah. So I'm just going through the sub stack. Number one, the David Luan one. - Yeah. - It's the most popular in 2024.

Why Google failed to make GPT-3? - I will take a little bit of credit for that, for the naming of that one, because I think that was the hack news thing. It's very funny because like, actually, obviously he wants to talk about debt, but then he spent half the episode talking about his time at OpenAI.

But I think it was a very useful insight that I'm still using today. Even in like the earlier post, I was still referring to what he said. And when we do podcast episodes, I try to look for that. I try to look for things that we'll still be referencing in the future.

And that concentrated badness, David talked about the brain compute marketplace, and then Ilya in his emails that I covered in the what Ilya saw essay, had the OpenAI side of this, where they were like one big training run is much, much more valuable than the hundred equivalent small training runs.

So we need to go big and we need to concentrate badness, not spread them. - Number two, how NotebookLM was made. - Yeah, that was fun. - Yeah, and everybody, I mean, I think that's like a great example of like just timeliness. You know, I think it was top of mind for everybody.

There were great guests. It just made the rounds on social media. - Yeah, and that one, I would say, Risa is obviously a star, but she's been on every episode, every podcast. But Usama, I think, you know, actually being the guy who worked on the audio model, being able to talk to him, I think was a great gift for us.

And I think people should listen back to how they trained the NotebookLM model. 'Cause I think you put that level of attention on any model, you will make it solder. - Yeah, that's true. - And it's specifically like, they didn't have evals. They just- - Vibes. - They had a group session with vibes.

The ultimate guide to prompting, that was number three. I think all these episodes that are like summarizing things that people care about, but they're disparate, I think always do very well. - This helps us save on a lot of smaller prompting episodes, right? If we interviewed individual paper authors with like a 10-page paper that is just a different prompt, like not as useful as like an overview survey thing.

I think the question is what to do from here. People have actually, I would say, I've been surprised by how well received that was. Should we do ultimate guide to other things? And then should we do prompting 201, right? Those are the two lessons that we can learn from the success of this one.

- I think if somebody does the work for us, that was the good thing about Sander. Like he had done all the work for us. - Yeah, Sander is very, very fastidious about this. So he did a lot of work on that. I'm definitely keen to have him on next year to talk more prompting.

Okay, then the next one is the not safe for work one. - No. - Or structured outputs. - The next one is brain trust. - Really? - Yeah. - Okay, we have a different list then, but yeah. - I'm just going on the sub stack. - I see, I see.

So that includes the number of likes, but I was going by downloads. It's fine. - I would say this is almost recency bias in the way that like the audience keeps growing and then like the most recent episodes get more views. So I would say definitely like the NSFW one was very popular.

What people were telling me they really liked because it was something people don't cover. Yeah, structural outputs. I think people liked that one. I mean, the same one. Yeah, I think that's like something I refer to all the time. I think that's one of the most interesting areas. No, the simulation.

- Oh, WebSim, really? - Yeah, not that use case, but like how do you use that for like model training and like agents learning and all of that. - Yeah, so I would definitely point to our newest seven hour long episode on Simulative Environments because it is the, let's say the scaled up, very serious AGI lab version of WebSim and WebSim.

If you take it very, very seriously, you get Genie 2, which is exactly what you need to then build Sora and everything else. So yeah, I think Simulative AI still in summer. - Yeah, still in summer. - Still coming. And I was actually reflecting on this. Like, would you say that the AI winter has like coming on or like was it never even here?

'Cause we did a "Winds of AI Winter" episode and I was like trying to look for signs. I think that's kind of gone now. - Yeah, I would say it was here in the vibes, but not really in the reality. You know, when you look back at the yearly recap, it's like every month there was like progress.

There wasn't really a winter. There was maybe like a hype winter, but I don't know if that counts as a real winter. - I think the scaling has hit a wall thing has been a big driving discussion for 2024. And, you know, with some amount of conclusion at NeurIPS that we were also kind of pointing to in the "Winds of AI Winter" episode, but like, it's not a winter by any means.

We know what winter feels like. It is not winter. So I think things are going well. I think every time that people think that there's like not much happening in AI, just think back to this time last year and understand how much has changed from benchmarks to frontier models, to market share between OpenAI and the rest.

And then also cover like, you know, the various coverage areas that we've marked out, how the discussion has evolved a lot and what we take for granted now versus what we did not have a year ago. - Yeah, and then just to like throw that out there, there've been 133 funding rounds over a hundred million in AI this year.

- Does that include Databricks, the largest venture around in history? - And that's $10 billion, sheesh. Well, that mosaic now has been bought for two something billion because it was mostly stock, you know? So price goes up. - I see. - Theoretically. - I see. So you just bought at a valuation of 40, right?

- Yeah, it was like 43 or something like that. - And at the time, I remember at the time, there was a question about whether or not the valuation was real. - Yeah, well, that's why everybody- - Snowflake was down. - Yeah. - And like Databricks was a private valuation that was like two years old.

It's like, who knows what this thing's worth. Now it's worth 60 billion. - It's worth more. It's worth more. That's what it's worth. It's worth more than what you thought. - Yeah, it's been a crazy year, but I'm excited for next year. I feel like this is almost like, you know, now the agent thing needs to happen.

And I think that's really the unlock. - Yeah, I think it needs to happen. - I mean- - I have to agree with you. Next year is the year of the agent in production. - Yeah, I don't, you know, it's almost like I'm not a hundred percent sure it will happen, but like it needs to happen.

Otherwise it's definitely the winter next year. Any other parting thoughts? - I'm very grateful for you. I think you've been a dream partner to build Lane Space with. And also the Discord community, the paper club people have been beyond my wildest dreams, so supportive and successful. It's amazing that the community has grown so much and the vibe has not changed.

- Yeah, that's true. - We're almost at 5,000 people. - Yeah, we started this Discord like four years ago. - Yeah. - And still people get it when they join. You post news here and then you discuss it in threads and you try not to self-promote too much. And mostly people obey the rules and sometimes you smack them down a little bit, but that's okay.

- We rarely have to ban people, which is great. But yeah, man, it's been awesome, man. I think we both started not knowing where this was going to go. And now we've done a hundred episodes. It's easy to see how we're going to get to 200. I think maybe when we started, it wasn't easy to see how we would get to 100, you know?

Yeah, excited for more. Subscribe on YouTube because I swear to God, we're doing so much work. We need to make that work. - It's very expensive for an unclear payoff as to what we're actually going to get out of it. But hopefully people discover us more there. I do believe in YouTube as a podcasting platform much more so than Spotify.

- Yeah, totally. Thank you all for listening. - Yeah, thank you for listening. - See you in the new year. - Bye-bye. (upbeat music)

2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents

Chapters

Transcript