Should you worry about using the latest AI model?
New AI models are released every week. Is it worth it to keep up with the latest and greatest? Eric says no, John says yes, and they debate to find a practical middle ground.
Subscribe to get notified of new episodes
Watch on YouTube
Show Notes
Summary
Across all AI labs, a new model is released every few days. The frontier labs (OpenAI, Anthropic, and Google) release a major new model every 2 weeks. Keeping up with the latest and greatest can feel exhausting, and that's before you dig in to the details of why a new model is better, and what it's better at.
People with a Claude or ChatGPT subscription get access to new models automatically when they are made available through the respective apps, but many people still don't know if it's worth it to burn through credits faster with a more powerful engine under the hood. For anyone who has moved beyond the apps, the questions get even harder. Are open-weight models as good as the frontier? Is switching models worth the cost of changing your setup?
Eric and John stage a debate about whether you should keep up with the latest models.
John says yes:
You fail to understand what is possible if you're only ever optimizing for using the cheapest model. You need to be able to understand what you can do now that you couldn't do before.
Eric says no:
Someone who is extremely good at using AI can use a less capable model and do more than someone who is following the path of least resistance with the latest model. Before worrying about using the latest model, I would focus on becoming the type of AI user who can notice the differences.
By the end, the episode finds the middle ground of reality, derived directly from Eric and John's years of using AI in their daily work. As you become a more proficient user of AI, the right pattern is intentionally using different models for different parts of your workflow. In order to do that, you need to use mid-tier models to build core skills, experiment with the frontier to understand the limits of what's possible, and develop the discernment to know the best application for each.
Key takeaways
- Optimizing for cost obfuscates the art of the possible: A proficient AI user can do extraordinary things with a mid-tier model, but that's because they understand how capable models are. Trying to reach the ceiling on the frontier changes the way you think about AI as a medium and how you can apply it.
- Optimizing for your own proficiency should be the default: When you begin to notice the differences in a new model, especially the subtle ones, you're starting to build the core skill necessary to get the most use out of any model.
- Advanced AI users employ different models for different jobs: The proven pattern is using expensive frontier models for more critical work and cheaper, faster models for more standard tasks. Smart users learn to have frontier models build and audit workflows of cheaper models to get the most bang for their buck.
Notable mentions and links
- AI Gateway is a model router from Vercel that gives you access to 100s of models at cost. Routers like these are a great way to experiment with different models from different labs.
- The AI Gateway Production Index (also from Vercel) is a monthly report that details trends in model usage across trillions of tokens per day, including adoption metrics for the latest releases.
- Kimi K3 from Moonshot AI made big news when it achieved near-frontier performance at a fraction of the cost.
- Ben Thompson writes Stratechery, and his recent article on the economics of the frontier labs vs open-weight models breaks down how the labs make money (hint, it's not by training the most powerful models).
- "The medium is the message" is a principle developed by Marshall McLuhan in which he argues that the medium through which content is delivered cannot be separated from the content and is, in fact, a part of the content itself.
Transcript
00:00:00,600 --> 00:00:37,920 [Eric] [instrumental music] Welcome back to the Token Intelligence Show. AI is changing the way that we work. And here on the Token Intelligence Show, we help you cut through the hype. We show you the state-of-the-art and help you apply wisdom to be a great leader in this new environment that we're all, uh, trying to figure out. And today, we have a really interesting question, John, and it's going to be fun for our listeners because 00:00:39,260 --> 00:00:42,400 [Eric] I answer this question with a no, and you answer this question- 00:00:42,400 --> 00:00:42,410 [John] Yeah 00:00:42,410 --> 00:00:43,260 [Eric] ... with a yes. 00:00:43,260 --> 00:00:43,480 [John] Yeah. 00:00:43,480 --> 00:00:47,230 [Eric] So it's a little bit of a, uh, a little bit of a bake-off here, 00:00:48,340 --> 00:00:49,470 [Eric] uh, or a difference of opinion. 00:00:50,480 --> 00:01:03,460 [Eric] Here's the question. We'll get right into it. Should you be worried about using the latest AI model? And before we get into a very heated debate, uh [laughs]- 00:01:03,460 --> 00:01:05,200 [John] [laughs] 00:01:05,200 --> 00:01:26,080 [Eric] ... let's just define what we mean by the latest model, which I think is helpful, and the re- and the pace of change there, just to level set, because I think we'll have listeners who will know exactly what we mean when we say that, and then other listeners who may not even be familiar with that concept because it's baked into the- 00:01:26,080 --> 00:01:26,170 [John] Right 00:01:26,170 --> 00:01:27,660 [Eric] ... AI app that they're using. 00:01:27,660 --> 00:01:28,100 [John] Right. 00:01:28,100 --> 00:01:28,960 [Eric] Uh, so 00:01:30,020 --> 00:01:35,900 [Eric] the latest model. 1st of all, what is the latest model, John? Do you know the answer to that question? 00:01:35,900 --> 00:01:36,130 [John] Yeah. 00:01:37,360 --> 00:01:41,580 [John] But, you know, by the [laughs] by the time we release this, it might be different. 00:01:41,580 --> 00:01:42,600 [Eric] That's true. 00:01:42,600 --> 00:01:51,900 [John] Um, so the latest frontier models would be from Grok 4.5, from Anthropic would be Fable 5 or Mythos, if you have access- 00:01:51,900 --> 00:01:51,910 [Eric] Yep 00:01:51,910 --> 00:01:56,100 [John] ... to Mythos. And then, um, from ChatGPT, 00:01:57,240 --> 00:02:01,660 [John] OpenAI is a new, a whole GPT-5.6, and there's- 00:02:01,660 --> 00:02:01,800 [Eric] Yep 00:02:01,800 --> 00:02:02,840 [John] ... 3 versions of it. 00:02:02,840 --> 00:02:08,660 [Eric] Yep. And those are from Frontier Labs, so give a quick definition on Frontier Labs. 00:02:09,699 --> 00:02:13,100 [John] Yeah. Yeah, somebody... I used that term the other day and people were like, "What are you talking about?" 00:02:13,100 --> 00:02:13,579 [Eric] Hmm. 00:02:13,580 --> 00:02:20,660 [John] Yeah, so Frontier Labs, um, I don't, I don't know if there's a precise definition, 'cause I do think it changes. 00:02:20,660 --> 00:02:20,670 [Eric] Mm-hmm. 00:02:20,670 --> 00:02:21,160 [John] I think 00:02:22,200 --> 00:02:27,700 [John] for sure most peop- everyone, I think, would consider Anthropic and OpenAI Frontier Labs. 00:02:27,700 --> 00:02:28,680 [Eric] Yep, absolutely. 00:02:28,680 --> 00:02:39,040 [John] And outside of that, I don't know. Do you think l- w- are, are Google and Meta and, and, um, XAI, like, s- still frontier? 00:02:39,040 --> 00:02:42,160 [Eric] I, I would, I would definitely consider Google on the frontier. 00:02:42,160 --> 00:02:42,400 [John] Yeah. 00:02:42,400 --> 00:02:49,380 [Eric] Um, and xAI I think is very close, uh, but it's not evenly distributed, actually. 00:02:49,380 --> 00:02:49,920 [John] Right. 00:02:49,920 --> 00:02:53,740 [Eric] Uh, we do report at Vercel called the AI Production Index. 00:02:53,740 --> 00:02:54,200 [John] Okay. 00:02:54,200 --> 00:03:03,320 [Eric] AI Gateway Production Index, which is a bunch of telemetry that we collect from a product that we have called AI Gateway that actually gives you access to any model that you want. 00:03:03,320 --> 00:03:03,960 [John] Right. 00:03:03,960 --> 00:03:14,300 [Eric] And you just pay cost, which is cool. But, uh, what's interesting is that, uh, xAI leads in certain types of work. 00:03:14,300 --> 00:03:14,420 [John] Sure. 00:03:14,420 --> 00:03:16,340 [Eric] So video in particular. 00:03:16,340 --> 00:03:17,300 [John] Sure. 00:03:17,300 --> 00:03:19,280 [Eric] Uh, it's very, very good at video. 00:03:19,280 --> 00:03:19,540 [John] Right. 00:03:19,540 --> 00:03:29,320 [Eric] Right? Whereas sort of standard text generation, um, it's the, like spend and adoption is not nearly as high as Anthropic or OpenAI- 00:03:29,320 --> 00:03:29,780 [John] Yeah 00:03:29,780 --> 00:03:35,320 [Eric] ... or even Google, right? So it's not easy to say... It's actually difficult to say who's a clear winner. 00:03:35,320 --> 00:03:42,980 [John] Well, and to me, frontier would mean you're, you are the winner in something, right? Like, you're the top for, for some meaningful 00:03:44,040 --> 00:03:45,420 [John] type of work you might do with AI. 00:03:46,640 --> 00:03:49,260 [Eric] Yes, but again, that's not evenly distributed, which is- 00:03:49,260 --> 00:03:50,109 [John] No, it's not, yeah 00:03:50,109 --> 00:03:52,020 [Eric] ... it, it's good for us to spend just a moment on this- 00:03:52,020 --> 00:03:52,030 [John] Right 00:03:52,030 --> 00:03:56,530 [Eric] ... 'cause I think it's really helpful. This is not obvious from reading- 00:03:56,530 --> 00:03:56,660 [John] Right 00:03:56,660 --> 00:03:57,600 [Eric] ... the news. But- 00:03:57,600 --> 00:03:57,800 [John] Right 00:03:57,800 --> 00:04:00,420 [Eric] ... the data actually says 00:04:01,840 --> 00:04:02,170 [Eric] that 00:04:03,620 --> 00:04:07,540 [Eric] Anthropic, which is, tends to be the most expensive- 00:04:07,540 --> 00:04:07,570 [John] Mm-hmm 00:04:07,570 --> 00:04:09,030 [Eric] ... it can kind of vary depending on- 00:04:09,030 --> 00:04:09,030 [John] Right 00:04:09,030 --> 00:04:09,980 [Eric] ... model releases- 00:04:09,980 --> 00:04:10,100 [John] Mm-hmm 00:04:10,100 --> 00:04:11,829 [Eric] ... and the workloads that people are running. But 00:04:13,000 --> 00:04:15,560 [Eric] on Vercel's AI Gateway, Anthropic 00:04:16,899 --> 00:04:19,420 [Eric] actually consumes most of the money. 00:04:19,420 --> 00:04:19,440 [John] Yeah. 00:04:19,440 --> 00:04:22,330 [Eric] And when I say most, I mean, we're talking 60 to 70% of- 00:04:22,330 --> 00:04:22,330 [John] Yeah 00:04:22,330 --> 00:04:23,990 [Eric] ... all the money that passes through the gateway. 00:04:23,990 --> 00:04:24,040 [John] Yeah. 00:04:24,040 --> 00:04:24,160 [Eric] Right? 00:04:25,460 --> 00:04:31,880 [Eric] And that's not only because of the volume. It does have a large volume share, but it's because- 00:04:31,880 --> 00:04:31,950 [John] Right 00:04:31,950 --> 00:04:33,240 [Eric] ... the price per token is- 00:04:33,240 --> 00:04:33,370 [John] It's higher 00:04:33,370 --> 00:04:34,340 [Eric] ... literally higher, right? 00:04:34,340 --> 00:04:34,460 [John] Right. 00:04:35,880 --> 00:04:40,479 [Eric] Google, for a long time, was the leader in sort of, like, mid to low-cost workloads. 00:04:40,480 --> 00:04:41,280 [John] Yeah. 00:04:41,280 --> 00:04:48,500 [Eric] And so, but I would consider both of them on the frontier because Google is what we would call, like, the enterprise workhorse, right? 00:04:48,500 --> 00:04:48,720 [John] Right. 00:04:48,720 --> 00:04:51,910 [Eric] There are a lot of workloads where you don't need the most expensive price- 00:04:51,910 --> 00:04:51,910 [John] Right 00:04:51,910 --> 00:05:02,420 [Eric] ... per token 'cause you don't need to use the absolute most intelligent model, right? And so, uh, which will actually be part of my argument. 00:05:02,420 --> 00:05:02,520 [John] Okay. 00:05:02,520 --> 00:05:03,000 [Eric] Um, 00:05:04,340 --> 00:05:07,120 [Eric] but I would consider the frontier, 00:05:09,200 --> 00:05:20,420 [Eric] uh, the labs, the AI labs who train these models, um, who are at the bleeding edge of what's possible with them. 00:05:20,420 --> 00:05:20,680 [John] Right. 00:05:20,680 --> 00:05:31,060 [Eric] Which really, the frontier is a, it's rare air because you have to be extremely well capitalized on the order of billions to have enough- 00:05:31,060 --> 00:05:31,070 [John] Right 00:05:31,070 --> 00:05:36,060 [Eric] ... money to buy the GPU capacity in order to 00:05:37,080 --> 00:05:40,599 [Eric] just run the training runs and actually do research on- 00:05:40,600 --> 00:05:40,960 [John] Right 00:05:40,960 --> 00:05:46,979 [Eric] ... how to improve the models. And then you also have to hire extremely expensive scientists and some of the best thinkers in the world on this stuff. 00:05:46,980 --> 00:05:47,900 [John] Right. 00:05:47,900 --> 00:05:53,180 [Eric] But those models also tend to be the best, which is why Anthropic can command a really- 00:05:53,180 --> 00:05:53,400 [John] Right 00:05:53,400 --> 00:05:54,180 [Eric] ... high price per token. 00:05:54,180 --> 00:05:54,880 [John] Right. 00:05:54,880 --> 00:06:05,500 [Eric] There's a flip side to this equation because all of the frontier, uh, model providers that we just talked about are closed. So you can think about this- 00:06:05,500 --> 00:06:05,510 [John] Right 00:06:05,510 --> 00:06:06,912 [Eric] ... in terms of- 00:06:06,912 --> 00:06:23,912 [Eric] ... open source software, right? So you may go purchase, you know, software like Jira. It's closed source. You pay them a licensing fee. They give you an account to use your software, to use their software, but you don't get to look at their source code. Open source software, you get to look at the source code. You can run it yourself. 00:06:23,912 --> 00:06:23,962 [John] Right. 00:06:23,962 --> 00:06:25,112 [Eric] And there's hybrid models there. 00:06:25,112 --> 00:06:25,972 [John] Right. 00:06:25,972 --> 00:06:31,212 [Eric] We don't get to look under the hood of the frontier models from Anthropic or OpenAI or Google. 00:06:31,212 --> 00:06:31,572 [John] Right. 00:06:31,572 --> 00:06:35,272 [Eric] Right? But there's a whole nother class of model, which is- 00:06:35,312 --> 00:06:38,232 [John] Which Google does technically have an open- 00:06:38,232 --> 00:06:38,272 [Eric] An o- 00:06:38,272 --> 00:06:38,532 [John] Yeah. 00:06:38,532 --> 00:06:39,351 [Eric] Yes, an open model. 00:06:39,351 --> 00:06:43,672 [John] And OpenAI has some, has one too. I just, you just never hear about it. 00:06:43,672 --> 00:06:44,262 [Eric] Right. 00:06:44,262 --> 00:06:46,992 [John] 'Cause, you know, pretty far behind their- 00:06:46,992 --> 00:06:47,062 [Eric] Yeah 00:06:47,062 --> 00:06:48,652 [John] ... state of the art. Yeah. 00:06:48,652 --> 00:06:58,972 [Eric] So I think the reason I wanted to spend a, a minute talking about that is I think it's important to understand that the models on the frontier, or frontier model providers- 00:06:58,972 --> 00:06:59,572 [John] Yeah 00:06:59,572 --> 00:07:02,252 [Eric] ... frontier AI labs, these are all terms, they mean the same thing, 00:07:04,052 --> 00:07:08,152 [Eric] those are almost always closed- 00:07:08,152 --> 00:07:08,312 [John] Yeah 00:07:08,312 --> 00:07:08,682 [Eric] ... models. 00:07:08,682 --> 00:07:08,692 [John] Right. 00:07:08,692 --> 00:07:11,012 [Eric] They don't reveal what's happening under the hood. 00:07:11,012 --> 00:07:11,102 [John] Right. 00:07:11,102 --> 00:07:11,232 [Eric] Right? 00:07:11,232 --> 00:07:11,331 [John] Right. 00:07:11,331 --> 00:07:14,372 [Eric] And, and the term that's used there is weights. 00:07:14,372 --> 00:07:14,492 [John] Right. 00:07:15,722 --> 00:07:32,002 [Eric] Um, the weights that are used in the model training. Um, but there's this whole other class, which is why this conversation became really big in the last couple weeks, of models that perform close to the level of a frontier model 00:07:33,772 --> 00:07:35,882 [Eric] that have open weights, or would be- 00:07:35,882 --> 00:07:35,882 [John] Right 00:07:35,882 --> 00:07:37,202 [Eric] ... considered open models- 00:07:37,202 --> 00:07:37,202 [John] Right 00:07:37,202 --> 00:07:44,472 [Eric] ... or open source. Again, it's not, it's not quite as, it, it's not quite this exact same thing as an open, as open source software. 00:07:44,472 --> 00:07:44,932 [John] Right. 00:07:44,932 --> 00:07:46,852 [Eric] But that's the right model to think about it from. 00:07:46,852 --> 00:07:46,882 [John] Right. 00:07:46,882 --> 00:07:52,452 [Eric] Right? Um, you can actually download the weights and, and you could actually train the model yourself if you have the GPU space. 00:07:52,452 --> 00:07:53,342 [John] Right. Yeah. 00:07:53,342 --> 00:07:57,912 [Eric] Um, or the GPU capacity, um, which some people actually do. 00:07:57,912 --> 00:07:57,932 [John] Yeah. 00:07:57,932 --> 00:07:58,792 [Eric] People set up- 00:07:58,792 --> 00:07:58,822 [John] Mm-hmm. 00:07:58,822 --> 00:08:02,092 [Eric] People literally set up chips in their houses on a rack and- 00:08:02,092 --> 00:08:02,352 [John] Yeah 00:08:02,352 --> 00:08:03,642 [Eric] ... and, and do this, which is pretty cool. 00:08:03,642 --> 00:08:04,632 [John] Yeah. Yeah. 00:08:04,632 --> 00:08:04,792 [Eric] So 00:08:06,732 --> 00:08:15,612 [Eric] all of this boils down... A- and there were some big releases. So, uh, ZAI released a model called GLM 5.2, which was a huge- 00:08:15,612 --> 00:08:15,732 [John] Yeah 00:08:15,732 --> 00:08:23,432 [Eric] ... it was a huge deal because people said, "Wow, for certain types of work, this performs as well as Anthropic's-" 00:08:23,432 --> 00:08:24,352 [John] Yeah 00:08:24,352 --> 00:08:26,172 [Eric] ... you know, latest frontier model- 00:08:26,172 --> 00:08:26,262 [John] Right 00:08:26,262 --> 00:08:28,712 [Eric] ... which would be Opus 4.8 or Fable, which- 00:08:28,712 --> 00:08:28,762 [John] Right 00:08:28,762 --> 00:08:35,472 [Eric] ... you know, recently came back online. Uh, not comprehensively, but at, at a fraction of the cost. 00:08:35,472 --> 00:08:36,032 [John] Right. 00:08:36,032 --> 00:08:36,832 [Eric] Um... 00:08:36,832 --> 00:08:45,932 [John] Yeah. And, and as well in that there are aspects that are, that are as good. So depending on what you wanna use it for, it may perform as well. 00:08:45,932 --> 00:08:46,552 [Eric] Exactly. 00:08:46,552 --> 00:08:46,812 [John] Yeah. 00:08:46,812 --> 00:08:48,832 [Eric] And so for some people 00:08:49,992 --> 00:08:58,852 [Eric] who ha- are using those particular workloads, like, "Well, my cost just... I- I'm now paying a fraction of the cost for ba- for essentially the same type of performance." 00:09:00,132 --> 00:09:04,132 [Eric] Uh, and then, um, Kimi 3 was another big model released last week- 00:09:04,132 --> 00:09:04,492 [John] Right 00:09:04,492 --> 00:09:13,612 [Eric] ... uh, that ha- that went through the same cycle, right? This is an open weights model that's a fraction of the cost that, you know, in certain areas has performance that's as good as sort of the frontier- 00:09:13,612 --> 00:09:13,892 [John] Right 00:09:13,892 --> 00:09:14,292 [Eric] ... question. So 00:09:15,892 --> 00:09:17,292 [Eric] that's why I, 00:09:18,392 --> 00:09:19,732 [Eric] I wanted to- 00:09:19,732 --> 00:09:24,362 [John] Well, and to be fair, we're talking about US frontier 'cause the 2 you listed- 00:09:24,362 --> 00:09:25,752 [Eric] We are talking about US frontier, yes 00:09:25,752 --> 00:09:32,132 [John] ... are ch- both China-based and probably would be considered Chinese frontier, would you say? 00:09:33,652 --> 00:09:46,332 [Eric] Yes. There's a great article we'll link to in the show notes, um, from Stratechery, uh, which is a wonderful blog. Uh, but there's a article that, that talks about this. Um, Ben Thompson, 00:09:47,612 --> 00:09:57,152 [Eric] um, is his name. He's been writing the blog for a really long time and has a great podcast. But he talks about whether or not you could consider the, the Chinese models frontier. 00:09:57,152 --> 00:09:57,332 [John] Nice. 00:09:57,332 --> 00:10:00,512 [Eric] He doesn't because he talks about something called distillation- 00:10:00,512 --> 00:10:00,712 [John] Oh, okay 00:10:00,712 --> 00:10:04,992 [Eric] ... which is essentially a process where they use the frontier models from the US- 00:10:04,992 --> 00:10:05,372 [John] Mm-hmm 00:10:05,372 --> 00:10:07,781 [Eric] ... distill the learnings, and then roll out their own models- 00:10:07,781 --> 00:10:07,992 [John] Right 00:10:07,992 --> 00:10:09,582 [Eric] ... that have near frontier level performance. 00:10:09,582 --> 00:10:09,592 [John] Right. 00:10:09,592 --> 00:10:11,992 [Eric] But he doesn't consider them technically frontier. 00:10:11,992 --> 00:10:12,152 [John] Right. 00:10:12,152 --> 00:10:12,602 [Eric] So- 00:10:12,602 --> 00:10:12,992 [John] Okay 00:10:12,992 --> 00:10:14,832 [Eric] ... I don't know. Interesting topic. 00:10:14,832 --> 00:10:26,272 [John] Yeah. But it's helpful to, it's helpful to think, like, just more commentary on that. You have US tech companies, some of which that have been around a long time, like, like Google, for example- 00:10:26,272 --> 00:10:26,392 [Eric] Yep 00:10:26,392 --> 00:10:29,612 [John] ... competing. You've got new US tech companies, like- 00:10:29,612 --> 00:10:29,641 [Eric] Yep 00:10:29,641 --> 00:10:33,072 [John] ... OpenAI and Anthropic that have raised tons and tons and tons of money- 00:10:33,072 --> 00:10:33,452 [Eric] Yep 00:10:33,452 --> 00:10:40,612 [John] ... that are competing here. And then you have a few others, like we mentioned, like, like, um, like Grok, you know, ZAI, and, um, 00:10:41,872 --> 00:10:45,112 [John] um, Mic- Microsoft is, like, making their own now too. 00:10:45,112 --> 00:10:45,312 [Eric] Mm-hmm. 00:10:45,312 --> 00:10:50,612 [John] But as far as frontier, people are usually talking about, you know, the OpenAI and Anthropic. 00:10:50,612 --> 00:10:52,742 [Eric] Yes. US closed models. 00:10:52,742 --> 00:10:53,572 [John] But US closed sourced- 00:10:53,572 --> 00:10:53,792 [Eric] Yep 00:10:53,792 --> 00:11:21,822 [John] ... or usually. But it's worth talking about the, the 2 you mentioned, which is Kimi K3, which came out recently, and GLM 5.2. Those are Chinese-based, and they are open source and what... or open weight. And what that means is they... you have the ability to, not maybe you personally, but there is the ability to run them on hardware that's not associated with the model creator. 00:11:21,822 --> 00:11:22,292 [Eric] Hmm. 00:11:22,292 --> 00:11:23,532 [John] Right? That's the practical- 00:11:23,532 --> 00:11:23,542 [Eric] Yes 00:11:23,542 --> 00:11:24,572 [John] ... implication here. 00:11:24,572 --> 00:11:25,331 [Eric] Mm-hmm. 00:11:25,332 --> 00:11:40,272 [John] Um, 'cause sure, somebody might wanna, like, take it and has the right compute power to, to tweak it or whatever, but it's usually, "Oh, I can save money," because essentially they're giving away the tech. You just have to find servers to run- 00:11:40,272 --> 00:11:40,452 [Eric] Yes 00:11:40,452 --> 00:11:40,992 [John] ... the tech on. 00:11:40,992 --> 00:11:41,451 [Eric] Yep. 00:11:41,452 --> 00:11:49,092 [John] Which is true in general, but I don't think the weights have come out yet for Kimi K3, so I think you have to use- 00:11:49,092 --> 00:11:49,512 [Eric] Hmm 00:11:49,512 --> 00:11:54,372 [John] ... have to have an account with them and run it in their infrastructure right now. 00:11:54,372 --> 00:11:54,692 [Eric] Yep. Yep. 00:11:54,692 --> 00:11:56,682 [John] So there is still a little bit of a lag with these models, 00:11:57,892 --> 00:12:05,932 [John] um, which is interesting because it's... 'c- 'cause at this point, it's functionally close with the promise of being open. 00:12:05,932 --> 00:12:06,232 [Eric] Yes. 00:12:06,232 --> 00:12:08,792 [John] Right? Which so far it has been. 00:12:08,792 --> 00:12:09,211 [Eric] Yeah. 00:12:09,212 --> 00:12:19,696 [John] But you... But I don't know. I, I don't know enough about, like- I would assume there's no, there's no guarantees other than, like, that's what they've staked their, like, you know, reputation on, which is something. 00:12:19,696 --> 00:12:20,616 [Eric] Right. 00:12:20,616 --> 00:12:21,216 [John] Um... 00:12:21,216 --> 00:12:23,815 [Eric] So let's get to our debate. 00:12:23,816 --> 00:12:23,976 [John] Yes. 00:12:23,976 --> 00:12:24,916 [Eric] And back to the question, 00:12:25,996 --> 00:12:34,396 [Eric] because the, the reason this was on my mind a bunch, and we messaged back and forth about it a bunch, is that 00:12:35,596 --> 00:12:45,176 [Eric] if you just go by the news headlines, there are really big questions swirling around the economics of the entire AI model marketplace. 00:12:45,176 --> 00:12:45,596 [John] Right. 00:12:45,596 --> 00:13:10,216 [Eric] As you have these entrants that come in with frontier, near frontier performance at a fraction of the cost, is that going to... How is that going to impact the price for frontier models? How is that gonna impact, you know, their ability to generate revenue after raising so much money? How is that gonna impact the way that, you know, people think about using different models, um, for different types of work? Um, 00:13:11,336 --> 00:13:22,426 [Eric] and so I think that, that those dynamics can create some FOMO or questions around, "Well, am I just way behind if I'm not migrating everything- 00:13:22,426 --> 00:13:22,436 [John] Right 00:13:22,436 --> 00:13:24,316 [Eric] ... over to Kemi- [laughs] 00:13:24,316 --> 00:13:25,136 [John] K3, yeah. 00:13:25,136 --> 00:13:28,496 [Eric] [laughs] MiK3 or GLM 52? Um, 00:13:29,536 --> 00:13:32,986 [Eric] and I think for a lot of people who maybe even use AI every day, 00:13:34,596 --> 00:13:39,676 [Eric] it, that's kind of a difficult question to answer, right? There's a lot of ambiguity there, right? 00:13:39,676 --> 00:13:39,716 [John] Right. 00:13:39,716 --> 00:13:41,255 [Eric] What's the advantage of switching, et cetera? 00:13:42,296 --> 00:13:42,536 [Eric] So 00:13:44,496 --> 00:13:47,166 [Eric] I'm gonna give you the 1st word on this. 00:13:47,166 --> 00:13:47,596 [John] Excellent. 00:13:47,596 --> 00:13:56,236 [Eric] Uh, and you think that people should worry about using the latest AI models, and I'm really interested in why. 00:13:56,236 --> 00:13:56,396 [John] Yeah. 00:13:57,656 --> 00:14:02,776 [John] I think, I think I have 3 reasons. One of them, I'm, I'm gonna use one of your own weapons against you. 00:14:02,776 --> 00:14:05,356 [Eric] [laughs] I love this. Okay, great. 00:14:05,356 --> 00:14:09,895 [John] All right. So Eric, if you don't know, loves mental models. 00:14:09,896 --> 00:14:10,256 [Eric] Hmm. 00:14:11,316 --> 00:14:11,906 [John] One of the ones- 00:14:11,906 --> 00:14:13,786 [Eric] Yes, as evidenced by our library here. 00:14:13,786 --> 00:14:17,496 [John] [laughs] Yeah, that's right. What, do you... So you remember the medium is the message, right? 00:14:17,496 --> 00:14:18,576 [Eric] Oh, yes, Marshall McLuhan. 00:14:18,576 --> 00:14:26,816 [John] Yeah. Um, so it suggests the form of communication shapes society and human behavior, behavior far more than the actual content of the information conveyed. 00:14:26,816 --> 00:14:27,636 [Eric] Hmm. Yes. 00:14:27,636 --> 00:14:35,685 [John] So I think with these models, I think the medium of, is the message applies in a sense, where because, 00:14:36,916 --> 00:14:40,596 [John] like, there's a, there's a broader medium where the medium's the same as AI. 00:14:40,596 --> 00:14:40,816 [Eric] Yep. 00:14:40,816 --> 00:14:46,016 [John] But there's a more subtle medium where if you're working with a different model, the medium slightly shifts. 00:14:47,916 --> 00:14:55,236 [Eric] Okay. Yes. I now slight- [laughs] I now slightly regret picking my side because I see what you did there. [laughs] 00:14:55,236 --> 00:14:55,716 [John] And, and, and I mean, 00:14:57,016 --> 00:15:00,516 [John] and I think one of the, one of the key things here, 00:15:01,716 --> 00:15:06,116 [John] and this is my practical workflow. I do not just use the latest and greatest for everything that I do. 00:15:06,116 --> 00:15:06,836 [Eric] Mm-hmm. Mm-hmm. 00:15:06,836 --> 00:15:18,076 [John] But my reason to answer yes here is if you, you fail to understand what is possible if you're only ever, like, optimizing for, like, using the- 00:15:18,076 --> 00:15:18,286 [Eric] Yep 00:15:18,286 --> 00:15:19,856 [John] ... cheapest whatever model. 00:15:19,856 --> 00:15:24,116 [Eric] I'm smiling 'cause it's such a good point. [laughs] And I agree with it. [laughs] 00:15:25,396 --> 00:15:30,916 [John] And, and, and it's funny. I mean, I, I did a training yesterday on this with, with, um, like a non-technical team. 00:15:30,916 --> 00:15:31,656 [Eric] Yeah. 00:15:31,656 --> 00:15:34,716 [John] And they don't think about the model picker thing. 00:15:34,716 --> 00:15:34,866 [Eric] Yeah. 00:15:34,866 --> 00:15:47,076 [John] Like, they didn't... Which is fi- Like, it's, it's not buried, but it's not... Like, it's, you know, you kind of launch a new task and you just chat with it, and it's built in a way where, like, I don't know what it's set to. Like, it's just whatever it was set to, right? 00:15:47,076 --> 00:15:47,696 [Eric] Yep. 00:15:47,696 --> 00:15:52,516 [John] So we talked about, like, you can actually adjust it. You know, you can adjust how hard it works. Um- 00:15:52,516 --> 00:15:53,376 [Eric] Yep 00:15:53,376 --> 00:16:06,876 [John] ... which is helpful. But, but my main point in yes is this, this, number one was the you're changing the medium at least slightly. 2, you need to be a- like, and because of that, you need to be able to understand, like, what can you do now that you couldn't do before? 00:16:06,876 --> 00:16:07,736 [Eric] Hmm. 00:16:07,736 --> 00:16:20,256 [John] And then 3 would be as you're using it, as the models get more intelligent, you are, um, prompting... Your behavior changes, right? 00:16:20,256 --> 00:16:20,536 [Eric] Yep. 00:16:20,536 --> 00:16:32,856 [John] Because you need to give it less instruction, typically. Too much, um, noise, too... Yeah. Too much steering, I think is the word that people would use- 00:16:32,856 --> 00:16:33,076 [Eric] Yep 00:16:33,076 --> 00:16:36,026 [John] ... makes it worse. So there's all these weird things 00:16:37,176 --> 00:16:37,486 [John] that, 00:16:38,946 --> 00:16:45,176 [John] that you have to think about and adapt your behavior where it makes it, it actually makes it a little challenging- 00:16:45,176 --> 00:16:45,786 [Eric] Yep 00:16:45,786 --> 00:16:54,836 [John] ... to switch, to switch back and forth. Um, and then my last point is the newest models often do things 00:16:56,136 --> 00:16:58,136 [John] that are surprising. 00:16:58,136 --> 00:16:58,436 [Eric] Hmm. 00:16:58,436 --> 00:17:02,356 [John] And, and, or get better at things that you, maybe you didn't know you needed. 00:17:02,356 --> 00:17:02,436 [Eric] Hmm. 00:17:02,436 --> 00:17:03,226 [John] That's what I would say. 00:17:03,226 --> 00:17:06,086 [Eric] Give, give, give a specific example. 00:17:06,086 --> 00:17:23,015 [John] Yeah. I have an example. So, um, Fable is the latest as of today, Fable 5 from Anthropic, and one of the things it's really good at is calling and instructing of, like, other instances of itself or of a slightly less intelligent model. So we call it orchestration. 00:17:23,016 --> 00:17:23,496 [Eric] Mm-hmm. 00:17:23,496 --> 00:17:26,756 [John] But just imagine you had a really big task. 00:17:26,756 --> 00:17:27,716 [Eric] Yep. 00:17:27,716 --> 00:17:37,556 [John] And rather than, A, giving it to a model and being like, "Hey, do all these things," and it would fall apart. It's just too much. B, if you give it to this new model that's really, really good at planning- 00:17:37,556 --> 00:17:38,436 [Eric] Mm-hmm 00:17:38,436 --> 00:17:47,515 [John] ... it can write a better plan, one, which is cool. But 2, it's actually way more intelligent about bringing out the plan and then divvying out task to- 00:17:47,516 --> 00:17:47,896 [Eric] Yep 00:17:47,896 --> 00:17:51,835 [John] ... workers or sub, like, whatever you, whatever you wanna call them. Let's call them workers. 00:17:52,866 --> 00:17:55,976 [John] And then reading the response back of, like, "You did this, you did this." 00:17:55,976 --> 00:17:56,216 [Eric] Mm-hmm. 00:17:56,216 --> 00:17:59,416 [John] Synthesizing and then come, and then having a result. So- 00:17:59,416 --> 00:17:59,516 [Eric] Yep 00:17:59,516 --> 00:18:04,156 [John] ... so that would be one. Like, you didn't know that you needed that. You didn't know that, you know? 00:18:04,156 --> 00:18:04,636 [Eric] Yes. 00:18:04,636 --> 00:18:05,196 [John] So that's- 00:18:05,196 --> 00:18:07,446 [Eric] Although, yes, that is, that is 00:18:09,236 --> 00:18:14,056 [Eric] a cost optimization measure on the part of Anthropic- 00:18:14,056 --> 00:18:14,316 [John] Mm-hmm 00:18:14,316 --> 00:18:19,296 [Eric] ... because it's cheaper for them to delegate tasks. This is ultimately a good thing because- 00:18:19,296 --> 00:18:19,676 [John] Yeah 00:18:19,676 --> 00:18:37,700 [Eric] ... that theoretically- ... as the, you know, as the market, um, you know, as the market settles into a state or, or progresses towards a state of equilibrium, ideally as Anthropic gets more cost-efficient, you know, the... that means- 00:18:37,700 --> 00:18:37,710 [John] Right 00:18:37,710 --> 00:18:40,860 [Eric] ... that consumers get better. They get more intelligent about routing- 00:18:40,860 --> 00:18:40,870 [John] Right 00:18:40,870 --> 00:18:43,180 [Eric] ... to different models, and so the consumer sort of benefits from that- 00:18:43,180 --> 00:18:43,540 [John] Yeah 00:18:43,540 --> 00:18:45,620 [Eric] ... ultimately in the price that they pay. Um- 00:18:45,620 --> 00:18:56,080 [John] Yeah. And there's one other weird thing, non-intuitive thing about it, is the smartest models know a lot about the less smart models. 00:18:56,080 --> 00:18:56,380 [Eric] Yes. 00:18:56,380 --> 00:18:56,820 [John] Even about- 00:18:56,820 --> 00:18:57,060 [Eric] Yeah 00:18:57,060 --> 00:18:58,160 [John] ... their capabilities. 00:18:58,160 --> 00:18:58,740 [Eric] Yep. 00:18:58,740 --> 00:19:00,480 [John] Which is really interesting. 00:19:00,480 --> 00:19:00,500 [Eric] Yep. 00:19:00,500 --> 00:19:12,940 [John] Because you can give it the flexibility of like, "Hey, I've got this lot of work to do. You figure out, like, A, how to break out the work, B, how to work in parallel, and C, match the model to the work." 00:19:12,940 --> 00:19:13,140 [Eric] Yep. 00:19:14,260 --> 00:19:16,880 [John] And, and the smartest model is very, very, very good at that. 00:19:16,880 --> 00:19:17,320 [Eric] Yep. 00:19:17,320 --> 00:19:30,180 [John] Whereas if you were to ask, uh, one of the less intelligent ones, it's not gonna, it's not gonna do as well. And, and think about hierarchy, right? Like, which works better, like delegating work up or work down? 00:19:30,180 --> 00:19:30,640 [Eric] Yep. 00:19:30,640 --> 00:19:30,790 [John] Like, 00:19:31,860 --> 00:19:32,340 [John] yeah. So- 00:19:33,440 --> 00:19:33,630 [Eric] All right 00:19:33,630 --> 00:19:34,220 [John] ... there you go. 00:19:34,220 --> 00:19:37,360 [Eric] So you should worry about the [laughs]... So you should worry- 00:19:37,360 --> 00:19:38,780 [John] Shouldn't worry about anything. 00:19:38,780 --> 00:19:38,820 [Eric] [laughs] 00:19:38,820 --> 00:19:40,060 [John] But, um [laughs]- 00:19:40,060 --> 00:19:41,720 [Eric] Yes, that's actually a great- 00:19:41,720 --> 00:19:42,790 [John] But, but I think 00:19:44,280 --> 00:19:45,260 [John] experiment, right, is that. 00:19:45,260 --> 00:19:45,820 [Eric] Yeah. Yeah, yeah. 00:19:45,820 --> 00:20:01,240 [John] I would experiment with it. But, but, but from a really practical standpoint, after you feel like you understand what it does, I then practically work with the more mid-tier models on a daily basis and then wait for a failure mode to push me into using- 00:20:01,240 --> 00:20:01,270 [Eric] Hmm 00:20:01,270 --> 00:20:02,280 [John] ... a smarter one. 00:20:02,280 --> 00:20:02,580 [Eric] Okay. 00:20:02,580 --> 00:20:05,700 [John] So I don't actually, like, work that way. I just experiment that way. 00:20:05,700 --> 00:20:05,740 [Eric] Interesting. 00:20:05,740 --> 00:20:06,880 [John] And I think that's- 00:20:06,880 --> 00:20:06,970 [Eric] Yeah 00:20:06,970 --> 00:20:08,400 [John] ... worth differentiating. 00:20:08,400 --> 00:20:08,700 [Eric] Yeah. 00:20:10,580 --> 00:20:15,450 [Eric] All right. Counterargument to, uh, your argument, which I largely [laughs] agree with. 00:20:15,450 --> 00:20:15,480 [John] [laughs] 00:20:15,480 --> 00:20:16,080 [Eric] This is great. 00:20:18,120 --> 00:20:19,060 [Eric] I don't think, 00:20:20,480 --> 00:20:22,190 [Eric] by and large, that, 00:20:24,580 --> 00:20:27,060 [Eric] that the average person should worry about using the latest model. 00:20:28,370 --> 00:20:28,440 [John] Yeah. 00:20:28,440 --> 00:20:34,810 [Eric] And I'll tell you... I'll give you an example of why. And this is... I'm just gonna puff you up, um- 00:20:34,810 --> 00:20:34,810 [John] [laughs] 00:20:34,810 --> 00:20:35,490 [Eric] ... in this example. 00:20:36,820 --> 00:20:42,000 [Eric] But okay, if we think about we're... so we're on f- from Anthropic, we're on Fable. 00:20:42,000 --> 00:20:42,660 [John] Right. 00:20:42,660 --> 00:20:45,180 [Eric] The previous model family was called Opus. 00:20:45,180 --> 00:20:45,780 [John] Right. 00:20:45,780 --> 00:20:48,920 [Eric] The previous model before that was called Sonnet. 00:20:48,920 --> 00:20:49,720 [John] Right. 00:20:49,720 --> 00:20:51,120 [Eric] And the, and then the one- 00:20:51,120 --> 00:20:51,260 [John] Right 00:20:51,260 --> 00:20:53,120 [Eric] ... the family before that was called Haiku. 00:20:53,120 --> 00:20:55,699 [John] That one hasn't been updated in almost a year. Isn't that crazy? 00:20:55,700 --> 00:20:56,220 [Eric] Sonnet? 00:20:56,220 --> 00:20:56,640 [John] Yeah. No. 00:20:56,640 --> 00:20:56,980 [Eric] Is that surprising? 00:20:56,980 --> 00:20:57,840 [John] No, no. Haiku. 00:20:57,840 --> 00:20:58,720 [Eric] Oh, Haiku. Yeah, yeah. 00:20:58,720 --> 00:20:58,800 [John] Yeah. 00:20:58,800 --> 00:20:59,600 [Eric] That's not surprising. 00:20:59,600 --> 00:20:59,700 [John] Yeah. 00:21:00,800 --> 00:21:04,609 [Eric] Uh, for no extra charge, Haiku is really good at certain things. 00:21:04,609 --> 00:21:04,620 [John] Yeah. 00:21:04,620 --> 00:21:08,380 [Eric] It's actually, it's actually great for writing, but we'll do another episode on that. 00:21:08,380 --> 00:21:08,780 [John] Yeah. That'd be good. 00:21:08,780 --> 00:21:09,060 [Eric] Um, 00:21:10,860 --> 00:21:14,140 [Eric] okay. Let's just go back 3 families, which is significant. 00:21:14,140 --> 00:21:14,300 [John] Mm-hmm. 00:21:14,300 --> 00:21:21,200 [Eric] I mean, Sonnet is, is a very different model in terms of capabilities than Fable is, okay? 00:21:21,200 --> 00:21:22,160 [John] Right. 00:21:22,160 --> 00:21:22,320 [Eric] But 00:21:23,820 --> 00:21:24,900 [Eric] if I 00:21:26,480 --> 00:21:31,460 [Eric] came up with a task and gave it to you and then just someone random who- 00:21:31,460 --> 00:21:31,580 [John] Mm-hmm 00:21:31,580 --> 00:21:35,699 [Eric] ... we pulled off the street, and we gave them Fable and I gave you Sonnet- 00:21:35,700 --> 00:21:35,880 [John] Mm-hmm 00:21:36,980 --> 00:21:54,590 [Eric] ... my guess would be that 9 times out of 10, you would, you would do way better with Sonnet than the random person off the street would do with Fable, which is dramatically, objectively, like, a more powerful, more intelligent- 00:21:54,590 --> 00:21:55,760 [John] Hmm 00:21:55,760 --> 00:21:58,690 [Eric] ... model. And the reason that, the reason I believe that 00:22:00,040 --> 00:22:00,940 [Eric] is because 00:22:02,200 --> 00:22:08,640 [Eric] you have built core skill in using AI, 00:22:09,840 --> 00:22:26,030 [Eric] uh, very well and very intelligently. And I say that intelligently in terms of understanding the value of thinking really well before you ask the model to do something, right? Fable's very- 00:22:26,030 --> 00:22:26,030 [John] Right 00:22:26,030 --> 00:22:27,199 [Eric] ... good at planning. 00:22:27,200 --> 00:22:27,540 [John] Right. 00:22:27,540 --> 00:22:31,960 [Eric] But it's incredible at planning if you're really good at planning. 00:22:31,960 --> 00:22:32,360 [John] Right. 00:22:32,360 --> 00:22:33,200 [Eric] Right? Um, 00:22:34,980 --> 00:22:36,360 [Eric] and so 00:22:37,860 --> 00:22:50,680 [Eric] I, I firmly believe that there's this economy of scale that you get in using AI as a, as a... like, understanding how to use AI as a craft, right? To perform- 00:22:50,680 --> 00:22:50,890 [John] Yeah 00:22:50,890 --> 00:22:51,520 [Eric] ... your craft. 00:22:51,520 --> 00:22:52,260 [John] Right. 00:22:52,260 --> 00:23:04,900 [Eric] And it's a skill that you develop. And we've talked a lo- a lot on the show, and we'll continue to talk a lot on the show, about the, the fact that AI itself does not lead you to that outcome. 00:23:04,900 --> 00:23:05,220 [John] Right. 00:23:05,220 --> 00:23:15,140 [Eric] That that takes a lot of personal discipline and a lot of really hard thinking outside of the use of AI to become a real power user of AI. 00:23:16,200 --> 00:23:16,580 [Eric] And so 00:23:17,760 --> 00:23:27,000 [Eric] the, the reason that I don't think it's worth worrying about, which we'll have some caveats when we land the plane, is that 00:23:28,780 --> 00:24:08,000 [Eric] someone who is extremely good at using AI can use a less capable model and do more than someone who is following the path of least resistance with the latest model available. And in fact, I actually think that using less capable models can even sharpen that skill, right? Is, like, doing things and using things to actually develop ultimately, which this is, this is interesting. You, you brought this up in the way that you use it, in the way that you use different models. But I, I think if I had to summarize, or, or one way to summarize the point that I'm making 00:24:09,080 --> 00:24:09,580 [Eric] is that 00:24:10,620 --> 00:24:27,588 [Eric] what's really important about w- the way that you talked about using sort of mid-level models in your day-to-day work and experimenting with, um, and experimenting with the latest models, the latest frontier models or the latest open weight models- 00:24:27,588 --> 00:24:38,408 [Eric] ... is that you can notice the difference. And, and I think that before worrying about using the latest model, I would really focus on, 00:24:39,928 --> 00:24:40,128 [Eric] on 00:24:42,188 --> 00:25:00,008 [Eric] becoming the type of AI user who can notice the differences. That will truly unlock what... That, I think, is eye-opening when you use one of the more capable models, because you can actually notice, like, the fundamental differences or ultimately fundamental improvements in it. 00:25:00,008 --> 00:25:00,528 [John] Right. 00:25:00,528 --> 00:25:00,948 [Eric] Um, 00:25:02,268 --> 00:25:07,108 [Eric] and so there you go. That's my argument for why I don't think you need to worry about the using the latest model. 00:25:07,108 --> 00:25:07,208 [John] Yeah. 00:25:07,208 --> 00:25:10,768 [Eric] I think you need to get really good with, with really any model. Now- 00:25:10,768 --> 00:25:10,868 [John] Right 00:25:13,288 --> 00:25:19,948 [Eric] ... there are a couple of caveats here. So one is that it... I mean, in large part it depends on what you're doing, right? 00:25:19,948 --> 00:25:20,948 [John] Right. 00:25:20,948 --> 00:25:21,288 [Eric] If 00:25:22,508 --> 00:25:27,088 [Eric] you're just having AI, like, read your Gmail and summarize, like, it doesn't really matter. 00:25:27,088 --> 00:25:27,108 [John] Yeah. 00:25:27,108 --> 00:25:27,548 [Eric] You know? 00:25:27,548 --> 00:25:27,568 [John] Right. 00:25:27,568 --> 00:25:30,608 [Eric] Like, it just... I mean, it's an LLM. It's a large language model. 00:25:31,628 --> 00:25:34,838 [Eric] Use whatever. Like, it's probably gonna do, like, a similar job, right? 00:25:34,838 --> 00:25:36,048 [John] Right. 00:25:36,048 --> 00:25:36,288 [Eric] Um, 00:25:38,048 --> 00:25:41,868 [Eric] the more complex the work and the more things involved, 00:25:42,968 --> 00:25:45,988 [Eric] uh, the, the more capable a, 00:25:47,208 --> 00:25:53,438 [Eric] sort of the latest, like, frontier model is going to be at executing that type of work, right? 00:25:53,438 --> 00:25:53,527 [John] Right. 00:25:53,528 --> 00:25:53,658 [Eric] Um, 00:25:54,808 --> 00:26:10,928 [Eric] but also, the less j- j- like we said, the path of least resistance is that it will just sort of do a bunch of stuff under the hood and do planning and all of that, which can be very helpful. But y- it, it obfuscates a lot of that, um, process under the hood for you. 00:26:10,928 --> 00:26:11,068 [John] Right. 00:26:11,068 --> 00:26:20,568 [Eric] Right? Which again, can limit your ability to, to, to really develop craft in using AI well. So by and large, the... 00:26:21,948 --> 00:26:24,908 [Eric] By and large, I would say, like, I don't think you need to worry. Now, 00:26:27,188 --> 00:26:29,828 [Eric] to land the plane a little bit here, the, 00:26:31,348 --> 00:26:34,808 [Eric] uh... Or to put the... I'll put the landing gear down and then you can actually touch down. 00:26:34,808 --> 00:26:36,628 [John] [laughs] Perfect. 00:26:38,728 --> 00:26:43,038 [Eric] We've been talking about this in the context of, like, choosing a model to do something. 00:26:44,088 --> 00:27:01,488 [Eric] When you start to build AI workflows and AI agents that perform long-running tasks or a series of tasks that sort of ultimately ladder up to a pretty complex job- 00:27:01,488 --> 00:27:01,868 [John] Right 00:27:01,868 --> 00:27:02,138 [Eric] ... um, 00:27:03,388 --> 00:27:35,508 [Eric] you, you intentionally use different models for different things, right? So I'll give you an example. So we, we have a content agent at Vercel that my team, um, has built out and works on, and it does a lot of different things. So for example, it may go do a bunch of research, right? It may, um, draft an outline. It may go try to validate, um, the accuracy of, you know, technical details of a product, right? 00:27:36,798 --> 00:27:36,868 [John] Mm-hmm. 00:27:36,868 --> 00:27:42,218 [Eric] And you don't need to just throw Fable 5 at all of those tasks, right? 00:27:42,218 --> 00:27:42,228 [John] Right. 00:27:42,228 --> 00:27:45,598 [Eric] And in fact, like, you, you don't want to do that because it's- 00:27:45,598 --> 00:27:45,598 [John] Right 00:27:45,598 --> 00:27:54,578 [Eric] ... it's, it is very expensive, right? So if I wanna generate an outline from research that has been collected, I don't need Fable 5 to do that, right? 00:27:54,578 --> 00:27:54,788 [John] Right. 00:27:54,788 --> 00:28:13,517 [Eric] That is a, like, block and tackle LLM. Like, read the skill, all the research is there, and, like, generate a good outline, right? Um, if we're doing complex research, right, like from multiple different s- you know, sources and then trying to summarize that, um, 00:28:14,668 --> 00:28:32,588 [Eric] or even trying to decide, you know, or even having AI make decisions about where to look for additional information or do research publicly in the market, uh, you know, or in the public market or public internet, um, competitive research, et cetera, Fable 5's gonna do a much better job of that, right, because it's a much more capable model. 00:28:32,588 --> 00:28:32,948 [John] Right. 00:28:32,948 --> 00:28:35,848 [Eric] It can delegate sub-agents and all that sort of stuff. And so 00:28:37,148 --> 00:28:42,848 [Eric] I think the... I say that selfishly to reinforce my point in that 00:28:44,367 --> 00:28:53,368 [Eric] you don't need to worry about the latest model because as you build more and more complex agents and workflows, you're going to use different models- 00:28:53,368 --> 00:28:53,608 [John] Right 00:28:53,608 --> 00:29:01,708 [Eric] ... for different things anyways. And that's why I think it's so important to understand and be able to notice the differences between the models so that you can make- 00:29:01,768 --> 00:29:01,778 [John] Right 00:29:01,778 --> 00:29:05,328 [Eric] ... intelligent decisions about which model makes sense for which part of the workflow. 00:29:05,328 --> 00:29:09,958 [John] Yeah. So my mental model for this... So 2 things. One, 00:29:11,028 --> 00:29:17,368 [John] um, my, 1st my mental model is thinking about who would I hire to do this if it were a person? 00:29:17,368 --> 00:29:18,648 [Eric] Hmm. 00:29:18,648 --> 00:29:24,567 [John] And if I think like, "Oh, this is, like, entry level," like, that's pretty good guidance to, like, which model to use. 00:29:24,568 --> 00:29:25,348 [Eric] It's great. 00:29:25,348 --> 00:29:25,438 [John] Um, 00:29:26,548 --> 00:29:28,788 [John] 2, reaction to what you were saying with research or whatever, 00:29:30,128 --> 00:29:35,668 [John] I do think it's still worth trying latest frontier models on things that you're doing successfully. 00:29:35,668 --> 00:29:35,768 [Eric] Yep. 00:29:37,108 --> 00:29:47,028 [John] Mostly as far as auditing, 'cause it's a... Like, Fable knows a lot of the capabilities, like probably more than you or I, of the models below it. 00:29:47,028 --> 00:29:47,448 [Eric] Yep. 00:29:47,448 --> 00:29:59,708 [John] So it's great to be like, "Hey, audit this skill or workflow or whatever. These are the models I'm running it with. Like, where are the gaps? How could I tighten it? How could I make it more efficient?" Whatever. That's really useful. 00:29:59,708 --> 00:30:00,127 [Eric] Yep. 00:30:00,128 --> 00:30:11,728 [John] And then occasionally I've run into things where I'm like, "You know what? I don't think I need, like, a frontier, like the latest and greatest on this thing, but I'm gonna run it and just see if I think it's better." 00:30:11,728 --> 00:30:12,127 [Eric] Mm-hmm. 00:30:12,128 --> 00:30:19,728 [John] And if it's better, I'll ask it why it's better and we'll try to apply the learning back down, 'cause I don't actually wanna spend, you know, the extra money- 00:30:19,728 --> 00:30:20,108 [Eric] Yes 00:30:20,108 --> 00:30:20,848 [John] ... on, on running it. 00:30:20,848 --> 00:30:21,377 [Eric] Yep. 00:30:21,377 --> 00:30:30,440 [John] Um, so I, I think there's some really interesting trade-offs there. Um, and I'm starting to think through this a lot as far as who... 00:30:30,440 --> 00:30:33,640 [John] Who do you want training? Who do you want that training role? 00:30:33,640 --> 00:30:33,760 [Eric] Hmm. 00:30:33,760 --> 00:30:36,199 [John] That should be... That's a smart role. And if that trainer's- 00:30:36,200 --> 00:30:37,160 [Eric] Like, in your business? 00:30:37,160 --> 00:30:37,930 [John] Yeah, in your business. 00:30:37,930 --> 00:30:37,940 [Eric] Yep. 00:30:37,940 --> 00:30:41,800 [John] And if that trainer's continuing to get better and knows a lot about the capabilities of the- 00:30:41,800 --> 00:30:42,180 [Eric] Yep 00:30:42,180 --> 00:30:53,640 [John] ... the things underneath it. So that's one. And then a practical one, back to, like, I'm an everyday... Like, maybe I consider myself a power user, but, but I'm more of a knowledge worker than a- 00:30:53,640 --> 00:30:54,200 [Eric] Mm-hmm 00:30:54,200 --> 00:31:07,260 [John] ... you know, programmer. Um, I did this yesterday and, and it's very simple. Did it with a client, and it went awesome. So we had... I was like, "Hey, give me an idea that you're thinking about right now." 00:31:07,260 --> 00:31:08,020 [Eric] Mm-hmm. 00:31:08,020 --> 00:31:20,020 [John] He's like, "Oh, I've got a presentation this, you know, next week." I was like, "Okay, great. Like, give me the presentation." So it's pulled up a slide with a diagram of, like, a system that they were developing. 00:31:20,020 --> 00:31:20,340 [Eric] Mm-hmm. 00:31:20,340 --> 00:31:26,840 [John] And they talked about the slide for 10 minutes. I was like, "Okay, watch this." So clipped, like, 10 minutes of the transcription from the meeting notes. 00:31:26,840 --> 00:31:27,680 [Eric] Mm-hmm. 00:31:27,680 --> 00:31:32,069 [John] Gave the image of the slide. And I was like, "Build this. Fake the data." 00:31:32,069 --> 00:31:32,100 [Eric] [laughs] 00:31:32,100 --> 00:31:33,200 [John] Like, it didn't have to be real data. 00:31:33,200 --> 00:31:33,560 [Eric] Yeah, yeah. 00:31:33,560 --> 00:31:34,280 [John] You know, in 10 minutes- 00:31:34,280 --> 00:31:34,350 [Eric] Yeah 00:31:34,350 --> 00:31:35,979 [John] ... later they're like, "How did it do that?" 00:31:35,979 --> 00:31:36,030 [Eric] [laughs] 00:31:36,030 --> 00:31:36,520 [John] Like, what? 00:31:36,520 --> 00:31:37,260 [Eric] Yeah, totally. 00:31:37,260 --> 00:31:38,280 [John] Like, what? You know. 00:31:38,280 --> 00:31:38,530 [Eric] Totally. 00:31:38,530 --> 00:31:42,450 [John] Like, e- even, like, specific words. Like, "I didn't say that." Like, how did... It was like, "Oh, it was in your slide." 00:31:42,450 --> 00:31:42,460 [Eric] Yep. 00:31:42,460 --> 00:31:44,120 [John] I was like, "Oh." Like- 00:31:44,120 --> 00:31:44,380 [Eric] Yep 00:31:44,380 --> 00:31:46,280 [John] ... that's... It's pretty incredible. 00:31:46,280 --> 00:31:58,240 [Eric] Yeah. I agree. I agree. I, I think the last, the last thought I'll share in terms of practical advice is, is circling back to the cost point that you just mentioned. 00:31:59,480 --> 00:32:00,810 [Eric] You know, if you're just on a plan that 00:32:02,040 --> 00:32:05,900 [Eric] gives you access to the latest frontier models and you pay the same rate every month- 00:32:05,900 --> 00:32:06,120 [John] Mm-hmm 00:32:06,120 --> 00:32:08,670 [Eric] ... right, there's some sort of subsidization happening there. 00:32:08,670 --> 00:32:08,940 [John] Yeah. Right. 00:32:08,940 --> 00:32:15,290 [Eric] And if that's the case, then you can probably just set it to use the latest and greatest and it doesn't really matter, right? 00:32:15,290 --> 00:32:15,340 [John] Yeah. 00:32:15,340 --> 00:32:16,740 [Eric] If you're not hitting your usage limits. 00:32:16,740 --> 00:32:17,640 [John] Yeah, sure. Yeah. 00:32:17,640 --> 00:32:22,320 [Eric] This becomes... I, I still would stand by what I said in terms of noticing the differences- 00:32:22,320 --> 00:32:22,330 [John] Right 00:32:22,330 --> 00:32:23,500 [Eric] ... I think is important. 00:32:23,500 --> 00:32:23,839 [John] Yeah, for sure. 00:32:23,840 --> 00:32:25,900 [Eric] Like, for your long-term ability to use- 00:32:25,900 --> 00:32:26,120 [John] For sure 00:32:26,120 --> 00:32:26,840 [Eric] ... AI well. 00:32:26,840 --> 00:32:27,410 [John] Mm-hmm. 00:32:27,410 --> 00:32:27,820 [Eric] Um, 00:32:28,940 --> 00:32:32,320 [Eric] but this really becomes interesting when cost does- 00:32:32,320 --> 00:32:32,330 [John] Yeah 00:32:32,330 --> 00:32:34,720 [Eric] ... become an issue, right? Because just pointing the latest- 00:32:34,720 --> 00:32:34,920 [John] Yeah 00:32:34,920 --> 00:32:39,240 [Eric] ... you know, closed weight frontier model at everything is very expensive. 00:32:39,240 --> 00:32:39,760 [John] Yeah. 00:32:39,760 --> 00:32:43,540 [Eric] Right? Um, and, and so 00:32:44,800 --> 00:32:58,600 [Eric] I think that as you sort of progress in your level of complexity and you re- reach, like, usage limits, at that point, it becomes incredibly valuable to understand the differences between the models, know when it makes sense to use- 00:32:58,600 --> 00:32:58,710 [John] Right 00:32:58,710 --> 00:33:07,820 [Eric] ... a frontier model, and then, like you said, also get into a cycle of experimenting where the models are themselves getting better at, at delegating- 00:33:07,820 --> 00:33:07,980 [John] Right 00:33:07,980 --> 00:33:09,040 [Eric] ... like, across models, right? 00:33:09,040 --> 00:33:09,720 [John] Right. 00:33:09,720 --> 00:33:13,040 [Eric] The other thing I would say is that, um, 00:33:14,940 --> 00:33:20,760 [Eric] there are new models released all the time. And so when we think about Fable delegating tasks to, 00:33:22,100 --> 00:33:24,640 [Eric] uh, you know, earlier model families- 00:33:24,640 --> 00:33:24,660 [John] Yeah 00:33:24,660 --> 00:33:25,720 [Eric] ... or less capable models, right? 00:33:25,720 --> 00:33:26,240 [John] Right. 00:33:26,240 --> 00:33:29,360 [Eric] You're sort of locked into the Anthropic ecosystem there, right? 00:33:29,360 --> 00:33:29,810 [John] Sure. 00:33:29,810 --> 00:33:34,460 [Eric] So another thing that I think makes a lot of sense is to use something like Vercel's AI Gateway 00:33:35,500 --> 00:33:41,020 [Eric] that gives you access to all of the models, right, so that you can also test this across- 00:33:41,020 --> 00:33:41,069 [John] Yeah 00:33:41,069 --> 00:33:54,790 [Eric] ... other frontier closed weight providers like GPT, um, or, you know, OpenAI's GPT families, or, uh, of course, open weight models, you know, like, um, KIMI or GLM as well. 00:33:54,790 --> 00:33:54,880 [John] Yeah. 00:33:54,880 --> 00:34:03,760 [Eric] Um, because if you get to the point where you're looking at optimizing cost, y- almost certainly you're going to expand beyond just using a single- 00:34:03,760 --> 00:34:03,980 [John] Yeah 00:34:03,980 --> 00:34:13,840 [Eric] ... lab. Um, which again, that's not bad, but, uh, you can get a lot more bang for your buck if you give yourself the option to experiment across different providers- 00:34:13,840 --> 00:34:13,849 [John] Right 00:34:13,849 --> 00:34:15,160 [Eric] ... and not just sticking with, like, one- 00:34:15,160 --> 00:34:15,349 [John] Right 00:34:15,349 --> 00:34:16,340 [Eric] ... single provider. 00:34:16,340 --> 00:34:23,820 [John] Yeah, and one really practical example, like if I was a security researcher right now and didn't have clearance to use, like, Mythos or- 00:34:23,820 --> 00:34:24,020 [Eric] Mm-hmm 00:34:24,020 --> 00:34:29,060 [John] ... GPT's CyberSec model, I would absolutely be using KIMI K3- 00:34:29,060 --> 00:34:29,480 [Eric] Yep 00:34:29,480 --> 00:34:33,200 [John] ... through something like a Vercel gateway or some sort of gateway. 00:34:33,200 --> 00:34:33,380 [Eric] Yep. 00:34:33,380 --> 00:34:37,340 [John] Um, because that has the least guardrails right now- 00:34:37,340 --> 00:34:37,480 [Eric] Yep 00:34:37,480 --> 00:34:40,300 [John] ... as far as doing real cybersecurity research. 00:34:40,300 --> 00:34:43,940 [Eric] Totally. Totally. And I think the, uh, I think the other... 00:34:46,280 --> 00:34:51,389 [Eric] For sure, there are less guardrails. I think the other thing is that you may discover that 00:34:52,460 --> 00:34:57,960 [Eric] for certain tasks, models that are a 10th of the price do just as good of a job for you. 00:34:57,960 --> 00:34:58,320 [John] Oh, definitely. 00:34:58,320 --> 00:34:58,960 [Eric] Right? 00:34:58,960 --> 00:34:59,380 [John] Yeah. 00:34:59,380 --> 00:34:59,690 [Eric] Um, 00:35:01,380 --> 00:35:04,730 [Eric] and so ultimately, I think we're saying experiment, right? 00:35:04,730 --> 00:35:04,760 [John] Sure. 00:35:04,760 --> 00:35:06,330 [Eric] Maintain the craft of using AI- 00:35:06,330 --> 00:35:06,330 [John] Right 00:35:06,330 --> 00:35:11,600 [Eric] ... and experiment. So there you go. Don't worry about the latest model, but definitely experiment with it. [laughs] 00:35:11,600 --> 00:35:12,520 [John] Right. 00:35:13,760 --> 00:35:14,660 [John] Great conclusion. 00:35:14,660 --> 00:35:15,940 [Eric] All right. We'll catch you on the next one. 00:35:22,240 --> 00:35:24,690 [Eric] [upbeat music]
