
About this episode
I think the future is cheaper and Open Source SOTA models combined with context, not custom, narrow models.
Become a Member: https://danielmiessler.com/upgrade
See omnystudio.com/listener for privacy information.
Get every episode summarized
Each time Unsupervised Learning publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
41 searchable segments. Every word is indexed and playable.
Full transcript
Unsupervised Learning — Why I Believe in SOTA Models Over Custom Ones. Machine-transcribed; use the interactive transcript above to jump the player to any line.
I'm not completely sure, I'm right about this, but I've never been a big believer in training custom models. I've also never believed in fine tuning. Going all the way back to 2023, my intuition has always pushed me towards the best state-of-the-art model possible combined with context management. I just finally crystallized my reasoning around this. Anytime you think you're using a small model for a small task, there's usually a whole lot more going into a given decision than just that individual area of expertise. For example, labeling emails, writing reports, processing security events, searching for threats on our network. On one hand, I think these are specialized, but the fact is the smarter and more experienced a human is, who has this expertise, the better job they're gonna do. This is because most specialized tasks still benefit from the general life experience of the person doing the execution. This is why I think the future is not a whole bunch of extremely small specialized models
throughout the enterprise. I think what's far more likely is more of an opus on a high-cube model, where the best of the best just keeps coming down in price, including going into open source. And those smaller models are using conjunction with context to perform all the different tasks in an organization at much lower cost. But I think there'll still be extremely general models, not tiny and narrow custom ones. I think the TLDR here is, when you think you're doing a narrow task, that narrow task is actually benefiting from a ton of general experience. And I think this applies to humans, and I think it also applies to models. I'm not completely convinced of this. I'm about 70% sure. But yeah, I think this is the way it's gonna go.
More episodes
More from Unsupervised Learning

Most Companies Aren't Anywhere Near Ready for AI
Unsupervised Learning

We're All Building a Single Digital Assistant
Unsupervised Learning

Why AI Will Replace Knowledge Workers
Unsupervised Learning

AI Quality Inversion
Unsupervised Learning