
Why Scaling Prediction Cannot Create Intelligence - Alexander Mattick
Get every episode summarized
Each time Machine Learning Street Talk (MLST) publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
Alexander Mattick is a researcher at Fraunhofer IIS and a PhD researcher at the University of Technology Nuremberg (UTN), and a regular on Yannic Kilcher's Discord. He first came on MLST in 2022, after helping research the Yann LeCun and Randall Balestriero episode on interpolation.
SPONSOR:
---
Cyber Fund built the Monastery to help founders ship products that were impossible a year ago. Applications for Batch 1 are now open.
Apply now: https://cyber.fund
---
Alexander treats inference as the thread running through modern machine learning: once you have a model, what does it cost to get an answer out of it? He works through Monte Carlo, GFlowNets, energy-based models, diffusion, normalising flows and flow matching, with four short explainers he recorded himself. He is blunt about energy-based models: you can sample from them in principle, but it is rarely worth the compute. JEPA and "world model", he says, are closer to branding than to technical categories.
Next: theories of deep learning, none of which he thinks predicts enough yet to guide practice, then reinforcement learning.
---
0:00 Cold open: information is expensive
0:51 Welcome back, Alexander Mattic
2:08 Alexander's research background
2:50 Inference: densities, sampling and Monte Carlo
6:42 GFlowNets, energy functions and MCMC
9:45 Explainer: energy-based models
11:03 Why model a density at all?
17:30 From learned energies to flow matching
25:08 Explainers: diffusion and normalising flows
28:33 Are energy-based models generative?
33:22 JEPA, contrastive learning and collapse
41:13 Why non-language modalities need flows
44:51 Inference as search: branch and bound
49:43 Q-learning and delayed consequences
55:14 Flow matching, optimal transport, Fokker-Planck
1:00:03 Explainer: flow matching
1:01:49 AlphaFold, latents and scale versus architecture
1:07:52 Two families of deep learning theory
1:15:04 What a good theory would predict
1:23:53 The manifold hypothesis and compression
1:28:25 Is reward enough?
1:32:01 Control theory versus reinforcement learning
1:37:22 The Bitter Lesson and expensive information
1:42:08 Constrained RL: the constrained MDP toolbox
1:50:12 Creativity as constrained search
1:55:44 Reality is protean: when abstractions hold
2:00:32 What is a world model?
2:04:38 Prediction is not control
2:08:13 Robot demos, MPC and reliability
---
REFERENCES:
[6:55] GFlowNets (Bengio et al., 2021)
https://arxiv.org/abs/2106.04399
[38:46] Contrastive Self-Supervised Learning (Anand, 2020)
https://ankeshanand.com/blog/2020/01/26/contrative-self-supervised-learning.html
[38:56] LeJEPA (Balestriero and LeCun, 2025)
https://arxiv.org/abs/2511.08544v3
[47:10] RL for Node Selection in Branch-and-Bound (Mattick)
https://openreview.net/forum?id=0ez68a5UqI
[56:20] Flow Matching for Generative Modeling
https://arxiv.org/abs/2210.02747v2
[1:12:41] Disentangling feature and lazy training in deep neural networks
https://arxiv.org/abs/1906.08034v4
[1:31:05] Reward is enough (Silver)
https://doi.org/10.1016/j.artint.2021.103535
[1:35:12] Learning ReLU networks to high uniform accuracy is intractable (Berner et al.)
https://arxiv.org/abs/2205.13531v2
[1:40:20] Dota 2 with Large Scale Deep RL
https://arxiv.org/abs/1912.06680v1
[1:45:41] Constrained Update Projection for Safe Policy Optimization (Yang et al., 2022)
https://arxiv.org/abs/2209.07089
[1:46:11] SafeMPO (ICLR 2026)
https://openreview.net/forum?id=1m0EU6QXj6
[1:50:17] Why Creativity Cannot Be Interpolated
https://archive.mlst.ai/paper/why-creativity-cannot-be-interpolated/
[1:51:39] Invalid Action Masking (Huang and Ontañón)
https://arxiv.org/abs/2006.14171
[2:00:04] Probability Theory: The Logic of Science (Jaynes, 2003)
https://www.cambridge.org/core/books/probability-theory/9CA08E224FF30123304E6D8935CF1A99
[2:01:53] Training Agents Inside of Scalable World Models (Hafner et al., 2025)
https://arxiv.org/abs/2509.24527v1
[2:03:43] World Models (Ha and Schmidhuber, 2018)
https://arxiv.org/abs/1803.10122v4
Get every episode summarized
Each time Machine Learning Street Talk (MLST) publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from Machine Learning Street Talk (MLST)

When AI Research Starts Moving Faster Than Human Research - Zhengyao Jiang
Machine Learning Street Talk (MLST)

How Deep Learning Finally Cracked Messy Tables - Frank Hutter
Machine Learning Street Talk (MLST)

How Physical AI Learns Across Language, Video and Action — Ming-Yu Liu
Machine Learning Street Talk (MLST)

Speech Recognition Is Not a Solved Problem — Pavan Kumar Reddy
Machine Learning Street Talk (MLST)