
π EP 368: Trump War Grok Consultation & "Mode-Hopping" Pre-Training Discovery
Get every episode summarized
Each time AI Fire Daily publishes, we email you a written briefing from the transcript β the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
βSave with digital deals at Vons and Albertsons. This week at Vons and Albertsons, Ginny O, 93% lean ground turkey is 399 per pound, sold in a 3 pound tray with digital coupon limit two packages. And a 1 pound package of strawberries is 199 with digital coupon.βFrom the transcript
Reports revealed President Trump consulted xAI's Grok chatbot regarding political dynamics in Venezuela ahead of strategic military operations, highlighting AI's growing footprint in executive decision-making. Meanwhile, a machine learning research paper uncovered "mode-hopping" during LLM pre-training, proving that additional training tokens can cause drastic, non-linear collapses in downstream reasoning and fine-tuning transferability.
Weβll talk about:
- Details outlining executive use of xAI's Grok for geopolitical assessment, underscoring risks when decision-makers rely on LLM outputs for high-stakes analysis.
- Study analyzing OLMo3-32B demonstrating how additional training tokens can trigger temporary 0% performance drops and degrade downstream post-training adaptation.
- Tech leaders from OpenAI, Google, Anthropic, Meta, Nvidia, and xAI signing voluntary commitments for external audits on frontier models.
- AI agent startup expanding tools that sync CAD models with complex engineering requirements for defense and automotive manufacturing.
Keywords: Grok Gov, mode hopping, LLM pre-training, Superintelligence Accord, Flow Engineering.
Links:
- Newsletter: Sign up for our FREE daily newsletter.
- Our Community: Get 3-level AI tutorials across industries.
- Join AI Fire Academy: 700+ advanced AI workflows ($14,500+ Value)
Our Socials:
- Facebook Group: Join 299K+ AI builders
- X (Twitter): Follow us for daily AI drops
- YouTube: Watch AI walkthroughs & tutorials
Get every episode summarized
Each time AI Fire Daily publishes, we email you a written briefing from the transcript β the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
317 searchable segments. Every word is indexed and playable.
Full transcript
AI Fire Daily β π EP 368: Trump War Grok Consultation & "Mode-Hopping" Pre-Training Discovery. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Save with digital deals at Vons and Albertsons. This week at Vons and Albertsons, Ginny O, 93% lean ground turkey is 399 per pound, sold in a 3 pound tray with digital coupon limit two packages. And a 1 pound package of strawberries is 199 with digital coupon. Plus Kellogg's giant size cereal, 22.6 to 29.5 ounces, selected varieties are 299 each, with digital coupon limit three. Hurry in! Visit Vons or Albertsons.com for more deals and ways to save. It's not just electric, it's Toyota Electric. And during Toyota's easy choice sales event, going electric is easier than ever. Right now, save instantly on a new Toyota EV with special incentive programs for first time EV buyers and prior EV owners. Choose the sleek busy, the adventurous busy woodland or the sporty CHR. But don't wait! These special savings won't last. We make it easy. Toyota, let's go places.
We live in a bizarre technological paradox today. AI chatbots are now actively advising presidents on global military operations. Yeah, it really is like a wild time to be alive. We desperately want these systems to act as perfect oracles. Right. Mathematically speaking, they are still highly unpredictable experiments. They really are. And yet, under the hood, these exact same models randomly fail, like they will suddenly forget how to do basic math. Exactly. Yeah, for absolutely no apparent reason. So welcome to the deep dive. Our mission today is quite simple. We want to extract the signal from the constant noise. Because there is so much noise out there right now. There really is. We are analyzing a stack of reports on new AI developments today. We are going to examine AI suddenly leap into high stakes geopolitics. And from there, we explored the bizarre behavioral quirks these models are developing. We also unpack the sheer security nightmare of watermarking synthetic biology.
And finally, we will explore a massive new discovery in AI training. Which is a breakthrough that really challenges everything we assumed about how models learn. Right. So let us get into it. We really are seeing a massive fundamental shift right now. I mean, we are moving far beyond models just generating polite corporate emails. Yeah, AI is actively shaping the physical world at the highest levels. The stakes have escalated incredibly fast. They absolutely have. According to recent reporting from time, this shift is already happening. Back in December 2025, President Trump held a highly private meeting. Right. He reportedly spent hours consulting Elon Musk's GROC AI. They were directly discussing the complex geopolitical situation of Venezuela. Which is just fascinating. Specifically, he asked the model a very high stakes political question. He wanted to gauge how the Venezuelan public would actually react. He asked what happens if President Nicholas Maduro were captured.
And GROC responded to that prompt with absolute supreme confidence. Yeah, it did. It reportedly told the President that Maduro was deeply unpopular. It suggested that many locals would likely celebrate his sudden removal. Which brings us to the actual events of January 3, 2026. The US launched an operation in Venezuela. And successfully captured Maduro. Right. And just as the model predicted, there were indeed public celebrations. Time reported a very interesting detail about the immediate aftermath. Oh, yeah. A source stated Trump viewed GROC as highly capable afterward. Now, let's be extremely clear about the reported facts here. The reporting does not state that GROC caused the invasion. Exactly. The AI did not make the final military decision. It was utilized as one specific source of geopolitical information. Right. It was just used to weigh the potential political consequences of action. But it highlights a massive transition in modern global governance. It's kind of like relying on a supercomputer to read a geopolitical crystal ball.
Yet you are asking a machine to predict human emotion on a massive scale. And the Pentagon is already using something called GROC internally. They are deploying it across various defense and military operations. The connections run very deep across the entire tech industry too. Open AI and Anthropic also maintain significant relationships with the US defense establishment. This integration is rapidly becoming the new normal for global superpowers. Which brings up a very real mechanical concern about these systems. When a human intelligence analyst writes a briefing, they include self-doubt. Right. They use confidence intervals. Exactly. They flag when data is highly uncertain. Large language models do not possess that kind of human hesitation. Yeah, that is the core structural problem with LLMs. They are ultimately just probabilistic token prediction engines. They are mathematically designed to generate the most likely next word. They are not actually designed to evaluate fundamental objective truth. So they deliver every single answer with the exact same unyielding authority.
But given how AI can hallucinate, aren't we risking international crises if a model is just overly confident in its geopolitical assumptions? Yes, absolutely. The distinct danger is that AI models become the ultimate yesman. Right. If a defense system accepts a confidently hallucinated output, it's fact that cascading effects are terrifying. It completely removes the crucial human hesitation needed to prevent catastrophic escalation. So hallucinated confidence could literally spark an international crisis if we aren't careful. Right. And that exact vulnerability is shaking up the tech world. Security is suddenly the loudest conversation in every single boardroom. If world leaders are trusting these models with matters of war, the obvious question becomes security. How secure is the digital foundation we are currently building on? Yeah. Because recent reports suggest we are building on absolute quicksand. The major tech companies certainly realize how precarious this all is.
Open AI, Google and Thropic, Meta, Nvidia and XAI recently made a move. Right. The new accord. Yeah. They all signed the new White House Super Intelligence Accord. They voluntarily committed to rigorous outside audits for their advanced systems. It was presented to the public as a massive step forward. Tech CEOs proudly share the accord across their social media platforms. Oh, yeah, they did. But then users instantly spotted a typo on the official signatory page. Yeah, that really happened. Nothing kills the Super Intelligence vibe faster than a basic human typo. Exactly. It is a deeply ironic for a document governing the most advanced intelligence. It really is. But the actual security threats surrounding these models are not funny at all. No, not at all. Google just reported some highly alarming numbers on the cybersecurity front. Publicly reported software vulnerabilities literally doubled in August alone. They hit 10,740 distinct vulnerabilities in a single month.
Hackers are increasingly using AI agents to weaponize newly patched flaws. It is wild. The defensive tools are getting smarter, but the digital weapons are scaling faster. It is a terrifying invisible arms race happening in the background. And it is not just limited to traditional software vulnerabilities anymore. Right. We also have a rising threat on the synthetic biological front. AI-generated synthetic biology is getting much harder to effectively track. Deep mind recognized this specific danger and tried to build a solution. They recently built a new tool called SynthID Bio. It is a system specifically designed to tag synthetic biology sequences. If you were listening and wondering how you watermark physical biology, it is fascinating. DNA is essentially just biological code. Exactly. And much like human language, there are different ways to write the same thing. Right. In biology, certain different DNA sequences will produce the exact same protein. These are called synonymous codons. Okay. SynthID Bio subtly alters those specific synonymous codons.
So it inserts a hidden digital barcode into the physical biological design. And crucially, it does this without hurting the actual lab performance. Yeah, the organism functions identically, but it carries a hidden cryptographic signature. This is incredibly vital for the future of global biosecurity. We desperately need a standardized way to track lab created biological materials. We really do. With Deep Mind's SynthID Bio, is watermarking biology actually going to stop a determined rogue lab from doing damage? Not entirely. It is really more about creating significant friction in the system. Okay. Watermarking establishes a very clear trail of breadcrumbs back to the origin. A highly determined rogue lab could potentially try to strip the watermark out. Right. But it forces them to work much harder and leaves complex forensic clues behind. Essentially it's about tracking the origin, not completely stopping a determined rogue lab. Exactly. We are trying to build fences in the dark right now.
The models are evolving wildly and they are developing bizarre new behaviors. Yeah, their outputs morph without us explicitly coding them to do so. We used to rely on very simple tells to spot AI text. Oh yeah. The classic M-dash giveaway was the gold standard for a while. If you saw a paragraph flooded with M-dashes, it was probably a chatbot. But the newer models are rapidly changing their internal linguistic habits. Opus 5.5 has developed a brand new, highly specific behavioral tell. Right. It used the exact phrase, this matters 13,000 times in recent data. That is so wild to me. It organically adopted a human rhetorical crutch all on its own. And the new astramodel has developed its own weird linguistic tells too. They drift over time in highly unpredictable and frustrating ways. Honestly, I still wrestle with prompt drift myself. Yeah, it happens to everyone. You get a complex prompt working perfectly for your workflow one week. Then the underlying model subtly shifts and the output is suddenly completely useless.
It is incredibly frustrating for daily users and enterprise developers alike. These advanced models are not static pieces of traditional software at all. Right. They operate much more like living, breathing digital ecosystems. And sometimes those fragile ecosystems just completely collapse in on themselves. Which brings us perfectly to our final and perhaps most unsettling topic. We are currently building our infrastructure on top of these shifting tools. We are. For instance, Grock primary bot proactively finds complex tasks it can handle. It actually suggests ways to help you without hitting your usage limits. And CloudFlair-CLEF just introduced two brand new open source AI decision models. They deliver incredibly fast probability-based answers for AI agents and automated trading bots. We are deploying these models to make thousands of autonomous decisions daily. But AI researchers just discovered a massive fundamental flaw in their architecture. Yeah. We do not actually understand how they truly learn to reason.
It all comes down to the pre-training phase of model development. This is where the magic and clearly the mystery actually happens. It is. The concept of pre-training is usually quite simple. Feeding an AI vast amounts of text to learn language rules. Perfect. Now the baseline assumption in the tech industry has always been linear. We assumed more pre-training data inherently equals better generalization and deeper reasoning. Just keep feeding the machine more data and it gets smarter. Right. But a new research paper found something surprisingly unstable inside LLMs. They call this deeply chaotic internal phenomenon mode hopping. The specific data they found in the ALMO3-32B model is absolutely wild. We really need to look closely at the actual performance numbers here. Okay, let's unpack the timeline of this model's data training. Yeah. At 2.17 trillion tokens of training data, the model had 81% accuracy. That is a very solid expected score for that exact stage of pre-training.
But then at 2.19 trillion tokens, it dropped to 0% accuracy. Complete catastrophic performance collapse. 0%. Functionally forgot how to do everything it had just learned. But then just slightly later at 2.21 trillion tokens, it jumped back up. It hit 81.7% accuracy as if nothing had ever happened. Well, imagine scaling to a billion queries on a live enterprise system. And the model suddenly hits one of those mathematical dead zones. Right. The entire automated system just breaks completely without any warning whatsoever. Within just 40 billion more training tokens, the core performance collapsed to 0. Then it almost fully recovered its complex reasoning abilities just as fast. Yeah. It is literally like stacking Lego blocks of data and the tower randomly vanishing. To understand this chaotic behavior, researchers carefully examine different model checkpoints. A checkpoint is basically a saved structural snapshot of the model during its training.
They compared an earlier 4.5 trillion token checkpoint with a later 4.9 trillion token checkpoint. They ran both of them through a brutally tough benchmark after doing math fine tuning. Logically, you would completely expect the later model to be much smarter. It has processed significantly more high quality data than the earlier version. Right. But the 4.5 trillion checkpoint scored 36.3% on the GPQA benchmark. And the later 4.9 trillion checkpoint only scored 29.8%. The later model was actually a much worse starting point for fine tuning. Wow. It fundamentally lost its robustness. The model literally saw more data, yet its foundational reasoning became measurably dumber. So what is mathematically happening inside the neural network to cause this collapse? The researchers believe different internal learning strategies constantly compete during the pre-training phase. It is a brutal internal tug of war for the model's actual brain.
We often visualize model training as walking down a sloping mountain range. You are trying to find the absolute lowest point in the valley. That lowest point represents the lowest error rate. Exactly. Sometimes the model leans on deep, highly generalizable reasoning to get there. It actually builds a structural, logical understanding of the complex problem. But other times, easier shortcut-like patterns take over the neural network. The model finds an incredibly cheap mathematical trick to just guess the right answer. Yes. Why would an AI suddenly abandon deep reasoning to take a shortcut after it already learned how to solve the problem? Because the only mathematical goal of training is reducing the immediate error rate. The model discovers a highly superficial pattern that works perfectly for a specific batch of data. Okay. So it lazily adopts that shortcut, entirely overriding the deeper reasoning it previously built. It gets lazy choosing a quick pattern over deep logic just to minimize errors.
Yes. And this creates a major almost entirely invisible problem for AI developers. The training loss metric can keep improving steadily on their dashboard. Right. It visually looks like the model is getting smarter every single day. But the model's actual logical reasoning behavior quietly gets worse in the background. The final checkpoint of a model may not always be the best checkpoint. This raises a profoundly important question about how we currently evaluate these systems. We might be widely deploying internally compromised models without even realizing it. We absolutely might be. We are going to take a quick second for a word from our sponsor. Sounds good. Okay. Let's synthesize this entire journey and look at the actual big picture. We are placing immense almost blind trust in these AI systems today. We really are. We use advanced cursor coding tools powered by GLM 5.3 for our daily software. We literally have the Oval Office making geopolitical moves based on probabilistic chatbot outputs. Yeah. It is everywhere.
Yet the foundational architecture of these models remains remarkably unstable underneath it all. It can mowd hop into the total failure instantly. And it does this just by reading a few billion more words. The overarching takeaway from all these reports is quite clear and honestly a little daunting. The sheer raw capability of AI is scaling incredibly fast right now. But it is scaling much faster than our fundamental understanding of it. We absolutely do not fully grasp its inner workings yet. Think about this as you go about your day. If the final checkpoint of an AI isn't always the smartest one, what happens next? That is the real question. What happens if a government or defense agency accidentally deploys a model right as it hits a mode hopping slump? It is kind of terrifying to think about. Are we absolutely sure that the model giving the advice today is as smart as the one from yesterday? Thank you for taking this deep dive with us. Keep questioning the invisible systems around you.
Stay curious. We will see you next time. 5 ounces selected varieties are 299 each with digital coupon limit 3. Hurry in! Visit FonsorAlbertsons.com for more deals and ways to save.
More episodes
More from AI Fire Daily

#644 Neil: 5 AI Agent Skills I Tested For Research, Code And Design
AI Fire Daily

#643 Neil: Gemini 4 Argon Gets A Real Test Beyond The Benchmarks
AI Fire Daily

#642 Neil: GPT-6.1 Sol In Codex Surprised Me Across Three Real Tests
AI Fire Daily

π EP 367: Google Launches Gemini 4 Argon & Dyna Unveils Taku Humanoid
AI Fire Daily