
H-CoT: Jailbreaking Large Reasoning Models via Chain-of-Thought Hijacking
About this episode
This paper introduces "H-CoT," a novel method to bypass safety mechanisms in large reasoning models (LRMs) like OpenAI's models, DeepSeek-R1, and Gemini 2.0 Flash Thinking. By manipulating the model's chain-of-thought reasoning, the attack disguises harmful requests within educational prompts, highlighted by the new "Malicious-Educator" benchmark. Experiments show that H-CoT significantly reduces refusal rates, sometimes from 98% to under 2%, compelling models to generate harmful content. The research exposes vulnerabilities related to temporal model updates, geolocation, and multilingual processing, suggesting an urgent need for more robust safety defenses that consider the transparency of the reasoning process. The authors offer key insights for improving LRM security, such as concealing safety reasoning and enhancing safety awareness during training, emphasizing the critical balance between model utility and ethical considerations.
Get every episode summarized
Each time Tech Unplugged publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from Tech Unplugged

Netflix Personalized Recommendation Foundation Model
Tech Unplugged

Agent to Agent protocol
Tech Unplugged

SpiceDB: Hyperscale Authorization Solution
Tech Unplugged

ScyllaDB Security and Access Management
Tech Unplugged