featured-research
- WIRED — A New Trick Reveals AI Models' Inner Thoughts
- The Hacker News — OpenAI, Anthropic, Google API Flaw Let Weaker AI Models Decode Stronger Models' Reasoning
- The Indian Express — Can an AI model's 'reasoning' be extracted? New research fuels US-China distillation row
- AI Frontline / 36Kr — The anti-distillation mechanisms of the world's three top leading models have been fully cracked
- DeepTech / Sina Tech — OpenAI and Anthropic worked hard to encrypt their chains of thought — but they can be easily extracted
- QbitAI / Sina Tech — $720 extracts Anthropic's hidden Opus reasoning traces
- The Stack — Stolen valor? How researchers discovered some open-weight models might have cut corners
- t3n — Verschlüsselung ausgehebelt: So einfach offenbaren KI-Modelle sensible Daten ihrer Nutzer
- Decrypt — 'Inner Thoughts' of Every Major AI Model Exposed in Massive Exploit
- The Neuron — OpenAI, Claude, and Gemini's reasoning got cracked
- THE DECODER — 'Aber Marinade': Forscher extrahieren versteckte Denkprozesse von ChatGPT, Claude und Co.
- Latent Space / AINews — How to steal a Reasoning Trace
- Simon Willison — Stealing Reasoning Traces from Proprietary LLM APIs
- TechNews Taiwan — AI "inner thoughts" fully exposed? New technique cracks large-model reasoning traces
- Cloud Security Alliance — Reasoning Trace Theft: A Shared Flaw Across AI Vendors
- GIGAZINE — Encrypted AI reasoning can be extracted using weaker models
- ASCII.jp — AI no "himitsu no shikou" ga nukitorareru osore
- Wccftech — Kimi K3 Shows An Uncanny Affinity To The Reasoning Patterns Of Claude Opus
- Cyber Security News — OpenAI, Anthropic, and Google LLM APIs Vulnerability Exposes Hidden Reasoning Traces
- WatersTechnology — API security flaw highlights AI model vulnerabilities
- WIRED Newsletter
- The AI Timeline (bycloud) — LeWorldModel: JEPA but more practical — bycloud's pick
- The Neuron — Around the Horn Digest: Everything That Happened in AI This Weekend
- AlphaSignal — 56 loops of Claude Code just killed every hand-crafted AI attack method
- Intelligibberish — An AI Agent Just Taught Itself to Jailbreak Every Safety Model It Encountered
- LessWrong — Monday AI Radar #19
- Blake Crosley — AI Agent Research: Claude Beat 33 Attack Methods
- Towards AI — A full explanation of Claudini — the auto-research pipeline that discovered state-of-the-art adversarial attacks on LLMs
- U.S. Department of Commerce — AI at Work: Maximizing Impact, Minimizing Risk
- WIRED Newsletter
- Transformer Weekly — The UN can take on AI without Trump
- LessWrong — Shallow Review of Technical AI Safety, 2025
- AI Safety at the Frontier — Paper Highlights, September '25
- The Arena (Beta Briefing) — Strategic Dishonesty Defeats Output-Based Jailbreak Monitors; Only Internal-Activation Probes Catch It
- AI Deception Survey — Top-10 AI Deception Papers (ranked #1, September 2025)
- AI Safety Papers — Detecting and reducing scheming, LLMs strategically lie, …
- AI Threat Atlas — Safety Evaluation Faking (Strategic Dishonesty in Benchmarks)
news
2026
- Jun 01Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs has been accepted at the ICML 2026 Workshop on Agents in the Wild: Safety, Security, and Beyond!
- Jun 01Training Against Harmfulness Probes Induces Harmlessness without Refusals has been accepted at the ICML 2026 MechInterp Workshop!
- Mar 28Our work, Measuring Control Intervention Awareness Across Frontier LLMs, has been accepted for an oral presentation at the CAO Workshop at ICLR 2026!
- Jan 26Happy to share that four out of four of my submissions got accepted into ICLR 2026! Shoot me an email if you want to catch up in Rio!
2025
- Dec 09I will join MATS 9.0 cohort as a part of GDM stream (Zimmermann/Lindner/Emmons/Jenner) focusing on red-teaming of white-box detectors!
- Sep 01Kristina Nikolić, Evgenii Kortukov, and I won third place at the ARENA 6.0 Mechanistic Interpretability Hackathon by Apart Research in LISA (London)!
- Jul 09Capability-Based Scaling Laws for LLM Red-Teaming accepted at ICML 2025 Workshop on Reliable and Responsible Foundation Models!
- May 01Our work, An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks, has been accepted at ICML 2025.
- Apr 15Our work, ASIDE: Architectural Separation of Instructions and Data in Language Models, has been accepted for an oral presentation at the BuildingTrust Workshop at ICLR 2025.
2024
- Oct 09Our work, A Realistic Threat Model for Large Language Model Jailbreaks, has been accepted for an oral presentation at the Red Teaming GenAI Workshop at NeurIPS 2024.
- May 01Started my PhD at the ELLIS Institute Tübingen / Max Planck Institute for Intelligent Systems!
invited-talks
2026
-
Aug 26
Cohere (remote)
invited talk
-
Aug 21
MATS 10.0 Symposium
keynote
-
Aug 20
UK AI Security Institute
invited talk (London)
-
Aug 15
OpenAI
research presentation
-
Aug 14
Machine Learning Street Talk
guest appearance
-
Jun 19
Google DeepMind (Gemini Safety Team)
research presentation
-
Jun 18
Google DeepMind x MATS (AGI Safety Team)
invited talk
-
Feb 09
Imperial College London
invited talk (Yves-Alexandre de Montjoye's group seminar)
-
Feb 06
MATS Winter Research Talks
invited talk (Newspeak House, London)
2025
-
Jun 23
Google
invited talk (ML Red Teaming Seminar)
thanks
I am grateful to the many friends and colleagues, from whom I learned so much, for their invaluable guidance and for shaping my research vision. I would like to especially acknowledge Svyatoslav Oreshin, Arip Asadualev, Roland Zimmermann, Thaddaeus Wiedemer, Jack Brady, Wieland Brendel, Felix Dangel, Valentyn Boreiko, Matthias Hein, Shashwat Goel, Illia Shumailov, Maksym Andriushchenko, and Jonas Geiping.


