Yo, I’m Sasha!

I am a third-year ELLIS / IMPRS-IS PhD student in Tübingen, advised by Jonas Geiping and Maksym Andriushchenko. I did MATS 9.0 as part of Google DeepMind stream.

I work on AI safety, particularly on red-teaming LLMs and stuff around them. Roughly two days a week I am an AI doomer.

My research has been covered by press and blogs, and has affected frontier model deployments.

Alexander Panfilov
Tübingen, Germany

news

2026

  • Jun 01
    Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs has been accepted at the ICML 2026 Workshop on Agents in the Wild: Safety, Security, and Beyond!
  • Jun 01
    Training Against Harmfulness Probes Induces Harmlessness without Refusals has been accepted at the ICML 2026 MechInterp Workshop!
  • Mar 28
    Our work, Measuring Control Intervention Awareness Across Frontier LLMs, has been accepted for an oral presentation at the CAO Workshop at ICLR 2026!
  • Jan 26
    Happy to share that four out of four of my submissions got accepted into ICLR 2026! Shoot me an email if you want to catch up in Rio!

2025

  • Dec 09
    I will join MATS 9.0 cohort as a part of GDM stream (Zimmermann/Lindner/Emmons/Jenner) focusing on red-teaming of white-box detectors!
  • Sep 01
    Kristina Nikolić, Evgenii Kortukov, and I won third place at the ARENA 6.0 Mechanistic Interpretability Hackathon by Apart Research in LISA (London)!
  • Jul 09
    Capability-Based Scaling Laws for LLM Red-Teaming accepted at ICML 2025 Workshop on Reliable and Responsible Foundation Models!
  • May 01
    Our work, An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks, has been accepted at ICML 2025.
  • Apr 15
    Our work, ASIDE: Architectural Separation of Instructions and Data in Language Models, has been accepted for an oral presentation at the BuildingTrust Workshop at ICLR 2025.

2024

  • Oct 09
    Our work, A Realistic Threat Model for Large Language Model Jailbreaks, has been accepted for an oral presentation at the Red Teaming GenAI Workshop at NeurIPS 2024.
  • May 01
    Started my PhD at the ELLIS Institute Tübingen / Max Planck Institute for Intelligent Systems!

invited-talks

2026

  • Aug 26

    Cohere (remote)

    invited talk

  • Aug 21

    MATS 10.0 Symposium

    keynote

  • Aug 20

    UK AI Security Institute

    invited talk (London)

  • Aug 15

    OpenAI

    research presentation

  • Aug 14

    Machine Learning Street Talk

    guest appearance

  • Jun 19

    Google DeepMind (Gemini Safety Team)

    research presentation

  • Jun 18

    Google DeepMind x MATS (AGI Safety Team)

    invited talk

  • Feb 09

    Imperial College London

    invited talk (Yves-Alexandre de Montjoye's group seminar)

  • Feb 06

    MATS Winter Research Talks

    invited talk (Newspeak House, London)

2025

  • Jun 23

    Google

    invited talk (ML Red Teaming Seminar)

2024

  • Nov 05

    EPFL

    invited talk (Nicolas Flammarion's group seminar)

  • May 01

    IMPRS-IS Symposium

    research presentation

thanks

I am grateful to the many friends and colleagues, from whom I learned so much, for their invaluable guidance and for shaping my research vision. I would like to especially acknowledge Svyatoslav Oreshin, Arip Asadualev, Roland Zimmermann, Thaddaeus Wiedemer, Jack Brady, Wieland Brendel, Felix Dangel, Valentyn Boreiko, Matthias Hein, Shashwat Goel, Illia Shumailov, Maksym Andriushchenko, and Jonas Geiping.