Research
We specify what we want from a language model through examples, while the properties themselves must hold far beyond them. I am interested in AI alignment that operates on a model's internal representations, shaping abstract dispositions rather than enumerating desired outputs, so that behavior follows from what a model is rather than from what it has been told to do. This motivates a parallel interest in AI introspection: what models encode about their own state, their plans, and the process of producing their own outputs, and whether those internal signals can be read out faithfully enough to support oversight. The same question underlies honesty and deception, where what matters is whether a model's stated reasoning reflects the computation that actually produced it, and whether its representations can expose a divergence that the output alone would hide. I come from a cybersecurity background: my PhD thesis was on adversarial attacks and defenses for AI-based network intrusion detection systems.
Papers
Efficient Safety Alignment of Language Models via Latent Personality Traits
Mohamed Amine Merzouk, Nolan Smyth, Damiano Fornasiere, Linh Le, David Williams-King, Adam Oberman
Conference on Language Modeling (COLM 2026)
How Much is Left? LLMs Linearly Encode Their Remaining Output Length
Mohamed Amine Merzouk, Dmitri Carpov, Mirko Bronzi, Damiano Fornasiere, Adam Oberman
Preprint (2026)
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
Linh Le, David Williams-King, Mohamed Amine Merzouk, Aton Kamanda, Adam Oberman
Trustworthy AI Workshop at the International Conference on Learning Representations (ICLR 2026)
Deep-Cover Agents: Long-Horizon Prompt Injections on Production LLM Systems
Neel Alex, Mohamed Amine Merzouk, David Krueger
Preprint (2026)
Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
Mina Taraghi, Yann Pequignot, Amin Nikanjam, Mohamed Amine Merzouk, Foutse Khomh
Preprint (2025)
Diffusion-based Adversarial Purification for Intrusion Detection
Mohamed Amine Merzouk, Erwan Beurier, Reda Yaich, Nora Boulahia-Cuppens, Frédéric Cuppens, Foutse Khomh
Conference on Data and Applications Security and Privacy (DBSec 2025)
Adversarial robustness of deep reinforcement learning-based intrusion detection
Mohamed Amine Merzouk, Christopher Neal, Joséphine Delas, Reda Yaich, Nora Boulahia-Cuppens, Frédéric Cuppens
International Journal of Information Security (2024)
Parameterizing poisoning attacks in federated learning-based intrusion detection
Mohamed Amine Merzouk, Frédéric Cuppens, Nora Boulahia-Cuppens, Reda Yaich
International Conference on Availability, Reliability and Security (ARES 2023)
Investigating the practicality of adversarial evasion attacks on network intrusion detection
Mohamed Amine Merzouk, Frédéric Cuppens, Nora Boulahia-Cuppens, Reda Yaich
Annals of Telecommunications (2022)
Evading deep reinforcement learning-based network intrusion detection with adversarial attacks
Mohamed Amine Merzouk, Joséphine Delas, Christopher Neal, Frédéric Cuppens, Nora Boulahia-Cuppens, Reda Yaich
International Conference on Availability, Reliability and Security (ARES 2022)
A deeper analysis of adversarial examples in intrusion detection
Mohamed Amine Merzouk, Frédéric Cuppens, Nora Boulahia-Cuppens, Reda Yaich
International Conference on Risks and Security of Internet and Systems (CRiSIS 2020)
Teaching
INF8102: Security of cloud infrastructures
Lecturer: Fall 2025
INF8085: Cybersecurity
Lecturer: Summer 2025, Winter 2026, Summer 2026
CR345: Security of servers
Lecturer: Winter 2025
INF4420A: Computer security
Lecturer: Summer 2023, Summer 2024
Lab instructor: Winter 2022, Fall 2022, Summer 2023, Summer 2024
INF6103: Cybersecurity of critical infrastructures
Lab instructor: Winter 2022, Fall 2022, Fall 2023, Fall 2024
INF8602: Cybersecurity of operating systems
Lab instructor: Winter 2022
INF1040: Introduction to Computer Engineering
Evaluator: Winter 2021, Fall 2021