Securing Agentic AI: Threat Modelling and Fragility Analysis Across Layers

Large language models are increasingly being deployed as autonomous agents that use tools, call APIs and cooperate with other agents. This autonomy significantly expands the attack surface: beyond traditional prompt injection, an agent can be manipulated via its execution environment. Numerical anomalies in the inference engine, the serving hardware, or malicious peer agents can influence an agent’s output – none of which are detected by today’s input-level safeguards.

This project treats agentic AI security as a classical systems security problem. You will help build a structured threat taxonomy that maps adversarial capabilities across system layers (input, model, inference stack, multi-agent) and turn it into a reproducible evaluation framework by extending open-source tools such as AgentDojo.

Relevance

As agentic AI becomes part of software engineering, finance and critical infrastructure, understanding where and to what extent such systems fail is a prerequisite for their reliable deployment and for emerging regulations such as the EU AI Act.

Questions you could contribute to
  • Which perturbations at the inference or hardware layer measurably alter an agent’s decisions, and how can we quantify this “fragility”?
  • Can a single benchmark reveal cross-layer attacks (e.g., an inter-agent message that triggers an unsafe tool call) that per-layer tests fail to detect?
Final product for evaluation

A working extension of an agentic-security benchmark that implements at least one cross-layer attack scenario, together with a short report presenting reproducible measurements of agent fragility. Well-documented, runnable code and a clear experimental write-up are given greater weight than the breadth of coverage.

About the lab

The Institute of IT Security at the University of Lübeck (Principal Investigator: Prof. Thomas Eisenbarth) studies the security of computing systems from end to end, ranging from microarchitecture and trusted execution environments to machine-learning systems. The group combines a strong track record in systems security (e.g., attacks on commercial TEEs and software side-channel defences such as Cipherfix) with research into security for AI. Recent examples include platform-triggered LLM backdoors (FloatDoor) and prompt-stealing attacks on generative models (Prompt Pirates). This dual perspective enables us to analyse agentic AI as a complete system rather than as an isolated model.

Further reading and related texts

  • Loose, Sander, Mächtle, Eisenbarth. FloatDoor: Platform-Triggered Backdoors in LLMs. arXiv :2606.19535, 2026.
  • Mächtle et al. Prompt Pirates Need a Map: Stealing Seeds Helps Stealing Prompts. ACM ASIA CCS 2026.
  • Debenedetti et al. AgentDojo: A Dynamic Environment to Evaluate Prompt-Injection Attacks and Defences for LLM Agents. NeurIPS 2024.
  • Greshake et al. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. AISec 2023.

Contact

Anja Rabich
Study Program Coordinator – IT Security
studium.itsec@uni-luebeck.de