Securing Agentic AI: Threat Modeling and Fragility Analysis Across Layers

Large language models are increasingly deployed as autonomous agents that use tools, call APIs, and cooperate with other agents. This autonomy sharply enlarges the attack surface: beyond classical prompt injection, an agent can be manipulated through its execution environment. Numerical quirks of the inference engine, the serving hardware, or malicious peer agents can influence an agent's output - none of which today's input-level guardrails detect.

This project treats agentic-AI security as a classical systems-security problem. You will help build a structured threat taxonomy that maps adversarial capabilities across system layers (input, model, inference stack, multi-agent) and turn it into a reproducible evaluation harness by extending open-source tooling such as AgentDojo.

Relevance

As agentic AI enters software engineering, finance, and critical infrastructure, understanding where and how badly such systems break is a prerequisite for trustworthy deployment and for emerging regulation such as the EU AI Act.

Questions you could contribute to
  • Which perturbations at the inference or hardware layer measurably change an agent's decisions, and how can we quantify this "fragility"?
  • Can a single benchmark expose cross-layer attacks (e.g., an inter-agent message that triggers an unsafe tool call) that per-layer tests miss?
Final product for evaluation

A working extension of an agentic-security benchmark implementing at least one cross-layer attack scenario, together with a short report presenting reproducible measurements of agent fragility. Well-documented, runnable code and a clear experimental write-up weigh more heavily than breadth of coverage.

About the lab

The Institute of IT Security at the University of Lübeck (PI Prof. Thomas Eisenbarth) studies the security of computing systems end to end, from microarchitecture and trusted execution environments up to machine-learning systems. The group pairs a strong systems-security track record (e.g., attacks on commercial TEEs and software side-channel defenses such as Cipherfix) with security-for-AI research. Recent examples include platform-triggered LLM backdoors (FloatDoor) and prompt-stealing attacks on generative models (Prompt Pirates). This dual perspective lets us analyze agentic AI as a complete system rather than as an isolated model.

Own and related readings

  • Loose, Sander, Mächtle, Eisenbarth. FloatDoor: Platform-Triggered Backdoors in LLMs. arXiv:2606.19535, 2026.
  • Mächtle et al. Prompt Pirates Need a Map: Stealing Seeds Helps Stealing Prompts. ACM ASIA CCS 2026.
  • Debenedetti et al. AgentDojo: A Dynamic Environment to Evaluate Prompt-Injection Attacks and Defenses for LLM Agents. NeurIPS 2024.
  • Greshake et al. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. AISec 2023.

Contact

Anja Rabich
Studiengangskoordinatorin IT-Sicherheit
studium.itsec@uni-luebeck.de