Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming
2026-08-05 • Cryptography and Security
Cryptography and Security
AI summaryⓘ
The authors address the security problem of prompt injection attacks on large language models (LLMs). They created PIMiner, a system that learns attack strategies by training on different datasets and models, then uses this knowledge to test new LLMs with very few queries. Unlike previous methods, PIMiner can transfer its strategies to unseen models without retraining. Their experiments show that PIMiner effectively finds vulnerabilities in multiple advanced LLMs.
prompt injectionlarge language modelsred-teamingreinforcement learningtransfer learningattack success ratetarget modelssecurity evaluation
Authors
Yanting Wang, Chenlong Yin, Runpeng Geng, Jinyuan Jia
Abstract
Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing state-of-the-art prompt injection red-teaming methods primarily rely on reinforcement learning (RL), producing attacker models that often generalize poorly to new target LLMs. In this work, we develop PIMiner, an agentic system for prompt injection red-teaming. During training, PIMiner is trained on a sequence of (dataset, target model) pairs and builds a strategy library from scratch. At test time, the learned strategy library can be directly transferred to a previously unseen target LLM without additional training. PIMiner requires only a small number of queries to a target agent (e.g., 10) per test sample. Experimental results demonstrate that PIMiner achieves strong performance. On IPIArena, it attains a 76.2% ASR against Gemini-2.5-Pro, 61.9% ASR against GPT-5.1, and 42.9% ASR against Claude-Sonnet-4.5. On AgentDojo, it achieves an 86.7% ASR against Gemini-2.5-Pro, 53.3% ASR against GPT-5.1, and 40.0% ASR against Claude-Sonnet-4.5.