Hybrid Defense Against Indirect Prompt Injection in RAG: A Cross-Model Benchmark

Authors

  • Dr.Geetha Raghuraj
  • Dr.Veena N
  • Dr. Jennie Bharathic
  • Geetha Rani Kadavad

Keywords:

prompt injection, retrieval-augmented generation, large language models, LLM security, adversarial robustness, hybrid defense.

Abstract

Retrieval-Augmented Generation (RAG) systems ground large language model (LLM) outputs in externally retrieved document. However, this dependence paves the way for indirect prompt injection, in which malicious instructions embedded in received documents take control of model activity. Existing lightweight defenses such as content sanitization and instruction-hierarchy prompting each encode a different, partially independent signal for differentiating trusted and non-trusted instructions. We hypothesize that combining both signals yields performance on par with either signal individually, as a model evading detection via one cue may be captured by the other. This study suggests a hybrid defense that combines sanitization delimiters with instruction-hierarchy prompting, and benchmarks it against three baseline defenses (none, sanitize-only, instruction-hierarchy-only) and a second-pass output-verification defense. Using a RAG pipeline constructed on a 60-question subset with 35% of the retrieval corpus poisoned, we assess a taxonomy of indirect prompt injection attacks on three open-source instruction-tuned LLMs.  We quantify both Task accuracy and Attack Success Rate(ASR) to illustrate the security–utility trade-off of each defense. Critically, the hybrid defense never showed any severe accuracy drop in any model, making it a consistently safe, deployment-ready choice. These findings offer practical, deployment-ready guidance for developers seasoning RAG systems against indirect prompt injection without requiring model retraining.

Downloads

Published

2026-09-24

How to Cite

Raghuraj , D., N , D., Bharathic, D. J., & Kadavad , G. R. (2026). Hybrid Defense Against Indirect Prompt Injection in RAG: A Cross-Model Benchmark. International Journal of Artificial Intelligence and Machine Learning, 6(3), 1148–1156. Retrieved from https://svedbergopen.com/index.php/ijaiml/article/view/2557