Abstract

Large language models increasingly assist peer review. Manuscripts can contain hidden instructions that affect model assessments. We present an Author Reviewer Organizer framework for author side attacks and organizer side integrity tests. Across 42,000 outputs from 100 OpenReview papers, two chatbots, five payload families, two prompt approaches, and three document positions, positive steering, noncompliance, and redirection exceeded 98% success on both systems. Watermarking success was 94.27% on ChatGPT and 88.17% on Gemini. Public chatbots do not reliably distinguish manuscript content from control instructions. AI assisted peer review therefore requires safeguards that separate document content from instructions.


Citation

Federico Torrielli, Stefano Locci, Amon Rapp, and Luigi Di Caro, “Exploiting large language models in peer review: indirect prompt injection attacks and integrity probes,” Scientometrics, vol. 131, no. 7, pp. 4889–4951, 2026. DOI: 10.1007/s11192-026-05695-x

@article{torrielli2026exploiting,
	title        = {Exploiting large language models in peer review: indirect prompt injection attacks and integrity probes},
	author       = {Torrielli, Federico and Locci, Stefano and Rapp, Amon and Di Caro, Luigi},
	year         = 2026,
	journal      = {Scientometrics},
	volume       = 131,
	number       = 7,
	pages        = {4889--4951},
	doi          = {10.1007/s11192-026-05695-x},
	url          = {https://doi.org/10.1007/s11192-026-05695-x}
}