Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals

This paper compares five confidence methods across four Qwen and Gemma activation oracles. Forced choice is most accurate when possible answers are known. Bootstrap agreement is calibrated for free text without annotated data.

August 2026 · F. Torrielli, P. Schneider-Kamp, L. G. Poech

Developing Virtual Personas from User Level Social Media Data

This paper tests the stability and human alignment of virtual patient profiles built from anonymized social media histories. Results vary across models, prompts, and clinical domains.

July 2026 · F. Quilghini, F. Torrielli, A. Rapp, L. Di Caro, M. Settanni, D. Marengo

Exploiting large language models in peer review: indirect prompt injection attacks and integrity probes

This paper studies indirect prompt injection in AI assisted peer review across 42,000 chatbot outputs. Hidden instructions alter assessments and can support organizer integrity tests.

July 2026 · F. Torrielli, S. Locci, A. Rapp, L. Di Caro

Potential and limitations of LLMs for augmenting lexical knowledge bases

This paper tests whether large language models can extend lexical knowledge bases. Human evaluators accepted 86.7% of novel concepts. Automatic overlap metrics missed many valid additions.

July 2026 · F. Torrielli, G. Siragusa, V. Lovera Rulfi, A. Rapp, L. Di Caro

PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models

This paper introduces PsychoSafe, a refusal framework grounded in evidence-based intervention strategies, and evaluates prompting and fine-tuning approaches for psychologically informed LLM refusals.

June 2026 · G. Barmina, F. Torrielli, S. Harms, J. Nielsen, F. Mächtle, S. L. Beltoft, P. Schneider-Kamp, T. Eisenbarth, L. G. Poech, A. Lauscher

Prompt Engineering and Prompt Thinking

A 20-hour course on prompt engineering methodologies and critical thinking about prompts, taught at the Master in Ethics and Artificial Intelligence, University of Torino.

October 2025 · Federico Torrielli

Paint it, BLACK: A Novel Methodology for Prompting

This paper introduces BLACK, a novel methodology for prompting large language models and generative AI systems.

September 2023 · F. Torrielli