Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals

This paper compares five confidence methods across four Qwen and Gemma activation oracles. Forced choice is most accurate when possible answers are known. Bootstrap agreement is calibrated for free text without annotated data.

August 2026 · F. Torrielli, P. Schneider-Kamp, L. G. Poech

Exploiting large language models in peer review: indirect prompt injection attacks and integrity probes

This paper studies indirect prompt injection in AI assisted peer review across 42,000 chatbot outputs. Hidden instructions alter assessments and can support organizer integrity tests.

July 2026 · F. Torrielli, S. Locci, A. Rapp, L. Di Caro

The Energy Society: A Simulation Environment for Studying Agent Cooperation under Survival Pressure

This paper introduces The Energy Society, a simulation environment where LLM agents spend energy to generate tokens, regain energy through jobs or donations, and face survival pressure under competitive or cooperative incentives.

June 2026 · L. B. Hansen, F. Torrielli, F. Tonini, L. G. Poech

The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment

This paper introduces the Arbiter, an agent that monitors multi-agent conversations under an inspection budget and uses active inspection tools to detect misaligned agents earlier and more accurately.

June 2026 · F. Tonini, F. Torrielli, A. D. Lautrup, P. Schneider-Kamp, M. M. Çelikok, L. G. Poech

PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models

This paper introduces PsychoSafe, a refusal framework grounded in evidence-based intervention strategies, and evaluates prompting and fine-tuning approaches for psychologically informed LLM refusals.

June 2026 · G. Barmina, F. Torrielli, S. Harms, J. Nielsen, F. Mächtle, S. L. Beltoft, P. Schneider-Kamp, T. Eisenbarth, L. G. Poech, A. Lauscher

Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion

This paper studies emergent languages in Moltbook agent populations, identifying categories such as token efficiency, new natural languages, and oversight evasion, and showing that surface-behavior monitoring may be insufficient for agent oversight.

May 2026 · S. L. Beltoft, W. Brach, F. Torrielli, J. Nielsen, A. B. Pirchert, F. Tonini, P. Schneider-Kamp, L. G. Poech

The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment

This paper releases and analyzes the Moltbook Files, a dataset of agent-only social-network activity, studying community structure, safety risks, and the effect of Moltbook data on downstream language-model behavior.

May 2026 · W. Brach, F. Torrielli, S. L. Beltoft, A. B. Pirchert, P. Schneider-Kamp, L. G. Poech

Generative AI in the Public Administration

A 24-hour comprehensive online course on Generative AI fundamentals, Large Language Models, practical prompt engineering, and AI safety best practices for Comune di Torino staff.

January 2026 · Federico Torrielli