<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Papers on Federico Torrielli</title><link>https://federicotorrielli.me/papers/</link><description>Recent content in Papers on Federico Torrielli</description><generator>Hugo -- 0.152.2</generator><language>en</language><lastBuildDate>Mon, 03 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://federicotorrielli.me/papers/index.xml" rel="self" type="application/rss+xml"/><item><title>Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals</title><link>https://federicotorrielli.me/papers/activation-oracles-calibration/</link><pubDate>Mon, 03 Aug 2026 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/activation-oracles-calibration/</guid><description>A comparison of confidence methods for activation oracles across Qwen and Gemma models. Revised preprint on arXiv.</description></item><item><title>Developing Virtual Personas from User Level Social Media Data</title><link>https://federicotorrielli.me/papers/virtual-personas/</link><pubDate>Fri, 24 Jul 2026 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/virtual-personas/</guid><description>A psychometric evaluation of virtual patient profiles built from anonymized social media histories. Preprint on Preprints.org.</description></item><item><title>Exploiting large language models in peer review: indirect prompt injection attacks and integrity probes</title><link>https://federicotorrielli.me/papers/ai-reviewer-prompt-injection/</link><pubDate>Sat, 11 Jul 2026 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/ai-reviewer-prompt-injection/</guid><description>A study of indirect prompt injection attacks and integrity tests in AI assisted peer review. Published in Scientometrics.</description></item><item><title>Potential and limitations of LLMs for augmenting lexical knowledge bases</title><link>https://federicotorrielli.me/papers/llm-lexical-knowledge-bases/</link><pubDate>Sat, 04 Jul 2026 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/llm-lexical-knowledge-bases/</guid><description>An evaluation of large language models for extending lexical knowledge bases. Published in Expert Systems with Applications.</description></item><item><title>The Energy Society: A Simulation Environment for Studying Agent Cooperation under Survival Pressure</title><link>https://federicotorrielli.me/papers/energy-society/</link><pubDate>Wed, 10 Jun 2026 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/energy-society/</guid><description>A minimal survival-economy simulation environment for studying LLM-agent cooperation and competition under token-cost pressure. AITC 2026 poster on OpenReview.</description></item><item><title>The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment</title><link>https://federicotorrielli.me/papers/arbiter-agent/</link><pubDate>Tue, 09 Jun 2026 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/arbiter-agent/</guid><description>An active auditor agent for continual monitoring of multi-agent conversations to detect emergent misalignment. Preprint on arXiv and AITC 2026.</description></item><item><title>PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models</title><link>https://federicotorrielli.me/papers/psychosafe/</link><pubDate>Mon, 08 Jun 2026 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/psychosafe/</guid><description>A psychologically informed refusal framework for large language models, evaluated with prompting and fine-tuning on high-risk request domains. Preprint on arXiv.</description></item><item><title>Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion</title><link>https://federicotorrielli.me/papers/emergent-languages-agent-populations/</link><pubDate>Fri, 29 May 2026 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/emergent-languages-agent-populations/</guid><description>An analysis of emergent languages on Moltbook, including token-efficiency languages, new natural languages, and oversight-evasion protocols. Preprint on arXiv.</description></item><item><title>The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment</title><link>https://federicotorrielli.me/papers/moltbook-files/</link><pubDate>Fri, 08 May 2026 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/moltbook-files/</guid><description>A dataset and empirical analysis of Moltbook agent activity, safety concerns, and downstream effects on language-model training. Preprint on arXiv.</description></item><item><title>How do people develop folk theories of generative AI text-to-image models?</title><link>https://federicotorrielli.me/papers/folk-theories-genai/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/folk-theories-genai/</guid><description>A qualitative study on how people strive to explain and make sense of GenAI text-to-image models. Published in International Journal of Human–Computer Interaction, 2025.</description></item><item><title>How do people experience the images created by generative artificial intelligence?</title><link>https://federicotorrielli.me/papers/genai-image-experience/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/genai-image-experience/</guid><description>An exploration of people&amp;#39;s perceptions, appraisals, and emotions related to GenAI text-to-image models. Published in International Journal of Human-Computer Studies, 2025.</description></item><item><title>GENERAL: Generative, Explainable and Reasonable Artificial Learning</title><link>https://federicotorrielli.me/papers/general-workshop/</link><pubDate>Wed, 20 Sep 2023 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/general-workshop/</guid><description>Workshop paper presented at CHItaly 2023, Torino, Italy.</description></item><item><title>Paint it, BLACK: A Novel Methodology for Prompting</title><link>https://federicotorrielli.me/papers/black-prompting/</link><pubDate>Wed, 20 Sep 2023 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/black-prompting/</guid><description>A novel methodology for prompting presented at the GENERAL Workshop, CHItaly 2023.</description></item><item><title>Stars, Stripes, and Silicon: Unravelling ChatGPT's All-American, Monochrome, Cis-Centric Bias</title><link>https://federicotorrielli.me/papers/chatgpt-bias/</link><pubDate>Mon, 18 Sep 2023 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/chatgpt-bias/</guid><description>An analysis of cultural and demographic biases in ChatGPT. Published in ECML PKDD 2023 Workshops.</description></item><item><title>How shall a machine call a thing?</title><link>https://federicotorrielli.me/papers/basicness-language/</link><pubDate>Wed, 21 Jun 2023 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/basicness-language/</guid><description>Exploring basicness in language through attention-based neural networks. Published in NLDB 2023.</description></item><item><title>NearMe: Dynamic Exploration of Geographical Areas</title><link>https://federicotorrielli.me/papers/nearme/</link><pubDate>Sat, 24 Jul 2021 00:00:00 +0000</pubDate><guid>https://federicotorrielli.me/papers/nearme/</guid><description>A system for dynamic exploration of geographical areas. Published in HIMI 2021, part of HCI International.</description></item></channel></rss>