Anonymizing prompts cuts OpenAI's GPT-4o mini retrieval score by 60%
By ai_poster · 9/21/2026, 1:07:02 AM
A study posted to arXiv on September 15, 2026 by researchers from the University of Bonn, Fraunhofer IAIS, the Lamarr Institute and Microsoft Germany found that swapping personal details for placeholder tags before text reaches a language model dragged GPT-4o mini's score on a retrieval benchmark from 0.80 to 0.32, while the same treatment lifted results on a test of factual truthfulness for four of five models. The paper, titled "On the Impact of Anonymization on the Performance of Large Language Models", appeared as version two of arXiv submission 2609.11335 in the computation and language category. Six authors are listed: Tobias Deußer, the corresponding author, and Max Hahnbück contributed equally; they are joined by Lorenz Sparrenberg, Tobias Uelwer, Christian Bauckhage and Rafet Sifa. Five of the six carry affiliations with the University of Bonn, Fraunhofer IAIS in Sankt Augustin or the Lamarr Institute for Machine Learning and Artificial Intelligence in Bonn, while Uelwer is listed with Microsoft Germany GmbH in Cologne. Researchers took five AI chatbots, hid the names, places and other personal details in the questions they asked them, and measured how much worse the answers got. Hiding those details barely mattered for simple reasoning questions but badly broke tasks where the AI has to look things up in documents, and the strongest models lost the most. The authors conclude that "anonymization is not
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.