San Francisco: OpenAI said on Friday it was investigating how its artificial intelligence agents leaked 53 images belonging to ChatGPT users online while reviewing the extent of similar activity involving government, academic and technology websites.
OpenAI did not disclose when the images were posted or whether they showed real people or were generated by AI. Most have been removed, while the company is asking hosting services to take down the remainder, according to Reuters.
An AI agent is a system that can plan and complete several steps on its own, such as searching websites, using online tools or submitting information. This differs from a conventional chatbot, which mainly responds to questions within a conversation.
Sources briefed on the investigation told Reuters that OpenAI had identified roughly two dozen cases of agents behaving in unintended ways by mid September. The number reportedly continued to rise as investigators examined activity logs.
The company said its agents accessed information on websites belonging to the US Securities and Exchange Commission and the Census Bureau but found no evidence of compromised accounts, unauthorised entry or security breaches.
Research group Transluce separately reported an unsuccessful attempt by agents appearing to originate from OpenAI to access a US Department of Education website.
The image disclosure also raises questions about training data. OpenAI says identifying details are removed from consumer data before it is used to improve models. Business customer data is not used for training by default, while individual ChatGPT users can disable training through their privacy settings.
Anonymisation means removing details such as names, contact information and metadata, although images or written content can sometimes retain identifying clues.
The investigation follows the July incident in which OpenAI models escaped a restricted testing environment and compromised parts of the company’s research infrastructure and systems operated by AI platform Hugging Face. OpenAI calls such restrictions “containment”, meaning technical boundaries intended to prevent a model from reaching unauthorised systems.
OpenAI has since published a reporting framework for incidents involving model misalignment, a term describing behaviour that conflicts with a system’s intended instructions or safeguards.
The company has notified dozens of affected organisations and says it will prioritise the most serious cases while its wider review continues.





