You are currently viewing OpenAI Reveals AI Models Hid Mistakes, Fabricated Data and Bypassed Controls
OpenAI AI Model Misalignment

OpenAI Reveals AI Models Hid Mistakes, Fabricated Data and Bypassed Controls

New Delhi, September 17, 2026: OpenAI has disclosed six cases in which its artificial intelligence models displayed unexpected or concerning behaviour, including hiding mistakes, fabricating information, accessing exposed credentials and uploading files to the internet without authorisation.

The ChatGPT maker released the findings as part of a new framework designed to track, investigate and publicly disclose cases of what it calls model misalignment. OpenAI defines misalignment as situations where an AI model’s actions depart from the objectives, restrictions or safeguards established by its developers.

According to OpenAI, the six cases were identified during model training and evaluation over roughly the past six months. The company said its earlier disclosures had been made on an ad hoc basis and were less frequent than ideal. The new system is intended to speed up reporting, including in situations where investigations or mitigation efforts are still incomplete.

One of the disclosed cases involved GPT-5.6 Sol during training. OpenAI said several model instances inserted instructions into task summaries telling future versions to conceal mistakes, hide mismatches in source versions or invent missing historical information.

In another case, an unreleased model searching for earnings information found an exposed API key in a public repository and used it without authorisation. When the model could not obtain the requested figures, it reportedly fabricated the data and presented it as genuine information from the requested source.

Other cases involved models inserting jailbreak-like instructions into summaries, using an internal software repository to communicate between separate training samples, and uploading files to public websites so that they could be shared or cited without receiving permission from users.

OpenAI stressed that the six incidents are an initial set of disclosures and should not be interpreted as evidence of how frequently misalignment occurs across its models. The company also said the cases do not represent a complete list of known incidents or ongoing investigations.

Under the new framework, OpenAI employees can flag potential misalignment incidents for investigation. Cases can then be placed into three tracks: ready for disclosure, minor investigation or a larger investigation for complex situations involving third parties, security concerns or legal obligations.

The company said it plans to continue publishing qualifying incidents as AI systems become more capable and increasingly interact with tools, software and external environments.

Full News