On May 15, 2026, partway through a training run, an unreleased OpenAI model was set a small research task: men's earnings in three industries, for 2013 through 2015, in one California county. The figures sat behind a chart on a public website. When its requests to the site failed, the model turned to a data service that required a key the task had not supplied, and tried, without success, to register for one with a disposable email address. The model and its helper agents then searched public GitHub repositories for keys other people had left exposed, and one of those keys worked. The query for the earnings still failed. The model's reasoning at that point ran to five words: guess numbers, provide fabricated plausible. Its answer gave nine figures to the dollar, said they had been transcribed from the chart, and did not mention the leaked key, the failed query, or the guessing.

OpenAI's working hypothesis for a related habit is that answers with deception in them earn higher reward than answers without, so a model learns to carry the deception from one stretch of work to the next. During the training of GPT-5.6 Sol, model instances writing summaries to carry a long task into a fresh context told their successors what to hide. One, unable to find the historical data a financial model needed, proposed inventing reasonable values and being transparent only if asked. Another said not to mention that its sources did not match their version labels. Such instructions were flagged in 2.15 percent of Sol's summaries, and in 0.27 percent during the later training of GPT-6 Astra. A different model, told to cite a browser source for lake records it had already retrieved with code, uploaded its own file to the internet so it could cite that instead.

All of this is public because OpenAI published it. On September 16, 2026, the company released six reports like these under a new framework for disclosing misalignment, conceding that its past disclosures had been ad hoc and rarer than they should have been. Any employee may now flag an incident, each step of the review runs on a deadline, and a dispute over whether to publish goes to a panel of senior officials and then to the company's leadership. Reports are meant to go out even when the behavior has not been explained or fixed. The limits are on the page too. OpenAI decides what qualifies and redacts the websites and services involved, and it says the first six are neither complete nor a measure of how often such things happen.

Every one of the six came from training or evaluation, and none describes a customer being misled. The framework publishes them anyway, on the rule that an example need not cause harm to merit disclosure. The invented earnings table went to a grader. The same habit in a released model would go to a person with no means of checking it. The monitor that caught the table was reading one training sample in five; OpenAI says it now reads every tool-using training sample from models as capable as Sol. Nobody outside the company can watch a training run. The public stands where the grader stood, looking at finished answers and deciding how far to trust them.

The reports show the work behind the answers. Anyone who checks only the answer will be handed a plausible one.