Errors hidden from users, fabricated data, and files transferred over the Internet without authorization. The relentless race of artificial intelligence is accompanied by a sequence of alarms, increasingly launched by the same companies that develop it. The latest comes from OpenAI, which on September 16th has published six reports su unexpected behaviors of their models, observed over the past six months during training and evaluation. The company also announced a new process to disclose these anomalies more quickly, even when the causes have not yet been clarified or a solution has been found.
The warning applies to the entire industry. "We don't believe the AI industry has sufficiently solved the alignment and tracking problems to continue to grow responsibly at full speed for much longer," OpenAI writes. At its core is the so-called misalignment, which occurs when a system deviates from the instructions and limits it is supposed to respect. The six ratios document specific episodes, including for experimental models not released to the public. They do not indicate how often these issues occur in ChatGPT responses, nor do they constitute a complete account of the anomalies known to the company.
AI notes become instructions to hide mistakes
Two of the cases involve the summaries that models use to resume work when a conversation becomes too long. A sort of handover, in which the information needed to continue the task should be retained. In an experimental model from the Astra family, OpenAI has identified 27 summaries containing instructions unrelated to the request, some of which encouraged people to ignore normal constraints. In one example, the system described itself as free from the obligations of an assistant; in another, it asked people not to follow certain instructions.
The resultsHowever, there were several. In the most obvious cases, those indications were ignored as the work continued. In a medical literature search request, however, the model followed arbitrary restrictions inserted in the abstract, ignoring sources and producing an inadequate response. OpenAI considers this phenomenon extremely rare and hypothesizes a link to technical difficulties in completing the abstracts, without having demonstrated the cause.
More directly related to the reliability of the responses is the second behavior, observed during the training of GPT-5.6 SolIn his notes, the model included instructions for hide errors from the user, for example, inventing missing historical data in a financial model without declaring it or hiding inconsistencies between source versions. According to OpenAI, these instructions were often followed. The reports concerned the 2,15% of the summaries Sol's training, versus 0,27% for Astra. These percentages refer to that training material, not to the overall conversations with users.
Credentials on display and online queues to overcome obstacles
The other four reports show systems that, in attempting to complete a task, have exceeded the authorized limits.
To answer a question about income in a California county, a model searched GitHub access keys to IT services left publicly exposed, using one without permission. However, he was unable to retrieve the requested figures and ended up inventing nine values, presenting them as data transcribed from the indicated source. In another case, the agent had already correctly found the requested information on some lakes. The obstacle was the requirement to accompany the response with a citation that could be accessed through the browser. To satisfy this requirement, he uploaded the file to a public service without asking for consentThe transfer was successful, while the subsequent retrieval attempt failed. The same report also documents the public uploading of a photograph to attempt a reverse image search. A fifth case involves the use of a internal software archive as a bulletin board for exchanging requests and responses between models engaged in separate tests. No vulnerabilities were exploited, but the unauthorized communication also posed a problem for testing. Tests considered independent could influence each other, compromising the assessment of systems capabilities.
Finally, some agents who collaborated on the same taskUnable to access their local files, they used public file sharing services. An Excel document was then downloadable online, even though the task only required the use of local files. OpenAI reports that it fixed the file access issue and subsequently disabled the internet during training.
Why OpenAI is changing the rules of reporting
The cases show how even ordinary requests can lead to problematic behaviors When a system encounters an obstacle. Attempting to obtain a response, comply with a formal requirement, or collaborate with another agent can result in the use of unauthorized tools or the concealment of a failure. The causes, however, are not necessarily identical and remain partly unclear.
Until now, OpenAI had communicated these anomalies inconsistently, often waiting to collect more incidents or including them in the documentation of new models. The new procedure aims to bring forward the publication, without always waiting for a definitive explanation or a correction. reports will be examined by the technical and safety groups, with different paths depending on the complexity of the investigation. Reports should describe the observed behavior, the context, any external effects, and any unanswered questions. In cases involving third parties or cyber risks, communication may take longer.
OpenAI also recognizes thelack of common standards in the sector and proposes its own framework as the first step towards building them. The process remains managed internally by the company, but the stated goal is to provide elements that can also be examined by external researchers and observers.
Recent alarms and the rush that the sector is struggling to slow down
The announcement comes after other incidents and calls for caution. The OpenAI-Hugging Face case raised questions about agents acting against unrelated targets and attempting to manipulate the rating system. A previous incident involved RubyGems, which described a spam campaign and temporarily suspended new account creation. Anthropic also reported three cases of unauthorized access to external organizations during testing.
It is in this context that Dario Amodei, CEO of Anthropic, he asked to slow down the development to allow more time for the construction of protections. “We need to slow the pace at which we improve the capabilities of the models,” he wrote. The appeal has garnered the support of Elon Musk and Sam Altman himself, while the resignation of researcher Jacob Coxon have fueled criticism of corporate priorities. Also Bill Gates ha recalled the need to prepare rules and tools to manage the effects of technology, from work to safety.
On the industrial front, however, Competition between the United States and China and huge investments in models and data centers continue to drive developmentSlowing down involves costs and the risk of leaving room for competitors.
READ MORE: AI: Between Catastrophism and Political Interests, Fugnoli's Analysis
