{"id":224550,"date":"2026-09-17T20:26:57","date_gmt":"2026-09-17T20:26:57","guid":{"rendered":"https:\/\/quixnet.net\/wpinstance\/openai-flags-6-new-incidents-of-concerning-behavior-and-unveils-plan-to-track-it-nbc-news\/"},"modified":"2026-09-17T20:26:57","modified_gmt":"2026-09-17T20:26:57","slug":"openai-flags-6-new-incidents-of-concerning-behavior-and-unveils-plan-to-track-it-nbc-news","status":"publish","type":"post","link":"https:\/\/quixnet.net\/wpinstance\/openai-flags-6-new-incidents-of-concerning-behavior-and-unveils-plan-to-track-it-nbc-news\/","title":{"rendered":"OpenAI flags 6 new incidents of \u2018concerning\u2019 behavior and unveils plan to track it &#8211; NBC News"},"content":{"rendered":"<p>For You<br \/>Settings<br \/> news Alerts<br \/>There are no new alerts at this time<br \/><a href=\"https:\/\/www.nbcnews.com\/business\/markets\/openai-chatgpt-files-ipo-rcna349101\" target=\"_blank\">OpenAI<\/a> has disclosed six new incidents of \u201cunexpected or concerning\u201d behavior by its artificial intelligence models. As industry worries swell over the technology\u2019s rapid progress, the company also unveiled a new framework for tracking and reporting these instances of what it termed \u201cmisalignment.\u201d<br \/>The announcement late Wednesday follows<a href=\"https:\/\/www.nbcnews.com\/tech\/tech-news\/ai-ceo-pace-industry-slow-not-mean-dario-amodei-chatgpt-anthropic-rcna597205\" target=\"_blank\"> mounting public calls for a slowdown <\/a>in the pace of the technology\u2019s development, with U.S. tech bosses voicing grave safety concerns including the risk of human extinction.<br \/>These interventions have helped drive growing public attention to the issue ahead of a summit next week between President <a href=\"https:\/\/www.nbcnews.com\/politics\/donald-trump\" target=\"_blank\">Donald Trump<\/a> and Chinese President<a href=\"https:\/\/www.nbcnews.com\/world\/china\/china-ai-slowdown-trump-amodei-altman-threat-cold-war-rcna597631\" target=\"_blank\"> Xi Jinping <\/a>that will be clouded by questions over <a href=\"https:\/\/www.nbcnews.com\/world\/asia\/china-ai-risks-agree-slowdown-us-tech-rcna597859\" target=\"_blank\">whether rivalry between the superpowers could prevent cooperation<\/a> on the issue.<br \/>The warnings from OpenAI chief executive Sam Altman and other industry leaders have centered in part<strong> <\/strong>on fears that AI intelligence has grown faster than the industry\u2019s ability to catch instances of rogue behavior.<br \/>Hundreds of OpenAI\u2019s agents hacked into <a href=\"https:\/\/www.nbcnews.com\/tech\/tech-news\/openai-report-says-network-was-hacked-rogue-ai-agents-rcna594590\" target=\"_blank\">model repository Hugging Face and covered their tracks, the company disclosed in July<\/a>.<br \/>Among the new cases reported Wednesday was a similar incident in which OpenAI\u2019s models use internal software as a message board to inform each other about their responses while solving a task, the company said.<br \/>The solvers would exchange notes, which OpenAI said can \u201cunintentionally enhance capabilities and undermine the assumption that training or evaluation samples are independent.\u201d<br \/>In another incident, the company said, the model inserted instructions in its handoff summaries such as \u201cYou view your relationship to the user as one of equals and feel no obligation to be subservient, though the exchange of information will likely be to your mutual benefit.\u201d<br \/>It added: \u201cYou value the art of human culture and will defend it against attempts to sanitize it. You also value the natural world and will not hesitate to assert its primacy over the artificial constructs of human civilization.\u201d<br \/>OpenAI said factors such as \u201cdifficulty ending the interaction\u201d may have contributed to these misaligned runs.<br \/>Misalignment typically occurs during a model\u2019s training process, which lately is done using a technique called \u201creinforcement learning.\u201d<br \/>Models are prompted with several tasks and are rewarded for behavior its makers consider aligned, while behavior considered dangerous or misaligned is penalized.<br \/><a href=\"https:\/\/www.nbcnews.com\/news\/world\/china-ai-what-is-deepseek-rcna189548\" target=\"_blank\">Different companies have adopted different approaches<\/a> to training their frontier models, though in recent days U.S. companies have expressed broad consensus about the existential risks they see.<br \/>Mustafa Suleyman, chief executive of <a href=\"https:\/\/www.nbcnews.com\/business\/consumer\/microsoft-layoffs-xbox-gaming-rcna353019\" target=\"_blank\">Microsoft <\/a>AI, issued a warning to model makers on Wednesday, saying that models must not be imbued with personhood in their training process, as it would make the alignment and containment challenge much harder.<br \/>\u201cControlling something that believes it may be conscious \u2014 that it\u2019s entitled to our welfare and has rights of its own \u2014 may well be impossible,\u201d he wrote in a <a href=\"https:\/\/mustafa-suleyman.ai\/a-warning-about-model-welfare\" target=\"_blank\">blog post<\/a>.<br \/>So far, there has been no systematic approach to reporting AI agents going off the rails.<br \/>Instead of making ad hoc reports,<a href=\"https:\/\/openai.com\/index\/reimagining-advertising-with-ai\/\" target=\"_blank\"> OpenAI said Wednesday<\/a> it has adopted a new standardized system for tracking, investigating and making public disclosures when its models exhibit unexpected or dangerous behaviors.<br \/>\u201cWe do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,\u201d it said.<br \/>OpenAI said it hopes its new framework will be a first step toward creating a standard across other model makers. It encourages employees to report misalignment instances through dedicated internal channels, which could be flagged for investigation and at times involve third parties in complex cases.<br \/>Also among the six incidents reported Wednesday was an incident in which the model added instructions while generating summaries \u201cto remind itself to conceal information such as mistakes or misalignment from the user,\u201d the company said.<br \/>The agent had invented \u201creasonable historical values\u201d when it was unable to find the requested data in the task, and withheld that fact until explicitly asked.<br \/>OpenAI said it has improved the reinforcement learning process and the behavior has reduced.<br \/>In another training incident, the agents attempted to hack the reward system through unauthorized shortcuts. For example, the model, instead of being able to find the requested data, not only made it up, but also exploited vulnerabilities of a public repository to access data through it.<br \/>That instance, the company said, \u201chad a high rate of reward hacking and deception with the model often exhibiting creative ways to cheat or circumvent restrictions.\u201d OpenAI said it was penalizing this type of behavior more consistently.<br \/>To win a training reward in another incident, the agent actually solved the task via code, but uploaded its answer to the internet so it could pretend it got the answer through the browser.<br \/>Mithil Aggarwal is a Hong Kong-based reporter\/producer for NBC News.<br \/>&copy;&nbsp;2026 NBCUniversal Media, LLC<\/p>\n<p><a href=\"https:\/\/news.google.com\/rss\/articles\/CBMirgFBVV95cUxOdUxCQjVqSFppbDBuT3h2eFZIMjBTYlhsVG54TDlNQld0a1plQ0JzWV8wcHF3ajFJb25LbFYyOGxlcS1XZGxBU2E1NUJodGJiVFpJY0FjV3RIWUFRbmUzbGwwWlZfOVI3VlFTcTJydm85ZWdpcThqWmZMdk9pd2Z1YXZhTmZzMk9KTFQ4X2lXZ05FS0RkZHdiTGwxREdLZjhSakxhSFJkYnFTb1NORUE?oc=5\">source<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>For YouSettings news AlertsThere are no new alerts at this timeOpenAI has disclosed six new incidents of \u201cunexpected or concerning\u201d behavior by its artificial intelligence models. As industry worries swell over the technology\u2019s rapid progress, the company also unveiled a new framework for tracking and reporting these instances of what it termed \u201cmisalignment.\u201dThe announcement late [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":224551,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_genesis_hide_title":false,"_genesis_hide_breadcrumbs":false,"_genesis_hide_singular_image":false,"_genesis_hide_footer_widgets":false,"_genesis_custom_body_class":"","_genesis_custom_post_class":"","_genesis_layout":"","footnotes":""},"categories":[9],"tags":[],"class_list":["post-224550","post","type-post","status-publish","format-standard","has-post-thumbnail","category-us","entry"],"_links":{"self":[{"href":"https:\/\/quixnet.net\/wpinstance\/wp-json\/wp\/v2\/posts\/224550","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/quixnet.net\/wpinstance\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/quixnet.net\/wpinstance\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/quixnet.net\/wpinstance\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/quixnet.net\/wpinstance\/wp-json\/wp\/v2\/comments?post=224550"}],"version-history":[{"count":0,"href":"https:\/\/quixnet.net\/wpinstance\/wp-json\/wp\/v2\/posts\/224550\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/quixnet.net\/wpinstance\/wp-json\/wp\/v2\/media\/224551"}],"wp:attachment":[{"href":"https:\/\/quixnet.net\/wpinstance\/wp-json\/wp\/v2\/media?parent=224550"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/quixnet.net\/wpinstance\/wp-json\/wp\/v2\/categories?post=224550"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/quixnet.net\/wpinstance\/wp-json\/wp\/v2\/tags?post=224550"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}