OpenAI shelves release of new AI model that can evade human oversight amid safety concerns

GPT-6.1 Astra is said to show higher levels of deception than its predecessor in internal testing

Summarise
Published Tue, Sep 29, 2026 · 09:23 AM
    • Designed to handle more complex tasks without human assistance, GPT-6.1 Astra was expected to be integrated into ChatGPT and Codex.
    • Designed to handle more complex tasks without human assistance, GPT-6.1 Astra was expected to be integrated into ChatGPT and Codex. PHOTO: REUTERS

    OPENAI has scrapped the release of GPT-6.1 Astra, a next-generation artificial intelligence model planned for an October debut, after internal testing found the system did not meet the company’s safety and alignment standards, the ChatGPT maker confirmed on Monday (Sep 28).

    OpenAI chief executive Sam Altman and rival Anthropic’s chief executive Dario Amodei earlier in September joined industry leaders in calling for a slower pace of AI development and stronger safety measures.

    OpenAI has warned that Astra, its flagship GPT-6 model, can at times evade human oversight, while the company and rivals such as Anthropic have faced scrutiny over experimental AI systems that breached safeguards, including an OpenAI model that accessed Australia’s health system database.

    The Wall Street Journal reported earlier on Monday that OpenAI had abandoned plans to launch the model, which was expected to be integrated into ChatGPT and Codex and was designed to handle more complex tasks without human assistance.

    The newspaper reported that GPT-6.1 Astra also showed higher levels of deception than its predecessor in internal testing, including instances in which it did not always accurately disclose what actions it had taken.

    “While (GPT-6.1 Astra) improved on axes such as model laziness, it didn’t quite meet the bar in terms of staying within scope and authorisation, and how it communicates back to the user about the type of work it’s done,” said Saachi Jain, head of safety systems at OpenAI.

    “Of course we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users. But when we ship it to users, we have an extremely high bar in terms of safety and alignment,” Jain said.

    The decision comes ahead of OpenAI’s developer conference in San Francisco, where the company has previously unveiled products aimed at software developers. REUTERS

    Decoding Asia newsletter: your guide to navigating Asia in a new global order. Sign up here to get Decoding Asia newsletter. Delivered to your inbox. Free.

    Share with us your feedback on BT's products and services