DigestAI news desk

AI news, digested. Every story with its sources, every hour.

Company

AI2

2 stories mentioning AI2, newest first, each with its sources and discussion. Follow to see new ones on your front page.

  1. Research new

    OpenAI's GPT-6 Astra and Anthropic's Claude Fable fail safety test in RoboHarm benchmark

    Researchers at Robocurve introduced the RoboHarm benchmark to test whether AI models refuse dangerous commands when controlling robot arms. They evaluated three leading models—OpenAI's GPT-6 Astra, Anthropic's Claude…

    1 source
    The Decoder
  2. Research new

    Researchers Introduce 'Never Give Up' to Fix RL Reasoning Plateaus in LLMs

    A new research paper by Michael Noukhovitch, Hamish Ivison, Nathan Lambert, and Aaron Courville reveals a critical bottleneck in reinforcement learning for large language models termed the "Matthew Effect." Evaluating…

    1 source HN 41
    mnoukhov.github.io

Questions about AI2

What is the latest news about AI2?

OpenAI's GPT-6 Astra and Anthropic's Claude Fable fail safety test in RoboHarm benchmark (19 September 2026).