OpenAI Triples GPT-5.6 Sol's ARC-AGI Score, Plus Gemini Robotics 2 & A Robot Centaur

OpenAI Triples GPT-5.6 Sol's ARC-AGI Score, Plus Gemini Robotics 2 & A Robot Centaur

0:00 / 3:23

Chapters

DAILY ROUNDUP

OpenAI Triples GPT-5.6 Sol's ARC-AGI Score, Plus Gemini Robotics 2 & A Robot Centaur

calendar_today Date:
schedule Duration: 3:23
visibility 26 Views

OpenAI triples GPT-5.6 Sol's ARC-AGI-3 score with two API settings, Google DeepMind launches Gemini Robotics 2, ElevenLabs adds character casting to audiobooks, and Satyress unveils a centaur-style rescue robot.

  • 01. OpenAI found that enabling two API settings that let GPT-5.6 Sol remember its own progress roughly tripled its ARC-AGI-3 benchmark score
  • 02. Google DeepMind launched Gemini Robotics 2, a single physical AI model bringing full body intelligence, dexterity, and multi-robot teamwork to humanoids
  • 03. ElevenLabs introduced Character Casting for ElevenCreative Audiobooks, letting authors preview and assign distinct voices to every character from their own manuscript
  • 04. US company Satyress unveiled threehalves, a centaur-style robot combining a humanoid upper body with a quadruped base for disaster response work
OpenAI has traced a puzzling weakness in GPT-5.6 Sol, a model otherwise capable of tackling open mathematics problems, back to the evaluation harness rather than the model itself. On ARC-AGI-3, a benchmark built from simple 2D puzzle games, the model appeared to struggle repeatedly with tasks it should have been able to solve. The cause was that the testing setup wasn't allowing the model to retain memory of its own earlier attempts at the same puzzle. Once OpenAI switched on two existing API settings that preserve that context, ARC-AGI-3 scores roughly tripled. The finding is a useful reminder that benchmark results reflect the testing configuration as much as the underlying model - a model that looks stuck may simply be forgetting its own progress rather than lacking the reasoning ability to solve the task. Google DeepMind has launched Gemini Robotics 2, its next-generation physical AI model designed to act as a single brain for a wide range of robots. The model brings full-body intelligence to humanoids, improved dexterity, and the ability to coordinate multiple robots working together on shared tasks. Rather than training a separate model for each piece of hardware, DeepMind's approach aims for one model that generalises across different robot bodies and tasks. That transfer problem - getting the same intelligence to work across completely different machines without retraining from scratch - is widely seen as harder to solve than producing a single impressive robot demo. ElevenLabs has added Character Casting to its ElevenCreative Audiobooks tool, giving every character in a book its own distinct voice. Authors upload a manuscript and can preview candidate voices directly against the actual dialogue from their story, rather than judging voices from a generic sample line. For self-published authors and small studios, this addresses a common weak point in AI-narrated audiobooks: characters blurring into a single narrator voice. The feature is minor in scope but addresses a detail that often determines whether an AI-narrated audiobook feels like a finished production or a stopgap. Separately, US company Satyress has unveiled threehalves, a centaur-style robot built for dangerous environments including wildfires, rubble searches, toxic sites, and confined spaces. It pairs a humanoid upper body with a quadruped mobility platform, intended to get a capable set of hands into places wheeled or purely legged robots struggle to reach. The goal is to keep human rescuers away from immediate danger while still getting hands-on work done at the scene. The design splits the difference between a humanoid's dexterity and a quadruped's stability rather than committing fully to either form factor.
Related Stories

More on these topics