Did OpenAI Just Achieve AGI?

GPT-6 Astra scored 99.9%. Then I looked at how.

Want to turn your AI knowledge into paid work?

Join my free live workshop. I will show you what businesses pay for, how I structure an AI workshop and how to get started.

Join the free workshop

GPT-6 Astra landed with a score designed to melt the internet: 99.9% on ARC-AGI-3.

Put AGI in the name, add 99.9% and you can imagine what happened next. "AGI has arrived" posts everywhere.

So…? We done here?

Not quite. As always the reality is a little more complex!

What ARC-AGI-3 tests

The ARC-AGI-3 test is a benchmark that has used as a rough proxy for our progress towards AGI. It is specifically about giving AIs a novel problem with no “rulebook” and seeing if the can solve it.

Astra is dropped into an unfamiliar little game with no instructions, rulebook or stated goal. It has to poke around, notice what changes, work out what success looks like, plan a route and correct itself.

Humans can solve all the environments. AIs have (until now) struggled. You can play the game here actually - I recommend it so you know what’s it actually is: https://arcprize.org/arc-agi/3

The score is a composite of level completion and rewards using fewer actions, so blindly clicking everything is a bad strategy.

More interestingly though is OpenAI’s new harness.

ARC Prize ran Astra through a standard, provider-neutral harness and it scored 62.7%. Very good but not the big number everyone is freaking out about.

They then ran it on OpenAI's new Provider Adapter and scored 99.9%. Hmm.

The adapter preserves private reasoning state between requests and compacts long conversations so the model can reuse earlier work. Astra built compact notes for objects, coordinates, rules and unfinished plans. With OpenAI's setup, it completed almost every level and used fewer actions than the human baseline on 96% of them.

It used symbolism. Created its own neuro symbolic view of the world and used this to solve. That’s the more exciting news here as genAI’s lack of symbolic reasoning is often pointed to as a failing - primarily by Gary Marcus, a skeptic who has been banging this drum for quite some time!

But in this case the model created its own neurosymbolic language and used it to ace the test.

So it’s the model's working environment that is doing a lot of the work. Not just MAKE BIGGER = MAKE BETTER

The right memory and tools can make the same model dramatically more capable. We saw a smaller version of this when ChatGPT got reusable Skills.

Oh and is this AGI? Nah. The ARC Prize is explicit that saturating its benchmark does not prove AGI. Despite… putting AGI right in the name and confusing people…

The ARC games are unfamiliar, but still bounded, deterministic and closed-ended. Life is messier. Real world problems are much more complex.

So I am leaving the AGI champagne in the fridge. Still haven’t decided if it’s a celebratory or commiserationary champagne.

Astra can use a computer properly

OK but Astra is still very good. And specifically for computer use.

On OpenAI's published tests, Astra scored 72.6% on OSWorld against Sol's 65.7%, while the work took roughly 40 minutes per task instead of 75. On AutomationBench, which tests professional workflows, Astra scored 41.4% against Sol's 18.1%.

You’ll probably see lots of fancy 3D models on social media that people use to show “wow so fancy, feel the AGI”. That’s because they are visual examples and great for going viral. But what do they tell us practically?

In these cases users handed it Final Cut, Blender and Unreal Engine. Big complex powerful tools.

The big oh shit moment is the increasing ability of the AI to stay oriented inside complicated software for hours or days. Using a range of tools we used to work on. But autonomously.

We are moving towards being able to basically hand off work. Any work we do on a computer. Which for some people is all their work..

If Astra is in your account (it’s rolling out this weekend), give it a job that crosses several apps or takes hours.

Maybe record yourself doing the job once. Give it the files, the result you need and a clear definition of done. Keep purchases, external messages and destructive changes behind an approval step. Generally a good idea!

Then leave it enough room to work. Leave it alone and go outside!

To the Task,

Kyle