- AI with Kyle
- Posts
- Claude Won. GPT Got Cheaper.
Claude Won. GPT Got Cheaper.
Opus 5.5 and GPT-6 Sol are solving different problems.

Watch the full video on YouTube: https://youtu.be/G5f7JApgotE
Still feel like you're winging it with AI?
I'm putting together the AI Confidence Accelerator: 30 days, about 30 minutes a day, covering Chat, Delegate, Build and Teach.
It's for people who use AI but want to understand what they're doing and feel confident explaining it to someone else.
Join the free waitlist and I'll send you first access and a code for 25% off the paid programme at launch.
More models!
Claude Opus 5.5 launched. Less than an hour later, OpenAI released GPT-6 Sol and Luna.
The internet immediately turned it into Claude versus ChatGPT.
Dumb comparison. As always.
Opus 5.5 has taken the lead on the Artificial Analysis Intelligence Index. It scored 58 at maximum effort and led six of the ten tests. It’s good. Real good.
It’s brought me back to the warm loving embrace of Claude. I’ve missed you Claude.
ChatGPT’s upgrades on the other hand were of a different type.
Sol made a much smaller jump in raw capability BUT its cost per task on the same index fell from about $1.99 to $1.06. Almost half.
Anthropic pushed model performance. OpenAI pushed the bill down.
One is going for the intelligence frontier. The other is aiming at cost-efficiency.
Those are both useful improvements. What you DO with AI determines which is better for you.

Standard API prices per million input/output tokens: Astra $10/$50, Opus $4/$20, Sol $2/$10 and Luna $0.10/$0.50.
Pick a job you know
A good way to test new models yourself (and you should) is to have a personal benchmark task you use them on.
Personally I test new models with a newsletter draft from one of my livestreams. I've done that job for years and written about 500,000 words manually, so I know what a good result looks like.
Finding your own version of a personal benchmark tells you waaaay more than another glossy 3D demo. These do great on Twitter but tell us nothing about our actual mileage from the new model.
Nope, your benchmark should come from your own work. Pick something repeatable with an output you can judge: a client proposal, a product brief, a research summary or one of the proper jobs I suggested for GPT-6 Astra.
And don’t just ramp it to the top level intelligence - if you do you’ll probably come away thinking" “yeah. Pretty good.” We can be smarter.
First run it at the lowest effort level in a fresh chat. Then run the same task one level higher. Keep going until extra effort stops producing an answer worth the extra time, usage or money.

Start with the default effort. Turn it up when the task earns it.
Maximum effort can be actively worse. In Simon Willison's pelican test, Opus 5.5 used all 128,000 output tokens on thinking and returned no SVG at all. It tripped over its own feet. Never seen that happen!
Make the switch earn its cost
Should you switch? Probably not.
If you already use Claude, try Opus 5.5 inside the setup you already have. The early evidence says it is a very good model, and it gives you Fable-class intelligence at a lower API price. No reason not to use it.
If you work in ChatGPT or Codex, move a few Astra jobs down to Sol. If Sol does them just as well, you've saved usage without rebuilding anything. Luna makes sense for focused, repeatable jobs where you can check the answer. Keep Astra for problems that justify the extra cost.
But switching from ChatGPT to Claude? Or from Claude to ChatGPT? Probably no need. Switching your whole working life between providers for a tiny benchmark or cost gain is a faff. You have to move files, rewrite instructions and reconnect permissions. That eats time. A shared AI vault gives Claude and ChatGPT the same working context, so you can test a model without moving everything around it. So…do that?
To the task,
Kyle
