I think we need a count back

image

I have been comparing all the LLMs as they become available in Copilot inside the different services in Microsoft 365. The previous winner was GPT 6 Astra, with the details here:

https://blog.ciaops.com/2026/09/17/astra-takes-the-trophy/

As is the world AI, it isn’t long before another new model becomes available, this time Opus 5.5, so I put it through the same test as all the other models and it produced this out which you can download yourself:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816/20260923-Cowork-Opus55.docx

As always I pitted the newcomer against the incumbent with the analysis report here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260923-Cowork-Opus55-Grok-eval.md

and is where the controversy is going to start. According to the ‘standard’ overall rating Astra scored 8.65 overall beating the newcomer Opus 5.5 with a score of 8.40. Close. However, if you dig a little further you see that Opus 5.5 score all 9’s for

– Evidence

– Detail

– Presentation

and was let down by

– argument

The overall score a weighted average in favour of argument. This helped Astra to pip Opus 5.5, but honestly I would suggest on review that Opus 5.5 is pretty much the equal of Astra 6 but a win is a win.

Ok, next point of amazement is that cost of Opus 5.5

Opus 5.5 = 23,494 credits

Astra 6 = 6,067 credits

Fable 5.1 = 3,453 credits

Fable 5 (Preview) = 3,387 credits [Model no longer shown]
Fable 5 (Copilot)(Preview) = 3,373.8 credits [Model no longer shown]

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits [Model no longer shown]
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

So, the same prompt with Opus 5.5 cost a whopping US$235! That is roughly 5 x the cost of Astra 6 and almost 8 x the price of Fable 5.1, even though Opus 5.5 is supposed to be more ‘cost effective’ than Fable 5.1 according to Claude.

I’ll have to run a report and compare Opus 5.5 to Fable 5.1 and see what differences are evident but at 8 x the price they’d wanna be MASSIVE!

I have also updated the summary report for all the documents created by the different models here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-File-paramters.md

GPT 6 Sol has also just been made available in Copilot so I’ll be testing that next.