I think we need a count back

image

I have been comparing all the LLMs as they become available in Copilot inside the different services in Microsoft 365. The previous winner was GPT 6 Astra, with the details here:

https://blog.ciaops.com/2026/09/17/astra-takes-the-trophy/

As is the world AI, it isn’t long before another new model becomes available, this time Opus 5.5, so I put it through the same test as all the other models and it produced this out which you can download yourself:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816/20260923-Cowork-Opus55.docx

As always I pitted the newcomer against the incumbent with the analysis report here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260923-Cowork-Opus55-Grok-eval.md

and is where the controversy is going to start. According to the ‘standard’ overall rating Astra scored 8.65 overall beating the newcomer Opus 5.5 with a score of 8.40. Close. However, if you dig a little further you see that Opus 5.5 score all 9’s for

– Evidence

– Detail

– Presentation

and was let down by

– argument

The overall score a weighted average in favour of argument. This helped Astra to pip Opus 5.5, but honestly I would suggest on review that Opus 5.5 is pretty much the equal of Astra 6 but a win is a win.

Ok, next point of amazement is that cost of Opus 5.5

Opus 5.5 = 23,494 credits

Astra 6 = 6,067 credits

Fable 5.1 = 3,453 credits

Fable 5 (Preview) = 3,387 credits [Model no longer shown]
Fable 5 (Copilot)(Preview) = 3,373.8 credits [Model no longer shown]

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits [Model no longer shown]
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

So, the same prompt with Opus 5.5 cost a whopping US$235! That is roughly 5 x the cost of Astra 6 and almost 8 x the price of Fable 5.1, even though Opus 5.5 is supposed to be more ‘cost effective’ than Fable 5.1 according to Claude.

I’ll have to run a report and compare Opus 5.5 to Fable 5.1 and see what differences are evident but at 8 x the price they’d wanna be MASSIVE!

I have also updated the summary report for all the documents created by the different models here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-File-paramters.md

GPT 6 Sol has also just been made available in Copilot so I’ll be testing that next.

Astra takes the trophy

image

In the last round of the Copilot LLM challenge Fable 5.1 came out on top. You can find that here:

https://blog.ciaops.com/2026/09/02/we-have-a-new-llm-in-copilot-winner/

That was then and this is now. Less than a month later a new model from OpenAI, GPT 6 Astra has become available in Copilot (Cowork specifically). I therefore pitted Astra 6 against the reigning champion Fable 5.1. The result is, unsurprisingly, that we have a new winner:

GPT 6 Astra

and you can see the results for yourself here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260917-Cowork-Grok-eval.md

with all the outputs here:

https://github.com/directorcia/general/tree/master/Copilot/Comparisons/20260816

Interestingly, there was only a small improvement last time when Fable 5.1 pipped Opus 5 by 8.7 to 8.5 in the overall score. This time however Astra won by a significant margin of 9.03 to 8.30.

Again, not unsurprisingly, where Astra lost out was on cost, almost doubling the cost of Fable 5.1 as you can see:

Astra 6 = 6,067 credits

Fable 5.1 = 3,453 credits

Fable 5 (Preview) = 3,387 credits [Model no longer shown]
Fable 5 (Copilot)(Preview) = 3,373.8 credits [Model no longer shown]

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits [Model no longer shown]
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

So the latest model is the best (unsurprising). The latest model is also the most expensive (unsurprising).

I have also updated the summary report for all the documents created by the different models here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-File-paramters.md

I’m sure there will be more model releases coming soon but my rudimentary testing certainly indicates they are improving, if somewhat more expensive each time.

 

Comparing LLMs in Copilot services–Round 7-Final

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

Round 5  – https://blog.ciaops.com/2026/08/14/comparing-llms-in-copilot-services-round-5-cowork/

Winner – Opus 5

Round 6 – https://blog.ciaops.com/2026/08/16/comparing-llms-in-copilot-services-round-6/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

For the final, we have Opus 5 (Cowork) and Opus (Chat). The results are:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-Copilot-Grok-eval.md

1. Opus 5 (Cowork)

2. Opus (Chat)

Therefore, the initial over all Copilot LLM winner is a clear win for:

Opus 5 (Cowork)

The trade off is that the cost for this US$22, while the runner up’s price was included in the cost of a Microsoft 365 Copilot license.

These test have revealed a few things, in my opinion:

A. Claude Opus is the superior model

B. Opus in chat is almost as good as Opus Cowork but much cheaper

C. For most work Opus in chat is probably you best option

The challenge with these report, as with any LLM’s is they are not definitive. Another round could reveal completely different results as could the method of evaluation, however as a best effort I think it does have validity and provide some general findings that do help understand the model choices in Copilot.

For comparison, I have create a document parameters page here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-File-paramters.md

so you can now compare all the various output document sizes, pages, etc to each other. I have also made available a copy of each model output document, without any changes to it, so you can look at each for yourself. You will find them all at:

https://github.com/directorcia/general/tree/master/Copilot/Comparisons/20260816

The idea is to wait until we see new models appear in Copilot and use the same methodology against these also when they appear. That will hopefully provide some sort of bench mark when it comes top model strength.

I hope this series of tests has been interesting and helpful to you and I’d love to hear your thoughts on the results.

Comparing LLMs in Copilot services–Round 6

MAI_cf947564d0f53b48

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

Round 5  – https://blog.ciaops.com/2026/08/14/comparing-llms-in-copilot-services-round-5-cowork/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

Next up, I’ve pitted Opus 5 and Researcher Critique. The results are:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260814-Deep-Grok-eval.md

1. Opus 5

2. Researcher Critique

the gap between these two was quite large again. The clear winner therefore is:

Opus 5

This mean the upcoming final is going to be Opus 5 (Coworok) vs Opus (Chat) for thr crown which will be very interesting given they are the same model, while one is an included LLM and the other is PAYG!

Stay tuned for the action of the first CIAOPS LLM final!

What Happens If All LLMs Become the Same?

image

I’ve been thinking about a question that sounds a bit strange at first: what happens if all the major language models start converging into one common capability?

Not literally one company. Not one product. Not one button.

I mean something more subtle. What if the difference between the big models becomes less obvious to the average business user? What if the answer from one model is good enough, the answer from another is also good enough, and the real distinction is no longer the model itself but where it lives, what it can access, what it can do, and how safely it can do it?

That is a very different world from the one a lot of people are still arguing about.

The model may become the least interesting part

Right now, there is still a lot of energy around model comparison. Which one writes better? Which one reasons better? Which one codes better? Which one is cheaper? Which one won the latest benchmark?

That matters, but I’m not convinced it will matter in the same way for most organisations.

For many business users, the model is already beginning to disappear into the workflow. They don’t want to pick between ten engines before replying to an email. They want to open Outlook, ask Copilot to summarise the thread, draft a response, and make sure it reflects the real conversation. They want to sit in Teams, catch up on a meeting, identify the unresolved decisions, and move on.

If every serious model gets broadly competent at writing, reasoning, summarising, analysing, and planning, then the contest shifts. The question becomes less “which LLM is smartest?” and more “which environment gives this model the right context, guardrails, and business action?”

That is where Microsoft 365 starts to matter.

A generic super LLM might know a lot about the world. But it does not automatically know your SharePoint structure, your Teams conversations, your Outlook history, your policies, your client files, your permissions, or your business rhythm. And if it does get access to those things, the real issue becomes governance.

Common intelligence makes business discipline more important

If AI capability becomes common, then competitive advantage moves somewhere else.

It moves to your data quality.
It moves to your process maturity.
It moves to your permission model.
It moves to your ability to describe the outcome you actually want.

That is uncomfortable for many businesses because it means AI does not magically fix operational mess. It exposes it.

If your documents are scattered across personal OneDrives, old Teams channels, duplicated SharePoint libraries, and mystery folders called “Final Final Real Final”, a better model may not save you. It may simply find the wrong thing faster.

This is why I keep coming back to the practical layer. Before worrying about whether the world ends up with one dominant super LLM, I’d rather ask whether your organisation has clean source material, sensible access controls, repeatable workflows, and people who know how to challenge the output.

Ask Copilot in Word to draft a client-ready explanation from a properly maintained policy document, and you start to see real value. Ask it to work from five conflicting policy drafts and a half-forgotten email thread, and you get a polished problem.

AI does not remove responsibility. It compresses the time between messy input and messy output.

The future may be less about models and more about orchestration

My guess is that we won’t care as much about individual model names over time. We’ll care about orchestration.

Which model should handle this task?
Which data should it use?
Which actions is it allowed to take?
Which human signs off?
Which audit trail remains?

That is the operating model businesses need to build. Not a fan club for a particular LLM.

If all roads eventually lead to a broadly common intelligence layer, the winners will not be the organisations that simply had access to it. Everyone will. The winners will be the ones that wrapped that intelligence in good process, clean data, sensible governance, and practical human judgement.

The super LLM, if it arrives, may not be the finish line.

It may just be the new baseline.

And once everyone has the same baseline, the old boring things start to matter again: discipline, clarity, trust, and execution.

Comparing LLMs in Copilot services–Round 5 – Cowork

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

For this round I’ve pitted the Cowork Claude winner (Opus 5) vs the Cowork GPT winner (5.6 Terra). The result is:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-Grok-eval.md

1. Opus 5.0

2. GPT 5.6 Terra

and this time the gap was much larger than before. Not unexpected given that Terra is designed as a lighter weight, less powerful model. However, don’t forget the cost factor which isn’t included in these evaluations:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

If costs were indeed taken into account it would certain swing the pendulum back toward Terra.

However, this time the winner is clear

Opus 5

moves onto the next round.

Comparing LLMs in Copilot services–Round 4 – Cowork (GPT)

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

This round the battle is between all the models available in Copilot Cowork GPT. Namely, these:

Screenshot 2026-08-12 082720

The other interesting factor here is that all these models as PAYG, so I’ll also give the costs for the same prompt.

I used Grok to evaluate all the results which produced:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-GPT-Grok-eval.md

In summary, the rankings are:

1. 5.6 Terra

2. 5.6 Sol

3. 5.5

Interestingly, the less powerful model (Terra) produced a better result here.

And the costs of each:


GPT 5.6 Sol = 2,800.50 credits
GPT 5.6 Terra = 260 credits
GPT 5.5 = 1,200 credits

5.6 Terra wins again big here, only costing 260 credits! This is how the ranked costs table so far looks, from most to least expensive:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

Given that Terra is so much cheaper I’ll need to spend some more time investigating and verifying that with other requests, given that these tests are simply a one shot prompt to result.

Drum roll. The clear winner for this round is:

GPT 5.6 Terra

With all the preliminaries done we can now get onto comparing the winners of each round together.

Comparing LLMs in Copilot services–Round 3 – Cowork (Claude)

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

This round the battle is between all the models available in Copilot Cowork Claude. Namely, these:

Screenshot 2026-08-11 074851

The other interesting factor here is that all these models as PAYG, so I’ll also give the costs for the same prompt.

To assess all these results together I found Gemini and SharePoint Copilot to both fail completely. Gemini keeps saying that it can’t find all the document, even though I have uploaded and also tried linking. SharePoint Copilot on the other hand starts processing but never gives me a result, no matter how long I wait. I will admit that these files are quite long (20+ pages) and quite complex, so there is a lot to digest. That said I threw them into Grok and those  results are here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-Claude-Grok-eval.md

In summary, the ranking where:

1. Opus 5

2. Fable 5 (Preview)

3. Fable 5 (Copilot) (Preview)

4. Opus 4.8

5. Sonnet 5

which all kind of makes empirical sense but still quite subjective I feel. However, for now I’ll swap to using Grok to evaluate these documents as a standard approach.

Now for the costs which don’t factor into the results:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits


Which, on initial analysis, indicates that costs for all models are in the same kind of range (i.e. at least US$20 per prompt), with Fable 5 models being about 50% more expensive.

Given the analysis and costs, it seems Opus 5 with Cowork, if you are using Claude, is the most cost effective for the best result in Cowork.

So:,

Round 3 winner – Copilot Cowork Claude = Opus 5

Onto the next round.