Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/
Winner – Opus
Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/
Winner – Critique
I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.
This round the battle is between all the models available in Copilot Cowork Claude. Namely, these:
The other interesting factor here is that all these models as PAYG, so I’ll also give the costs for the same prompt.
To assess all these results together I found Gemini and SharePoint Copilot to both fail completely. Gemini keeps saying that it can’t find all the document, even though I have uploaded and also tried linking. SharePoint Copilot on the other hand starts processing but never gives me a result, no matter how long I wait. I will admit that these files are quite long (20+ pages) and quite complex, so there is a lot to digest. That said I threw them into Grok and those results are here:
In summary, the ranking where:
1. Opus 5
2. Fable 5 (Preview)
3. Fable 5 (Copilot) (Preview)
4. Opus 4.8
5. Sonnet 5
which all kind of makes empirical sense but still quite subjective I feel. However, for now I’ll swap to using Grok to evaluate these documents as a standard approach.
Now for the costs which don’t factor into the results:
Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits
Which, on initial analysis, indicates that costs for all models are in the same kind of range (i.e. at least US$20 per prompt), with Fable 5 models being about 50% more expensive.
Given the analysis and costs, it seems Opus 5 with Cowork, if you are using Claude, is the most cost effective for the best result in Cowork.
So:,
Round 3 winner – Copilot Cowork Claude = Opus 5
Onto the next round.