Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/
Winner – Opus
Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/
Winner – Critique
Round 3 – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/
Winner – Opus 5
I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.
This round the battle is between all the models available in Copilot Cowork GPT. Namely, these:
The other interesting factor here is that all these models as PAYG, so I’ll also give the costs for the same prompt.
I used Grok to evaluate all the results which produced:
In summary, the rankings are:
1. 5.6 Terra
2. 5.6 Sol
3. 5.5
Interestingly, the less powerful model (Terra) produced a better result here.
And the costs of each:
GPT 5.6 Sol = 2,800.50 credits
GPT 5.6 Terra = 260 credits
GPT 5.5 = 1,200 credits
5.6 Terra wins again big here, only costing 260 credits! This is how the ranked costs table so far looks, from most to least expensive:
Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits
GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits
GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits
Given that Terra is so much cheaper I’ll need to spend some more time investigating and verifying that with other requests, given that these tests are simply a one shot prompt to result.
Drum roll. The clear winner for this round is:
GPT 5.6 Terra
With all the preliminaries done we can now get onto comparing the winners of each round together.