Comparing LLMs in Copilot services–Round 5 – Cowork

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

For this round I’ve pitted the Cowork Claude winner (Opus 5) vs the Cowork GPT winner (5.6 Terra). The result is:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-Grok-eval.md

1. Opus 5.0

2. GPT 5.6 Terra

and this time the gap was much larger than before. Not unexpected given that Terra is designed as a lighter weight, less powerful model. However, don’t forget the cost factor which isn’t included in these evaluations:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

If costs were indeed taken into account it would certain swing the pendulum back toward Terra.

However, this time the winner is clear

Opus 5

moves onto the next round.

Leave a comment