Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/
Winner – Opus
Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/
Winner – Critique
Round 3 – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/
Winner – Opus 5
Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/
Winner – GPT 5.6 Terra
Round 5 – https://blog.ciaops.com/2026/08/14/comparing-llms-in-copilot-services-round-5-cowork/
Winner – Opus 5
I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.
Next up, I’ve pitted Opus 5 and Researcher Critique. The results are:
https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260814-Deep-Grok-eval.md
1. Opus 5
2. Researcher Critique
the gap between these two was quite large again. The clear winner therefore is:
Opus 5
This mean the upcoming final is going to be Opus 5 (Coworok) vs Opus (Chat) for thr crown which will be very interesting given they are the same model, while one is an included LLM and the other is PAYG!
Stay tuned for the action of the first CIAOPS LLM final!