Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/
Round 1 Winner – Copilot Chat = Opus
This is round 2 of my Copilot LLM battle, with the aim to determine the ‘best’ result from all the AI services and all the available LLMs used with Copilot in Microsoft 365.
The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.
This round is all the possibilities with Researcher. As before, rathet than duplicate the reports here I have uploaded them to my Githuv repository in markdown format:
SharePoint Copilot assessment:
https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260630-Researcher.md
Gemini assessment:
Results (out of 10):
Critique – 9.30
Claude – 8.88
Council – 7.83
Auto – 7.63
Notes – SharePoint and Gemini produced very different rankings. I’m sticking with SharePoint’s assessment for consistency. I will also point out that gettign Gemini to do this comparison and produce a simple markdown file with the result was pretty much impossible. It does a really poor job on this simple comparison ask as you can see from the results.
So the Round 2 winner – Copilot Researcher = Critique
Onto Round 3