Inspired by the recent soccer world cup, I have decided to create a Copilot LLM output comparison challenge.
The plan is to use the same prompt with all examples of different services in Copilot and then different models available in each service. After that, the idea is them to compare the winner of each round to determine the overall winner and to continue to do this on a regular basis as new models and services are added over time.
Thus, the methodology is to use the same complex prompt to generate the result from the model (a report) and then use a standard prompt to evaluate all the results to determine a winner. The easiest comparison method is to use Copilot in SharePoint but the aim will also to be to compare using other models as well.
So, for round 1 I’m going to compare all the models available in Copilot chat. Comparison generated by Copilot for SharePoint.
Rather than try and fit the reports here I will upload them to my Github repository here in markdown format:
https://github.com/directorcia/general/tree/master/Copilot/Comparisons
Comparative Assessment – Live Writer Paste
This first report is now directly available at:
https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260630-Chat.md
The results where (out of 10):
1. Opus – 9.51
2. Sonnet – 9.45
3. GPT 5.6 Thinking – 8.76
4. GPT 5.5 Quick – 7.69
5. Auto – 6.93
So, the winner for Round 1 – Copilot Chat = Opus.
Onto Round 2