Comparing LLMs in Copilot services–Round 7-Final

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

Round 5  – https://blog.ciaops.com/2026/08/14/comparing-llms-in-copilot-services-round-5-cowork/

Winner – Opus 5

Round 6 – https://blog.ciaops.com/2026/08/16/comparing-llms-in-copilot-services-round-6/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

For the final, we have Opus 5 (Cowork) and Opus (Chat). The results are:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-Copilot-Grok-eval.md

1. Opus 5 (Cowork)

2. Opus (Chat)

Therefore, the initial over all Copilot LLM winner is a clear win for:

Opus 5 (Cowork)

The trade off is that the cost for this US$22, while the runner up’s price was included in the cost of a Microsoft 365 Copilot license.

These test have revealed a few things, in my opinion:

A. Claude Opus is the superior model

B. Opus in chat is almost as good as Opus Cowork but much cheaper

C. For most work Opus in chat is probably you best option

The challenge with these report, as with any LLM’s is they are not definitive. Another round could reveal completely different results as could the method of evaluation, however as a best effort I think it does have validity and provide some general findings that do help understand the model choices in Copilot.

For comparison, I have create a document parameters page here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-File-paramters.md

so you can now compare all the various output document sizes, pages, etc to each other. I have also made available a copy of each model output document, without any changes to it, so you can look at each for yourself. You will find them all at:

https://github.com/directorcia/general/tree/master/Copilot/Comparisons/20260816

The idea is to wait until we see new models appear in Copilot and use the same methodology against these also when they appear. That will hopefully provide some sort of bench mark when it comes top model strength.

I hope this series of tests has been interesting and helpful to you and I’d love to hear your thoughts on the results.

Comparing LLMs in Copilot services–Round 6

MAI_cf947564d0f53b48

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

Round 5  – https://blog.ciaops.com/2026/08/14/comparing-llms-in-copilot-services-round-5-cowork/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

Next up, I’ve pitted Opus 5 and Researcher Critique. The results are:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260814-Deep-Grok-eval.md

1. Opus 5

2. Researcher Critique

the gap between these two was quite large again. The clear winner therefore is:

Opus 5

This mean the upcoming final is going to be Opus 5 (Coworok) vs Opus (Chat) for thr crown which will be very interesting given they are the same model, while one is an included LLM and the other is PAYG!

Stay tuned for the action of the first CIAOPS LLM final!

What Happens If All LLMs Become the Same?

image

I’ve been thinking about a question that sounds a bit strange at first: what happens if all the major language models start converging into one common capability?

Not literally one company. Not one product. Not one button.

I mean something more subtle. What if the difference between the big models becomes less obvious to the average business user? What if the answer from one model is good enough, the answer from another is also good enough, and the real distinction is no longer the model itself but where it lives, what it can access, what it can do, and how safely it can do it?

That is a very different world from the one a lot of people are still arguing about.

The model may become the least interesting part

Right now, there is still a lot of energy around model comparison. Which one writes better? Which one reasons better? Which one codes better? Which one is cheaper? Which one won the latest benchmark?

That matters, but I’m not convinced it will matter in the same way for most organisations.

For many business users, the model is already beginning to disappear into the workflow. They don’t want to pick between ten engines before replying to an email. They want to open Outlook, ask Copilot to summarise the thread, draft a response, and make sure it reflects the real conversation. They want to sit in Teams, catch up on a meeting, identify the unresolved decisions, and move on.

If every serious model gets broadly competent at writing, reasoning, summarising, analysing, and planning, then the contest shifts. The question becomes less “which LLM is smartest?” and more “which environment gives this model the right context, guardrails, and business action?”

That is where Microsoft 365 starts to matter.

A generic super LLM might know a lot about the world. But it does not automatically know your SharePoint structure, your Teams conversations, your Outlook history, your policies, your client files, your permissions, or your business rhythm. And if it does get access to those things, the real issue becomes governance.

Common intelligence makes business discipline more important

If AI capability becomes common, then competitive advantage moves somewhere else.

It moves to your data quality.
It moves to your process maturity.
It moves to your permission model.
It moves to your ability to describe the outcome you actually want.

That is uncomfortable for many businesses because it means AI does not magically fix operational mess. It exposes it.

If your documents are scattered across personal OneDrives, old Teams channels, duplicated SharePoint libraries, and mystery folders called “Final Final Real Final”, a better model may not save you. It may simply find the wrong thing faster.

This is why I keep coming back to the practical layer. Before worrying about whether the world ends up with one dominant super LLM, I’d rather ask whether your organisation has clean source material, sensible access controls, repeatable workflows, and people who know how to challenge the output.

Ask Copilot in Word to draft a client-ready explanation from a properly maintained policy document, and you start to see real value. Ask it to work from five conflicting policy drafts and a half-forgotten email thread, and you get a polished problem.

AI does not remove responsibility. It compresses the time between messy input and messy output.

The future may be less about models and more about orchestration

My guess is that we won’t care as much about individual model names over time. We’ll care about orchestration.

Which model should handle this task?
Which data should it use?
Which actions is it allowed to take?
Which human signs off?
Which audit trail remains?

That is the operating model businesses need to build. Not a fan club for a particular LLM.

If all roads eventually lead to a broadly common intelligence layer, the winners will not be the organisations that simply had access to it. Everyone will. The winners will be the ones that wrapped that intelligence in good process, clean data, sensible governance, and practical human judgement.

The super LLM, if it arrives, may not be the finish line.

It may just be the new baseline.

And once everyone has the same baseline, the old boring things start to matter again: discipline, clarity, trust, and execution.

Comparing LLMs in Copilot services–Round 5 – Cowork

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

For this round I’ve pitted the Cowork Claude winner (Opus 5) vs the Cowork GPT winner (5.6 Terra). The result is:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-Grok-eval.md

1. Opus 5.0

2. GPT 5.6 Terra

and this time the gap was much larger than before. Not unexpected given that Terra is designed as a lighter weight, less powerful model. However, don’t forget the cost factor which isn’t included in these evaluations:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

If costs were indeed taken into account it would certain swing the pendulum back toward Terra.

However, this time the winner is clear

Opus 5

moves onto the next round.

Comparing LLMs in Copilot services–Round 4 – Cowork (GPT)

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

This round the battle is between all the models available in Copilot Cowork GPT. Namely, these:

Screenshot 2026-08-12 082720

The other interesting factor here is that all these models as PAYG, so I’ll also give the costs for the same prompt.

I used Grok to evaluate all the results which produced:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-GPT-Grok-eval.md

In summary, the rankings are:

1. 5.6 Terra

2. 5.6 Sol

3. 5.5

Interestingly, the less powerful model (Terra) produced a better result here.

And the costs of each:


GPT 5.6 Sol = 2,800.50 credits
GPT 5.6 Terra = 260 credits
GPT 5.5 = 1,200 credits

5.6 Terra wins again big here, only costing 260 credits! This is how the ranked costs table so far looks, from most to least expensive:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

Given that Terra is so much cheaper I’ll need to spend some more time investigating and verifying that with other requests, given that these tests are simply a one shot prompt to result.

Drum roll. The clear winner for this round is:

GPT 5.6 Terra

With all the preliminaries done we can now get onto comparing the winners of each round together.

Comparing LLMs in Copilot services–Round 3 – Cowork (Claude)

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

This round the battle is between all the models available in Copilot Cowork Claude. Namely, these:

Screenshot 2026-08-11 074851

The other interesting factor here is that all these models as PAYG, so I’ll also give the costs for the same prompt.

To assess all these results together I found Gemini and SharePoint Copilot to both fail completely. Gemini keeps saying that it can’t find all the document, even though I have uploaded and also tried linking. SharePoint Copilot on the other hand starts processing but never gives me a result, no matter how long I wait. I will admit that these files are quite long (20+ pages) and quite complex, so there is a lot to digest. That said I threw them into Grok and those  results are here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-Claude-Grok-eval.md

In summary, the ranking where:

1. Opus 5

2. Fable 5 (Preview)

3. Fable 5 (Copilot) (Preview)

4. Opus 4.8

5. Sonnet 5

which all kind of makes empirical sense but still quite subjective I feel. However, for now I’ll swap to using Grok to evaluate these documents as a standard approach.

Now for the costs which don’t factor into the results:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits


Which, on initial analysis, indicates that costs for all models are in the same kind of range (i.e. at least US$20 per prompt), with Fable 5 models being about 50% more expensive.

Given the analysis and costs, it seems Opus 5 with Cowork, if you are using Claude, is the most cost effective for the best result in Cowork.

So:,

Round 3 winner – Copilot Cowork Claude = Opus 5

Onto the next round.

Comparing LLMs in Copilot services–Round 1 – Chat

image

Inspired by the recent soccer world cup, I have decided to create a Copilot LLM output comparison challenge.

The plan is to use the same prompt with all examples of different services in Copilot and then different models available in each service. After that, the idea is them to compare the winner of each round to determine the overall winner and to continue to do this on a regular basis as new models and services are added over time.

Thus, the methodology is to use the same complex prompt to generate the result from the model (a report) and then use a standard prompt to evaluate all the results to determine a winner. The easiest comparison method is to use Copilot in SharePoint but the aim will also to be to compare using other models as well. 

So, for round 1 I’m going to compare all the models available in Copilot chat. Comparison generated by Copilot for SharePoint.

Rather than try and fit the reports here I will upload them to my Github repository here in markdown format:

https://github.com/directorcia/general/tree/master/Copilot/Comparisons

Comparative Assessment – Live Writer Paste

This first report is now directly available at:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260630-Chat.md

The results where (out of 10):

1. Opus – 9.51

2. Sonnet – 9.45

3. GPT 5.6 Thinking – 8.76

4. GPT 5.5 Quick – 7.69

5. Auto – 6.93

So, the winner for Round 1 – Copilot Chat = Opus.

Onto Round 2


LLMs Are Grown, Not Coded – And That Changes Everything

image

One of the biggest misunderstandings I still see in the market is the idea that large language models are “just software”. That they’re something you build, configure, and control in the same way you do an application, a script, or even a PowerShell module.

They’re not.

LLMs are not coded in the traditional sense. They are grown.

And once you understand that distinction, a lot of confusion around AI, risk, accuracy, and expectations suddenly makes sense.

Code Is Deterministic. LLMs Are Probabilistic.

Traditional software works because we tell it exactly what to do.

If this happens, do that.
If the value equals X, return Y.
If the script runs twice with the same inputs, you expect the same outputs.

LLMs don’t work like that.

They are trained on vast amounts of data and learn patterns, relationships, and probabilities. When you prompt an LLM, it isn’t “executing logic”. It is calculating the most likely next token based on everything it has seen before.

That’s not coding.
That’s cultivation.

Think of an LLM less like a calculator and more like a very well‑read human who answers based on experience, context, and probability. Sometimes they’re brilliant. Sometimes they’re confidently wrong. And sometimes they surprise you with insights you didn’t expect.

You Don’t Compile an LLM – You Train It

When we write code, we compile it. When there’s a bug, we fix the line of code and re‑run it.

With LLMs, you don’t fix bugs in the same way.

You:

  • Change the training data

  • Adjust the fine‑tuning

  • Improve the prompt context

  • Add guardrails

  • Supplement with retrieval (RAG)

  • Wrap it in agents, workflows, and policy

That’s why LLMs improve over time in jumps, not increments. A new model release isn’t a patch Tuesday update – it’s a new organism that has grown up on a bigger, cleaner, more structured diet.

This is also why the same prompt can give you slightly different answers on different days or across different models. You’re not calling a function. You’re having a conversation with a statistical engine.

Why This Matters for Business (and MSPs)

If you think LLMs are coded, you’ll expect certainty.

If you understand they’re grown, you’ll design for outcomes instead.

That means:

  • You validate outputs instead of blindly trusting them

  • You treat AI as an assistant, not an authority

  • You design processes that assume probabilistic answers

  • You put humans in the loop where it matters

  • You focus on reducing risk, not eliminating it (because you can’t)

This is exactly why raw “public AI” is dangerous in business contexts, and why platforms like Microsoft 365 Copilot matter. Copilot doesn’t magically make the LLM smarter – it feeds it better data, constrains its environment, applies identity, compliance, and security, and grounds responses in your organisation’s reality.

The model hasn’t changed. The nutrition has.

Prompts Are Fertiliser, Not Commands

Another symptom of the “coded mindset” is prompt obsession.

People ask for the perfect prompt as if it’s a magic incantation.

Prompts don’t control LLMs.
They nudge them.

A good prompt gives context, tone, constraints, and examples. A bad prompt starves the model and then complains about the output.

Again, this makes sense if you think in biological terms. You don’t shout instructions at a plant and expect it to grow differently overnight. You change the environment, the inputs, and the expectations.

Why AI Feels Uncomfortable to Traditional IT People

For those of us who grew up with servers, scripts, and systems that either worked or didn’t, LLMs are uncomfortable.

They live in the grey.

They’re not always right.
They’re not always wrong.
They’re useful far more often than they’re perfect.

And that’s the mental shift required.

The organisations that win with AI won’t be the ones who treat it like another application to deploy. They’ll be the ones who treat it like a junior staff member that:

  • Needs good information

  • Needs supervision

  • Improves with feedback

  • Gets more useful the more you work with it
The Bottom Line

LLMs aren’t coded.
They’re grown.

If you try to manage them like software, you’ll be frustrated. If you treat them like a system that learns, adapts, and responds to its environment, you’ll unlock real value.

This is why AI strategy isn’t about models. It’s about data, context, governance, and outcomes.

And it’s why the real competitive advantage won’t come from “which AI you use”, but from how well you grow it inside your business.

If you’re still treating AI like a tool, you’re already behind.

If you’re treating it like a capability, you’re finally asking the right questions.