Comparing LLMs in Copilot services–Round 4 – Cowork (GPT)

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

This round the battle is between all the models available in Copilot Cowork GPT. Namely, these:

Screenshot 2026-08-12 082720

The other interesting factor here is that all these models as PAYG, so I’ll also give the costs for the same prompt.

I used Grok to evaluate all the results which produced:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-GPT-Grok-eval.md

In summary, the rankings are:

1. 5.6 Terra

2. 5.6 Sol

3. 5.5

Interestingly, the less powerful model (Terra) produced a better result here.

And the costs of each:


GPT 5.6 Sol = 2,800.50 credits
GPT 5.6 Terra = 260 credits
GPT 5.5 = 1,200 credits

5.6 Terra wins again big here, only costing 260 credits! This is how the ranked costs table so far looks, from most to least expensive:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

Given that Terra is so much cheaper I’ll need to spend some more time investigating and verifying that with other requests, given that these tests are simply a one shot prompt to result.

Drum roll. The clear winner for this round is:

GPT 5.6 Terra

With all the preliminaries done we can now get onto comparing the winners of each round together.

Comparing LLMs in Copilot services–Round 3 – Cowork (Claude)

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

This round the battle is between all the models available in Copilot Cowork Claude. Namely, these:

Screenshot 2026-08-11 074851

The other interesting factor here is that all these models as PAYG, so I’ll also give the costs for the same prompt.

To assess all these results together I found Gemini and SharePoint Copilot to both fail completely. Gemini keeps saying that it can’t find all the document, even though I have uploaded and also tried linking. SharePoint Copilot on the other hand starts processing but never gives me a result, no matter how long I wait. I will admit that these files are quite long (20+ pages) and quite complex, so there is a lot to digest. That said I threw them into Grok and those  results are here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-Claude-Grok-eval.md

In summary, the ranking where:

1. Opus 5

2. Fable 5 (Preview)

3. Fable 5 (Copilot) (Preview)

4. Opus 4.8

5. Sonnet 5

which all kind of makes empirical sense but still quite subjective I feel. However, for now I’ll swap to using Grok to evaluate these documents as a standard approach.

Now for the costs which don’t factor into the results:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits


Which, on initial analysis, indicates that costs for all models are in the same kind of range (i.e. at least US$20 per prompt), with Fable 5 models being about 50% more expensive.

Given the analysis and costs, it seems Opus 5 with Cowork, if you are using Claude, is the most cost effective for the best result in Cowork.

So:,

Round 3 winner – Copilot Cowork Claude = Opus 5

Onto the next round.

Comparing LLMs in Copilot services–Round 1 – Chat

image

Inspired by the recent soccer world cup, I have decided to create a Copilot LLM output comparison challenge.

The plan is to use the same prompt with all examples of different services in Copilot and then different models available in each service. After that, the idea is them to compare the winner of each round to determine the overall winner and to continue to do this on a regular basis as new models and services are added over time.

Thus, the methodology is to use the same complex prompt to generate the result from the model (a report) and then use a standard prompt to evaluate all the results to determine a winner. The easiest comparison method is to use Copilot in SharePoint but the aim will also to be to compare using other models as well. 

So, for round 1 I’m going to compare all the models available in Copilot chat. Comparison generated by Copilot for SharePoint.

Rather than try and fit the reports here I will upload them to my Github repository here in markdown format:

https://github.com/directorcia/general/tree/master/Copilot/Comparisons

Comparative Assessment – Live Writer Paste

This first report is now directly available at:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260630-Chat.md

The results where (out of 10):

1. Opus – 9.51

2. Sonnet – 9.45

3. GPT 5.6 Thinking – 8.76

4. GPT 5.5 Quick – 7.69

5. Auto – 6.93

So, the winner for Round 1 – Copilot Chat = Opus.

Onto Round 2


LLMs Are Grown, Not Coded – And That Changes Everything

image

One of the biggest misunderstandings I still see in the market is the idea that large language models are “just software”. That they’re something you build, configure, and control in the same way you do an application, a script, or even a PowerShell module.

They’re not.

LLMs are not coded in the traditional sense. They are grown.

And once you understand that distinction, a lot of confusion around AI, risk, accuracy, and expectations suddenly makes sense.

Code Is Deterministic. LLMs Are Probabilistic.

Traditional software works because we tell it exactly what to do.

If this happens, do that.
If the value equals X, return Y.
If the script runs twice with the same inputs, you expect the same outputs.

LLMs don’t work like that.

They are trained on vast amounts of data and learn patterns, relationships, and probabilities. When you prompt an LLM, it isn’t “executing logic”. It is calculating the most likely next token based on everything it has seen before.

That’s not coding.
That’s cultivation.

Think of an LLM less like a calculator and more like a very well‑read human who answers based on experience, context, and probability. Sometimes they’re brilliant. Sometimes they’re confidently wrong. And sometimes they surprise you with insights you didn’t expect.

You Don’t Compile an LLM – You Train It

When we write code, we compile it. When there’s a bug, we fix the line of code and re‑run it.

With LLMs, you don’t fix bugs in the same way.

You:

  • Change the training data

  • Adjust the fine‑tuning

  • Improve the prompt context

  • Add guardrails

  • Supplement with retrieval (RAG)

  • Wrap it in agents, workflows, and policy

That’s why LLMs improve over time in jumps, not increments. A new model release isn’t a patch Tuesday update – it’s a new organism that has grown up on a bigger, cleaner, more structured diet.

This is also why the same prompt can give you slightly different answers on different days or across different models. You’re not calling a function. You’re having a conversation with a statistical engine.

Why This Matters for Business (and MSPs)

If you think LLMs are coded, you’ll expect certainty.

If you understand they’re grown, you’ll design for outcomes instead.

That means:

  • You validate outputs instead of blindly trusting them

  • You treat AI as an assistant, not an authority

  • You design processes that assume probabilistic answers

  • You put humans in the loop where it matters

  • You focus on reducing risk, not eliminating it (because you can’t)

This is exactly why raw “public AI” is dangerous in business contexts, and why platforms like Microsoft 365 Copilot matter. Copilot doesn’t magically make the LLM smarter – it feeds it better data, constrains its environment, applies identity, compliance, and security, and grounds responses in your organisation’s reality.

The model hasn’t changed. The nutrition has.

Prompts Are Fertiliser, Not Commands

Another symptom of the “coded mindset” is prompt obsession.

People ask for the perfect prompt as if it’s a magic incantation.

Prompts don’t control LLMs.
They nudge them.

A good prompt gives context, tone, constraints, and examples. A bad prompt starves the model and then complains about the output.

Again, this makes sense if you think in biological terms. You don’t shout instructions at a plant and expect it to grow differently overnight. You change the environment, the inputs, and the expectations.

Why AI Feels Uncomfortable to Traditional IT People

For those of us who grew up with servers, scripts, and systems that either worked or didn’t, LLMs are uncomfortable.

They live in the grey.

They’re not always right.
They’re not always wrong.
They’re useful far more often than they’re perfect.

And that’s the mental shift required.

The organisations that win with AI won’t be the ones who treat it like another application to deploy. They’ll be the ones who treat it like a junior staff member that:

  • Needs good information

  • Needs supervision

  • Improves with feedback

  • Gets more useful the more you work with it
The Bottom Line

LLMs aren’t coded.
They’re grown.

If you try to manage them like software, you’ll be frustrated. If you treat them like a system that learns, adapts, and responds to its environment, you’ll unlock real value.

This is why AI strategy isn’t about models. It’s about data, context, governance, and outcomes.

And it’s why the real competitive advantage won’t come from “which AI you use”, but from how well you grow it inside your business.

If you’re still treating AI like a tool, you’re already behind.

If you’re treating it like a capability, you’re finally asking the right questions.