Comparing LLMs in Copilot services–Round 7-Final

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

Round 5  – https://blog.ciaops.com/2026/08/14/comparing-llms-in-copilot-services-round-5-cowork/

Winner – Opus 5

Round 6 – https://blog.ciaops.com/2026/08/16/comparing-llms-in-copilot-services-round-6/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

For the final, we have Opus 5 (Cowork) and Opus (Chat). The results are:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-Copilot-Grok-eval.md

1. Opus 5 (Cowork)

2. Opus (Chat)

Therefore, the initial over all Copilot LLM winner is a clear win for:

Opus 5 (Cowork)

The trade off is that the cost for this US$22, while the runner up’s price was included in the cost of a Microsoft 365 Copilot license.

These test have revealed a few things, in my opinion:

A. Claude Opus is the superior model

B. Opus in chat is almost as good as Opus Cowork but much cheaper

C. For most work Opus in chat is probably you best option

The challenge with these report, as with any LLM’s is they are not definitive. Another round could reveal completely different results as could the method of evaluation, however as a best effort I think it does have validity and provide some general findings that do help understand the model choices in Copilot.

For comparison, I have create a document parameters page here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-File-paramters.md

so you can now compare all the various output document sizes, pages, etc to each other. I have also made available a copy of each model output document, without any changes to it, so you can look at each for yourself. You will find them all at:

https://github.com/directorcia/general/tree/master/Copilot/Comparisons/20260816

The idea is to wait until we see new models appear in Copilot and use the same methodology against these also when they appear. That will hopefully provide some sort of bench mark when it comes top model strength.

I hope this series of tests has been interesting and helpful to you and I’d love to hear your thoughts on the results.

Measure What Moved

image

I have sat through enough business debriefs to know the pattern.

Someone explains how busy they were. The late nights. The campaign tweaks. The tools tested. The sheer amount of motion involved.

Then I ask the only question that really matters.

What changed?

Activity feels comforting because it proves something happened. It gives people something to report. It fills meetings, updates, and weekly summaries. But activity is not progress. You can have hardworking people and still be drifting sideways.

Effort is not the scorecard

I am not dismissing hard work. Effort matters when it is pointed at the right target. The problem is when effort becomes the defence for poor results.

A marketing push that produces almost no qualified leads is not a success because the team spent days on it. A support process that burns hours but leaves customers waiting is not working because people are “doing their best”. A sales pipeline full of conversations but no movement is not healthy because everyone is busy.

That is why I keep coming back to measurement. Not because I love dashboards for the sake of dashboards, but because measurement forces honesty. It removes the storytelling that creeps into business discussions. When the number is sitting there in front of everyone, the conversation changes.

The question shifts from “who tried hard?” to “what actually improved?”

Make the work visible

One of the biggest mistakes I see businesses make is leaving performance hidden in private inboxes, personal spreadsheets, and half-remembered conversations.

If sales numbers live in one person’s Excel file, the business does not really have a sales view. If customer follow-ups are buried in Outlook, the business does not really have a follow-up process. If project blockers only surface during a meeting once a week, the business is accepting delay as normal.

Microsoft 365 gives you ways to bring that work into the open. Put the shared tracker in SharePoint. Pin it in Teams. Use Planner for ownership. Use Excel to track the result, then ask Copilot in Excel to identify trends, gaps, and outliers. Use Copilot in Teams after the meeting to summarise decisions and actions, then compare those actions against the numbers next time.

That is not more admin. It is fewer hiding places.

What gets reviewed gets improved

The real discipline is not building the dashboard. Anyone can throw together a colourful report and feel productive for an afternoon. The discipline is reviewing it consistently, asking uncomfortable questions, and changing behaviour because of what it shows.

If a campaign is not producing leads, stop admiring the effort and fix the offer, the audience, or the follow-up. If service tickets keep backing up, stop saying the team is flat out and find the bottleneck. If Copilot is being rolled out, do not just count licences. Measure whether proposal turnaround, meeting follow-up, reporting quality, or response times are improving.

That is where Copilot becomes useful. Not as another shiny thing to justify, but as a way to reduce the drag between seeing a problem and doing something about it. Summarise the data. Draft the follow-up. Build the first version of the report. Help the team inspect the work faster.

But the human still has to care about the outcome.

The blunt test

Do not tell me how busy you were. Show me what moved.

If the number improved, understand why and repeat it. If it did not, stop decorating failure with effort and make a better decision.

A good business does not reward invisible busyness. It rewards useful progress.

That only happens when the work is visible, the numbers are reviewed, and people are honest enough to act on what they see.

Comparing LLMs in Copilot services–Round 6

MAI_cf947564d0f53b48

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

Round 5  – https://blog.ciaops.com/2026/08/14/comparing-llms-in-copilot-services-round-5-cowork/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

Next up, I’ve pitted Opus 5 and Researcher Critique. The results are:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260814-Deep-Grok-eval.md

1. Opus 5

2. Researcher Critique

the gap between these two was quite large again. The clear winner therefore is:

Opus 5

This mean the upcoming final is going to be Opus 5 (Coworok) vs Opus (Chat) for thr crown which will be very interesting given they are the same model, while one is an included LLM and the other is PAYG!

Stay tuned for the action of the first CIAOPS LLM final!

Does Your AI Environment Need a Reset?

image

I used to rebuild my computer every so often.

Not because it was broken. Not because the hardware had failed. It was because, over time, the machine collected rubbish. Old drivers. Forgotten utilities. Startup items I no longer needed. Temporary files. Half-removed applications. Things that made sense at the time, but slowly turned a clean system into something sluggish and unpredictable.

I think we are heading toward the same problem with AI.

Not with the model itself. With the environment around it.

AI inherits your mess

When people talk about AI performance, they usually talk about the model. Bigger model. Newer model. Faster model. Better reasoning. More context.

That matters, but it is only part of the story.

Inside a business, Microsoft 365 Copilot works with the information, permissions, meetings, files, chats, and workflows already sitting in your tenant. If those foundations are messy, Copilot does not magically clean them. It works with what it finds.

If SharePoint is full of abandoned sites, Copilot has more noise to search through. If Teams channels contain years of half-finished conversations, it has more stale context to interpret. If OneDrive is full of duplicate drafts named “final-final-really-final.docx”, the AI is not the problem. The house is.

This is why I think an AI reset is becoming a real operational practice.

The reset is not a rebuild

I am not suggesting businesses wipe everything and start again. That would be ridiculous.

The better comparison is a deliberate clean-up cycle. A review of what AI can see, what it should use, what it should ignore, and what needs to be retired.

That means checking who has access to what. It means looking at old SharePoint sites, inactive Teams, forgotten guest accounts, broad sharing links, stale files, unused prompts, and agents nobody owns anymore.

It also means reviewing the human habits around AI.

Are people still using the same prompts they wrote six months ago? Are those prompts pointing at the right source material? Are users asking Copilot to summarise everything because they cannot be bothered choosing the right file? Are teams saving useful prompt patterns somewhere reusable, or are the good ones buried in private chat history?

That is cruft as well.

Performance is not just speed

When an old PC slowed down, the answer was often to remove junk and reduce the background load. With AI, performance is broader than speed.

A good AI environment gives better answers because the data is cleaner. It gives safer answers because permissions are tighter. It gives more useful answers because users know which sources to reference. It gives more predictable outcomes because prompts and workflows are standardised.

For example, if a user asks Copilot in Outlook to draft a client follow-up based on a meeting, the quality depends on more than Copilot. Was the meeting transcribed? Were the notes clear? Are the relevant project files in the right SharePoint location? Is the client information current? Has the user given Copilot a clear task, or just thrown a vague request at it?

The model may be smart. The environment still has to be tidy.

The monthly AI reset

I can see this becoming a monthly rhythm for MSPs and internal IT teams.

Review Copilot usage. Check new agents. Inspect oversharing. Look at sensitivity labels. Clean up old guests. Archive dead Teams. Refresh prompt libraries. Remove obsolete source documents. Confirm DLP policies still make sense. Ask whether AI is helping real work or just creating more output to manage.

That is not busywork. That is how you stop AI from drifting into the same state as an old Windows install full of utilities nobody remembers installing.

AI does not remove the need for operational discipline. It raises the price of not having it.

The organisations that get the most from AI will not be the ones constantly chasing the newest model. They will be the ones that keep their environment clean enough for AI to work with confidence.

Sometimes the best AI upgrade is not a new feature.

It is a clean-up.

New Microsoft image model

I have a standard image prompt that I use to test Ai models.

Previous iterations:

MAI-Image-2.5-Pro – https://blog.ciaops.com/2026/07/24/new-microsoft-image-model/

MAI-Image-2.5-Flash and MAI-Image-2.5 is here:

https://blog.ciaops.com/2026/06/04/latest-microsoft-image-models/

Before that with MAI-Image-1.5 and Flux.2 Flex is here:

https://blog.ciaops.com/2026/05/16/copilot-image-generation-in-powerpoint/

the previous attempts:

https://blog.ciaops.com/2026/05/05/revisiting-copilot-image-generation-analysis/

and the first attempt:

https://blog.ciaops.com/2026/03/07/image-generation-analysis/

Microsoft has just released a new models and here is what I got when I used them:

MAI-Image-2.6

MAI_0b7d8ab63529c641

Read more about this model here:

https://microsoft.ai/news/mai-image-2-6-launches-at-no-2-on-arena-ahead-of-google-meta-and-xai/

CIA Brief 20260815

image

Microsoft 365 & Copilot Productivity

  • What’s New in Copilot in SharePoint: August 2026

    Copilot in SharePoint can now turn a list, Excel file, or CSV into a live, interactive HTML dashboard that stays connected to the underlying data and refreshes each time it’s opened. The update also adds one-click “page buttons” that launch a saved Copilot prompt, plus chat improvements — shifting Copilot from simply answering questions to helping you build and act on content.

    https://techcommunity.microsoft.com/blog/spblog/whats-new-in-copilot-in-sharepoint-august-2026/4535421

  • What’s New in Excel (July 2026)

    The monthly Excel roundup is almost entirely about Copilot: new inline citations to verify AI responses, generally available synced connectors, and Power BI grounding that respects row-level security. Two frontier models — OpenAI’s GPT-5.6 and Anthropic’s Claude Opus 5 — are now selectable, and Copilot no longer requires AutoSave to be turned on.

    https://techcommunity.microsoft.com/blog/excelblog/whats-new-in-excel-july-2026/4523403

  • Link to a location in Word for Windows and Mac

    A new “Copy Link to Location” feature lets you highlight any text, right-click, and generate a shareable link that opens the document exactly at that spot — no heading, bookmark, or hyperlink required. It’s aimed at long documents and collaborative reviews, saving colleagues from scrolling to find the right section. Available in Word for Windows and Mac (links also open in the web).

    https://techcommunity.microsoft.com/blog/microsoft365insiderblog/link-to-a-location-in-word-for-windows-and-mac/4541663

Security & Threat Intelligence

AI & Agents

  • Building autonomous multi-agent workflows (AI Team)

    A video walkthrough shared in the AI Team on designing autonomous, multi-agent workflows — showing how agents can be chained to hand off and coordinate multi-step tasks rather than running as isolated, one-off bots. A useful primer for anyone moving toward production-grade agent automation in Copilot Studio.

    https://www.youtube.com/watch?v=UuJpNa_TbiI

After hours

The Obama Tan Suit | Season Finale of Life, Larry and the Pursuit of Unhappiness

https://www.youtube.com/watch?v=crNcV0k0aMQ

Editorial

If you found this valuable, the I’d appreciate a ‘like’ or perhaps a donation at https://ko-fi.com/ciaops. This helps me know that people enjoy what I have created and provides resources to allow me to create more content. If you have any feedback or suggestions around this, I’m all ears. You can also find me via email director@ciaops.com and on X (Twitter) at https://www.twitter.com/directorcia.

If you want to be part of a dedicated Microsoft Cloud community with information and interactions daily, then consider becoming a CIAOPS Patron – www.ciaopspatron.com.

Watch out for the next CIA Brief next week

What Happens If All LLMs Become the Same?

image

I’ve been thinking about a question that sounds a bit strange at first: what happens if all the major language models start converging into one common capability?

Not literally one company. Not one product. Not one button.

I mean something more subtle. What if the difference between the big models becomes less obvious to the average business user? What if the answer from one model is good enough, the answer from another is also good enough, and the real distinction is no longer the model itself but where it lives, what it can access, what it can do, and how safely it can do it?

That is a very different world from the one a lot of people are still arguing about.

The model may become the least interesting part

Right now, there is still a lot of energy around model comparison. Which one writes better? Which one reasons better? Which one codes better? Which one is cheaper? Which one won the latest benchmark?

That matters, but I’m not convinced it will matter in the same way for most organisations.

For many business users, the model is already beginning to disappear into the workflow. They don’t want to pick between ten engines before replying to an email. They want to open Outlook, ask Copilot to summarise the thread, draft a response, and make sure it reflects the real conversation. They want to sit in Teams, catch up on a meeting, identify the unresolved decisions, and move on.

If every serious model gets broadly competent at writing, reasoning, summarising, analysing, and planning, then the contest shifts. The question becomes less “which LLM is smartest?” and more “which environment gives this model the right context, guardrails, and business action?”

That is where Microsoft 365 starts to matter.

A generic super LLM might know a lot about the world. But it does not automatically know your SharePoint structure, your Teams conversations, your Outlook history, your policies, your client files, your permissions, or your business rhythm. And if it does get access to those things, the real issue becomes governance.

Common intelligence makes business discipline more important

If AI capability becomes common, then competitive advantage moves somewhere else.

It moves to your data quality.
It moves to your process maturity.
It moves to your permission model.
It moves to your ability to describe the outcome you actually want.

That is uncomfortable for many businesses because it means AI does not magically fix operational mess. It exposes it.

If your documents are scattered across personal OneDrives, old Teams channels, duplicated SharePoint libraries, and mystery folders called “Final Final Real Final”, a better model may not save you. It may simply find the wrong thing faster.

This is why I keep coming back to the practical layer. Before worrying about whether the world ends up with one dominant super LLM, I’d rather ask whether your organisation has clean source material, sensible access controls, repeatable workflows, and people who know how to challenge the output.

Ask Copilot in Word to draft a client-ready explanation from a properly maintained policy document, and you start to see real value. Ask it to work from five conflicting policy drafts and a half-forgotten email thread, and you get a polished problem.

AI does not remove responsibility. It compresses the time between messy input and messy output.

The future may be less about models and more about orchestration

My guess is that we won’t care as much about individual model names over time. We’ll care about orchestration.

Which model should handle this task?
Which data should it use?
Which actions is it allowed to take?
Which human signs off?
Which audit trail remains?

That is the operating model businesses need to build. Not a fan club for a particular LLM.

If all roads eventually lead to a broadly common intelligence layer, the winners will not be the organisations that simply had access to it. Everyone will. The winners will be the ones that wrapped that intelligence in good process, clean data, sensible governance, and practical human judgement.

The super LLM, if it arrives, may not be the finish line.

It may just be the new baseline.

And once everyone has the same baseline, the old boring things start to matter again: discipline, clarity, trust, and execution.

Comparing LLMs in Copilot services–Round 5 – Cowork

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

For this round I’ve pitted the Cowork Claude winner (Opus 5) vs the Cowork GPT winner (5.6 Terra). The result is:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-Grok-eval.md

1. Opus 5.0

2. GPT 5.6 Terra

and this time the gap was much larger than before. Not unexpected given that Terra is designed as a lighter weight, less powerful model. However, don’t forget the cost factor which isn’t included in these evaluations:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

If costs were indeed taken into account it would certain swing the pendulum back toward Terra.

However, this time the winner is clear

Opus 5

moves onto the next round.