New Microsoft image model

I have a standard image prompt that I use to test Ai models.

Previous iterations:

MAI-Image-2.5-Pro – https://blog.ciaops.com/2026/07/24/new-microsoft-image-model/

MAI-Image-2.5-Flash and MAI-Image-2.5 is here:

https://blog.ciaops.com/2026/06/04/latest-microsoft-image-models/

Before that with MAI-Image-1.5 and Flux.2 Flex is here:

https://blog.ciaops.com/2026/05/16/copilot-image-generation-in-powerpoint/

the previous attempts:

https://blog.ciaops.com/2026/05/05/revisiting-copilot-image-generation-analysis/

and the first attempt:

https://blog.ciaops.com/2026/03/07/image-generation-analysis/

Microsoft has just released a new models and here is what I got when I used them:

MAI-Image-2.6

MAI_0b7d8ab63529c641

Read more about this model here:

https://microsoft.ai/news/mai-image-2-6-launches-at-no-2-on-arena-ahead-of-google-meta-and-xai/

What Happens If All LLMs Become the Same?

image

I’ve been thinking about a question that sounds a bit strange at first: what happens if all the major language models start converging into one common capability?

Not literally one company. Not one product. Not one button.

I mean something more subtle. What if the difference between the big models becomes less obvious to the average business user? What if the answer from one model is good enough, the answer from another is also good enough, and the real distinction is no longer the model itself but where it lives, what it can access, what it can do, and how safely it can do it?

That is a very different world from the one a lot of people are still arguing about.

The model may become the least interesting part

Right now, there is still a lot of energy around model comparison. Which one writes better? Which one reasons better? Which one codes better? Which one is cheaper? Which one won the latest benchmark?

That matters, but I’m not convinced it will matter in the same way for most organisations.

For many business users, the model is already beginning to disappear into the workflow. They don’t want to pick between ten engines before replying to an email. They want to open Outlook, ask Copilot to summarise the thread, draft a response, and make sure it reflects the real conversation. They want to sit in Teams, catch up on a meeting, identify the unresolved decisions, and move on.

If every serious model gets broadly competent at writing, reasoning, summarising, analysing, and planning, then the contest shifts. The question becomes less “which LLM is smartest?” and more “which environment gives this model the right context, guardrails, and business action?”

That is where Microsoft 365 starts to matter.

A generic super LLM might know a lot about the world. But it does not automatically know your SharePoint structure, your Teams conversations, your Outlook history, your policies, your client files, your permissions, or your business rhythm. And if it does get access to those things, the real issue becomes governance.

Common intelligence makes business discipline more important

If AI capability becomes common, then competitive advantage moves somewhere else.

It moves to your data quality.
It moves to your process maturity.
It moves to your permission model.
It moves to your ability to describe the outcome you actually want.

That is uncomfortable for many businesses because it means AI does not magically fix operational mess. It exposes it.

If your documents are scattered across personal OneDrives, old Teams channels, duplicated SharePoint libraries, and mystery folders called “Final Final Real Final”, a better model may not save you. It may simply find the wrong thing faster.

This is why I keep coming back to the practical layer. Before worrying about whether the world ends up with one dominant super LLM, I’d rather ask whether your organisation has clean source material, sensible access controls, repeatable workflows, and people who know how to challenge the output.

Ask Copilot in Word to draft a client-ready explanation from a properly maintained policy document, and you start to see real value. Ask it to work from five conflicting policy drafts and a half-forgotten email thread, and you get a polished problem.

AI does not remove responsibility. It compresses the time between messy input and messy output.

The future may be less about models and more about orchestration

My guess is that we won’t care as much about individual model names over time. We’ll care about orchestration.

Which model should handle this task?
Which data should it use?
Which actions is it allowed to take?
Which human signs off?
Which audit trail remains?

That is the operating model businesses need to build. Not a fan club for a particular LLM.

If all roads eventually lead to a broadly common intelligence layer, the winners will not be the organisations that simply had access to it. Everyone will. The winners will be the ones that wrapped that intelligence in good process, clean data, sensible governance, and practical human judgement.

The super LLM, if it arrives, may not be the finish line.

It may just be the new baseline.

And once everyone has the same baseline, the old boring things start to matter again: discipline, clarity, trust, and execution.

Comparing LLMs in Copilot services–Round 5 – Cowork

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

For this round I’ve pitted the Cowork Claude winner (Opus 5) vs the Cowork GPT winner (5.6 Terra). The result is:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-Grok-eval.md

1. Opus 5.0

2. GPT 5.6 Terra

and this time the gap was much larger than before. Not unexpected given that Terra is designed as a lighter weight, less powerful model. However, don’t forget the cost factor which isn’t included in these evaluations:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

If costs were indeed taken into account it would certain swing the pendulum back toward Terra.

However, this time the winner is clear

Opus 5

moves onto the next round.

Comparing LLMs in Copilot services–Round 4 – Cowork (GPT)

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

This round the battle is between all the models available in Copilot Cowork GPT. Namely, these:

Screenshot 2026-08-12 082720

The other interesting factor here is that all these models as PAYG, so I’ll also give the costs for the same prompt.

I used Grok to evaluate all the results which produced:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-GPT-Grok-eval.md

In summary, the rankings are:

1. 5.6 Terra

2. 5.6 Sol

3. 5.5

Interestingly, the less powerful model (Terra) produced a better result here.

And the costs of each:


GPT 5.6 Sol = 2,800.50 credits
GPT 5.6 Terra = 260 credits
GPT 5.5 = 1,200 credits

5.6 Terra wins again big here, only costing 260 credits! This is how the ranked costs table so far looks, from most to least expensive:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

Given that Terra is so much cheaper I’ll need to spend some more time investigating and verifying that with other requests, given that these tests are simply a one shot prompt to result.

Drum roll. The clear winner for this round is:

GPT 5.6 Terra

With all the preliminaries done we can now get onto comparing the winners of each round together.

Why AI Charges More to Write Than to Read

image

I keep seeing people look at AI pricing pages and assume someone has made a typo.

Input tokens cost one amount. Output tokens cost another. Cached input may cost less again. At first glance that feels odd. A token is a token, surely?

Not really.

The easiest way to think about it is this: reading is cheaper than writing. AI can read a lot of your prompt in parallel. It can process the instructions, pasted document, previous conversation and system rules in large chunks. That is the input side.

Writing the answer is different. The model generates the response one token at a time. Each next word, fragment, line of code or JSON field depends on what came before it. The longer the response, the longer expensive compute is tied up producing it.

That is why output tokens usually cost more. You are not just paying for text. You are paying for generation.

This matters more than most people think

Once you move into agents, automation, Copilot Studio, Azure OpenAI, GitHub Copilot, or anything that runs repeatedly, the economics change quickly.

A short prompt that produces a long report can cost more than a large prompt that produces a tiny answer. That surprises people. They assume the big document is the expensive part. Sometimes it is not. The real cost can be the polished, verbose output they asked the model to create.

Ask an AI system to read a SharePoint policy library and return three risks, and you are probably dealing with an input-heavy, output-light workload. That can be relatively efficient.

Ask it to create a 25-page report, an executive summary, a remediation plan, a Teams post, a client email and a formatted table every time it runs, and you have created an output-heavy workload. That is where the bill starts to move.

The MSP lesson is simple: design the output

This is where MSPs need to stop treating AI as magic and start treating it as infrastructure.

When we built servers, we cared about CPU, RAM, disk and backup windows. With AI, we need to care about prompts, context, output length, caching and repeatability.

The bad habit is asking for everything every time. “Give me the full report.” “Include all the detail.” “Make it comprehensive.” That sounds harmless until the same workflow runs fifty times across fifty tenants.

A better approach is to be deliberate. Ask for the smallest useful output first. Use summaries where summaries are enough. Generate detailed reports only when there is a reason. Reuse stable instructions and context where caching is available. Put spending limits around anything consumption-based. Review what the agent writes, not just what it reads.

Inside Microsoft 365, this means being clear about the difference between ordinary Copilot use in Word, Excel, Outlook or Teams and consumption-based AI work that may be billed differently. A user drafting an email is one thing. An agent chewing through documents and producing long artefacts all day is another.

This is not a pricing trick

I do not see the input/output price split as some mysterious vendor tax. It reflects how the technology behaves.

The mistake is pretending it does not matter.

AI costs are not just about how many people have a licence. They are about what those people, agents and workflows ask the model to produce. The output is where the hidden weight often sits.

So the practical rule is this: do not just prompt for the result. Design the cost shape of the result.

That might be the difference between AI being a useful business tool and AI becoming the next cloud bill nobody wants to open.

Comparing LLMs in Copilot services–Round 3 – Cowork (Claude)

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

This round the battle is between all the models available in Copilot Cowork Claude. Namely, these:

Screenshot 2026-08-11 074851

The other interesting factor here is that all these models as PAYG, so I’ll also give the costs for the same prompt.

To assess all these results together I found Gemini and SharePoint Copilot to both fail completely. Gemini keeps saying that it can’t find all the document, even though I have uploaded and also tried linking. SharePoint Copilot on the other hand starts processing but never gives me a result, no matter how long I wait. I will admit that these files are quite long (20+ pages) and quite complex, so there is a lot to digest. That said I threw them into Grok and those  results are here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-Claude-Grok-eval.md

In summary, the ranking where:

1. Opus 5

2. Fable 5 (Preview)

3. Fable 5 (Copilot) (Preview)

4. Opus 4.8

5. Sonnet 5

which all kind of makes empirical sense but still quite subjective I feel. However, for now I’ll swap to using Grok to evaluate these documents as a standard approach.

Now for the costs which don’t factor into the results:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits


Which, on initial analysis, indicates that costs for all models are in the same kind of range (i.e. at least US$20 per prompt), with Fable 5 models being about 50% more expensive.

Given the analysis and costs, it seems Opus 5 with Cowork, if you are using Claude, is the most cost effective for the best result in Cowork.

So:,

Round 3 winner – Copilot Cowork Claude = Opus 5

Onto the next round.

Is Copilot Listening to Every Meeting? The Reality Is More Complicated

image

Every few weeks someone asks me a variation of the same question.

“Is Microsoft Copilot recording all my meetings?”

The concern is understandable. AI is becoming more capable, meeting summaries appear almost instantly, action items seem to materialise out of nowhere, and people naturally wonder whether Microsoft is quietly capturing everything that happens in Teams. Sometimes the question goes a step further: “Does it do this even if I don’t have a Copilot licence?”

In my experience, the answer is far less dramatic than many people expect.

The concern isn’t really about Copilot. It’s about visibility.

The Difference Between a Meeting and a Recording

I think a lot of confusion comes from people bundling several different technologies together and calling them all “Copilot”.

A Teams meeting can exist without being recorded.

A Teams meeting can be transcribed without being recorded.

A Teams meeting can be recorded and transcribed.

And Copilot may or may not be involved at all.

That’s an important distinction.

Copilot doesn’t magically create information from nowhere. It works with the information available to it. If your organisation has enabled meeting transcription or recording, that’s where much of the content comes from. Copilot can then use that information to answer questions, create summaries, identify action items and help participants catch up.

Without that underlying data, Copilot has far less to work with.

The key point is that recording and transcription settings are administrative decisions. They aren’t automatically triggered simply because somebody owns a Copilot licence.

What About People Without Copilot?

This is where things often become misunderstood.

I’ve spoken with organisations where only a handful of staff have Microsoft 365 Copilot licences, yet meeting transcripts still exist across the business. When employees see AI-generated meeting summaries, the assumption is that Copilot must be capturing everything.

In reality, those organisations may have enabled Teams transcription or recording policies independently of Copilot.

Think about it this way. Your ability to access information and Microsoft’s ability to store information are not the same thing.

Someone without a Copilot licence may still participate in a meeting that is being recorded or transcribed. The meeting data exists because the meeting organiser or organisation allowed that functionality. Whether an individual user has access to Copilot features is a separate licensing decision.

That’s a subtle difference, but an important one.

The Governance Question Most Businesses Miss

What I find interesting is that many organisations ask whether Copilot is recording meetings, but they rarely ask who should have access to the resulting information.

That’s the more important conversation.

If meeting recordings are stored in OneDrive or SharePoint, if transcripts are available through Teams, and if Copilot can surface information that already exists, then governance matters more than fear.

I’ve seen organisations spend hours debating AI risks while leaving years of meeting recordings scattered across Microsoft 365 with inconsistent permissions.

Copilot didn’t create that problem. It simply made the problem more visible.

A useful exercise is to ask Copilot in Teams or Microsoft 365 a question about a recent meeting and then investigate why the answer was possible. The answer usually points back to existing recordings, transcripts, shared files or meeting notes that were already being retained.

AI acts as a spotlight. It doesn’t necessarily create new information.

The Question I’m Watching Closely

As Copilot adoption grows, I think we’ll see a shift in how organisations think about meetings.

For years, many businesses treated meetings as temporary conversations. Once the meeting ended, most people assumed the discussion disappeared into the ether unless someone took notes.

That assumption no longer holds true.

Whether you’re using Copilot, Teams transcription, meeting recordings or a combination of all three, conversations increasingly become searchable organisational knowledge.

The question therefore isn’t whether Copilot is secretly listening to every meeting.

The better question is whether your organisation understands what meeting data is being captured, where it is stored, how long it is retained and who can access it.

That’s where the real risk lives.

And, just as importantly, that’s where the real value of Microsoft 365 Copilot begins.

Need to Know podcast–Episode 369

In this episode of the Need to Know Podcast, I cover recent Microsoft news, including strong financial results, continued Azure growth, rising Microsoft 365 Copilot adoption, and major security updates. Key security stories include supply chain compromise, cybercrime disruption, hotel Wi-Fi credential theft, and Project Perception, which points toward a future where security is increasingly managed by AI agents rather than manual review. The episode also highlights new Microsoft 365 Copilot capabilities, web grounding domain controls, Copilot in SharePoint improvements, plus two new CIAOPS simulators for Exchange/Defender policy flow and Microsoft 365 sign-in/Conditional Access testing.

The broader discussion focuses on how AI is changing software, security, and business productivity. I argue that cheaper open-weight AI models, Microsoft’s MAI models, and tools like GitHub Copilot make it easier for businesses to build their own dashboards, simulators, and lightweight applications instead of relying only on traditional software. He also introduces the idea of a “SharePoint gardener” to keep SharePoint information organised for AI use, and explains how AI loops can continuously test, improve, and refine code or business processes with human oversight where needed.

Brought to you by www.ciaopspatron.com

you can listen directly to this episode at:

https://ciaops.podbean.com/e/episode-369-developing/

Subscribe via iTunes at:

https://itunes.apple.com/au/podcast/ciaops-need-to-know-podcasts/id406891445?mt=2

or Spotify:

https://open.spotify.com/show/7ejj00cOuw8977GnnE2lPb

Don’t forget to give the show a rating as well as send me any feedback or suggestions you may have for the show

Resources

CIAOPS Need to Know podcast – CIAOPS – Need to Know podcasts | CIAOPS

X – https://www.twitter.com/directorcia

director@ciaops.com

CIAOPS Blog

Join my Teams Shared Channel – CIAOPS

CIAOPS Merch store – CIAOPS

Become a CIAOPS Patron

CIAOPS AI Dojo

CIAOPS weekly news update – CIA Brief – CIAOPS

CIAOPS Labs – The Special Activities Division of the CIAOPS

Support CIAOPS

Get your M365 questions answered via email

Join my email list

A special thanks to the CIAOPS Patron community for making this podcast possible. You can find the benefits of a subscription to the community and become a member at https://www.ciaopspatron.com

Security & Threat Intelligence

Microsoft 365 Copilot & AI

SharePoint & Content Management

Microsoft Corporate News & Strategy

CIAOPS:

Exchange Online + Defender for Office 365 Policy Flow Simulator

M365 Sign-In and Conditional Access Flow Simulator