New Microsoft image model

I have a standard image prompt that I use to test Ai models.

Previous iterations:

MAI-Image-2.5-Pro – https://blog.ciaops.com/2026/07/24/new-microsoft-image-model/

MAI-Image-2.5-Flash and MAI-Image-2.5 is here:

https://blog.ciaops.com/2026/06/04/latest-microsoft-image-models/

Before that with MAI-Image-1.5 and Flux.2 Flex is here:

https://blog.ciaops.com/2026/05/16/copilot-image-generation-in-powerpoint/

the previous attempts:

https://blog.ciaops.com/2026/05/05/revisiting-copilot-image-generation-analysis/

and the first attempt:

https://blog.ciaops.com/2026/03/07/image-generation-analysis/

Microsoft has just released a new models and here is what I got when I used them:

MAI-Image-2.6

MAI_0b7d8ab63529c641

Read more about this model here:

https://microsoft.ai/news/mai-image-2-6-launches-at-no-2-on-arena-ahead-of-google-meta-and-xai/

CIA Brief 20260815

image

Microsoft 365 & Copilot Productivity

  • What’s New in Copilot in SharePoint: August 2026

    Copilot in SharePoint can now turn a list, Excel file, or CSV into a live, interactive HTML dashboard that stays connected to the underlying data and refreshes each time it’s opened. The update also adds one-click “page buttons” that launch a saved Copilot prompt, plus chat improvements — shifting Copilot from simply answering questions to helping you build and act on content.

    https://techcommunity.microsoft.com/blog/spblog/whats-new-in-copilot-in-sharepoint-august-2026/4535421

  • What’s New in Excel (July 2026)

    The monthly Excel roundup is almost entirely about Copilot: new inline citations to verify AI responses, generally available synced connectors, and Power BI grounding that respects row-level security. Two frontier models — OpenAI’s GPT-5.6 and Anthropic’s Claude Opus 5 — are now selectable, and Copilot no longer requires AutoSave to be turned on.

    https://techcommunity.microsoft.com/blog/excelblog/whats-new-in-excel-july-2026/4523403

  • Link to a location in Word for Windows and Mac

    A new “Copy Link to Location” feature lets you highlight any text, right-click, and generate a shareable link that opens the document exactly at that spot — no heading, bookmark, or hyperlink required. It’s aimed at long documents and collaborative reviews, saving colleagues from scrolling to find the right section. Available in Word for Windows and Mac (links also open in the web).

    https://techcommunity.microsoft.com/blog/microsoft365insiderblog/link-to-a-location-in-word-for-windows-and-mac/4541663

Security & Threat Intelligence

AI & Agents

  • Building autonomous multi-agent workflows (AI Team)

    A video walkthrough shared in the AI Team on designing autonomous, multi-agent workflows — showing how agents can be chained to hand off and coordinate multi-step tasks rather than running as isolated, one-off bots. A useful primer for anyone moving toward production-grade agent automation in Copilot Studio.

    https://www.youtube.com/watch?v=UuJpNa_TbiI

After hours

The Obama Tan Suit | Season Finale of Life, Larry and the Pursuit of Unhappiness

https://www.youtube.com/watch?v=crNcV0k0aMQ

Editorial

If you found this valuable, the I’d appreciate a ‘like’ or perhaps a donation at https://ko-fi.com/ciaops. This helps me know that people enjoy what I have created and provides resources to allow me to create more content. If you have any feedback or suggestions around this, I’m all ears. You can also find me via email director@ciaops.com and on X (Twitter) at https://www.twitter.com/directorcia.

If you want to be part of a dedicated Microsoft Cloud community with information and interactions daily, then consider becoming a CIAOPS Patron – www.ciaopspatron.com.

Watch out for the next CIA Brief next week

What Happens If All LLMs Become the Same?

image

I’ve been thinking about a question that sounds a bit strange at first: what happens if all the major language models start converging into one common capability?

Not literally one company. Not one product. Not one button.

I mean something more subtle. What if the difference between the big models becomes less obvious to the average business user? What if the answer from one model is good enough, the answer from another is also good enough, and the real distinction is no longer the model itself but where it lives, what it can access, what it can do, and how safely it can do it?

That is a very different world from the one a lot of people are still arguing about.

The model may become the least interesting part

Right now, there is still a lot of energy around model comparison. Which one writes better? Which one reasons better? Which one codes better? Which one is cheaper? Which one won the latest benchmark?

That matters, but I’m not convinced it will matter in the same way for most organisations.

For many business users, the model is already beginning to disappear into the workflow. They don’t want to pick between ten engines before replying to an email. They want to open Outlook, ask Copilot to summarise the thread, draft a response, and make sure it reflects the real conversation. They want to sit in Teams, catch up on a meeting, identify the unresolved decisions, and move on.

If every serious model gets broadly competent at writing, reasoning, summarising, analysing, and planning, then the contest shifts. The question becomes less “which LLM is smartest?” and more “which environment gives this model the right context, guardrails, and business action?”

That is where Microsoft 365 starts to matter.

A generic super LLM might know a lot about the world. But it does not automatically know your SharePoint structure, your Teams conversations, your Outlook history, your policies, your client files, your permissions, or your business rhythm. And if it does get access to those things, the real issue becomes governance.

Common intelligence makes business discipline more important

If AI capability becomes common, then competitive advantage moves somewhere else.

It moves to your data quality.
It moves to your process maturity.
It moves to your permission model.
It moves to your ability to describe the outcome you actually want.

That is uncomfortable for many businesses because it means AI does not magically fix operational mess. It exposes it.

If your documents are scattered across personal OneDrives, old Teams channels, duplicated SharePoint libraries, and mystery folders called “Final Final Real Final”, a better model may not save you. It may simply find the wrong thing faster.

This is why I keep coming back to the practical layer. Before worrying about whether the world ends up with one dominant super LLM, I’d rather ask whether your organisation has clean source material, sensible access controls, repeatable workflows, and people who know how to challenge the output.

Ask Copilot in Word to draft a client-ready explanation from a properly maintained policy document, and you start to see real value. Ask it to work from five conflicting policy drafts and a half-forgotten email thread, and you get a polished problem.

AI does not remove responsibility. It compresses the time between messy input and messy output.

The future may be less about models and more about orchestration

My guess is that we won’t care as much about individual model names over time. We’ll care about orchestration.

Which model should handle this task?
Which data should it use?
Which actions is it allowed to take?
Which human signs off?
Which audit trail remains?

That is the operating model businesses need to build. Not a fan club for a particular LLM.

If all roads eventually lead to a broadly common intelligence layer, the winners will not be the organisations that simply had access to it. Everyone will. The winners will be the ones that wrapped that intelligence in good process, clean data, sensible governance, and practical human judgement.

The super LLM, if it arrives, may not be the finish line.

It may just be the new baseline.

And once everyone has the same baseline, the old boring things start to matter again: discipline, clarity, trust, and execution.

Comparing LLMs in Copilot services–Round 5 – Cowork

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

For this round I’ve pitted the Cowork Claude winner (Opus 5) vs the Cowork GPT winner (5.6 Terra). The result is:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-Grok-eval.md

1. Opus 5.0

2. GPT 5.6 Terra

and this time the gap was much larger than before. Not unexpected given that Terra is designed as a lighter weight, less powerful model. However, don’t forget the cost factor which isn’t included in these evaluations:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

If costs were indeed taken into account it would certain swing the pendulum back toward Terra.

However, this time the winner is clear

Opus 5

moves onto the next round.

The Client Journey Tells You Where the Business Really Is

image

There is a point in every service business where growth still looks busy, but it no longer feels healthy.

The calendar is full. Ads are running. Sales conversations are happening. New clients are arriving. From the outside, it looks like momentum.

Inside the business, something feels off.

You finish the month tired and the client count has hardly moved. Your team has handled more calls, answered more questions, chased more payments, and welcomed more people, yet the result barely reflects the effort. That is usually when I stop looking at the front door and start looking at the back one.

Because more leads will not fix a leaky experience.

Growth hides the problem until it does not

Early on, client loss can be easy to ignore. A few people leave and it does not feel significant. The business is small enough that enthusiasm covers rough edges.

Then the business gets bigger.

The same onboarding gaps become more expensive. The same unclear expectations create more friction. The same missed follow-ups turn into quiet exits. Nobody has to make a dramatic complaint for the damage to be real. Sometimes they just stop engaging, stop replying, stop showing up, and eventually stop paying.

This is where many owners make the obvious move. They spend more to bring more people in.

I understand why. Marketing feels active. Sales feels measurable. Pipeline gives everyone something to discuss in the weekly meeting. But if the client experience is not holding, all you have built is a costly replacement machine.

That is not growth. That is movement.

The first weeks set the whole relationship

I have become more interested in the first part of the client journey than almost any other part of the business.

Not because onboarding is glamorous. It is checklists, expectations, reminders, handovers, notes, meetings, small promises, and boring consistency.

But boring consistency is often where profit is hiding.

The first few weeks teach a client how your business works. They learn whether you are organised. They learn whether they need to chase you. They learn whether the promise they bought is becoming something tangible.

If that early experience is vague, the client starts filling in the gaps themselves. That is dangerous. Their version of the story may not be the one you intended.

This is where Microsoft 365 can do practical work. I would map the first client journey in a shared Word document or Loop workspace, turn the repeatable steps into Planner tasks, keep client context in a Teams channel, and use Copilot to summarise meeting notes, draft follow-up emails in Outlook, and identify action items that slipped.

None of that is magic. That is the point.

The value is not in making onboarding fancy. The value is in making it visible, repeatable, and harder to forget when the business gets busy.

Retention is an operating system

A lot of businesses treat retention as a customer service problem. I think that is too late.

Retention starts when the expectation is first set. It continues when the client understands what happens next. It improves when the team can see the same information, use the same process, and spot weak signals before they become cancellation emails.

You do not need a massive transformation project to start. Pick one client segment. Write down the first six weeks. Decide what each client should receive, what your team must do, what evidence shows progress, and where the handoffs fail.

Then put that into the tools your team already opens every day.

A better onboarding process will not remove every cancellation. Nothing does. But it changes the work. Instead of constantly buying attention from strangers, you start earning confidence from the people who already said yes.

That is the growth I trust more.

Not the noisy kind.

The kind that stays.

The Real Work Is Not Another Tactic

image

Most owners I meet are not short of effort. They are short of room.

Room to think. Room to make better decisions. Room to stop reacting to every urgent customer request, vendor announcement, staff issue, cashflow wobble, and half-finished idea sitting in the inbox.

That is why so many businesses end up looking busy but feeling fragile. The owner keeps tuning the visible parts of the machine. Better campaigns. Better processes. Better meetings. Useful, but none of it fixes the real constraint if the owner is still the bottleneck.

The business usually rises to the level of the person leading it.

That is an uncomfortable sentence. It should be.

The easy work looks productive

It is tempting to stay in work that gives visible proof of progress. Rewrite the website. Change the offer. Hire another person. Buy another app. Create another spreadsheet. Push harder on social. Start another initiative.

I have done versions of this myself. Most business owners have. It feels like discipline because there is activity everywhere.

But activity is not always advancement.

The harder work is looking at the recurring patterns and asking, “Why does this keep happening around me?” Why do the same decisions come back to my desk? Why do I avoid the awkward conversation until it becomes expensive? Why do I keep saying yes to work that does not fit?

That work rarely produces a neat announcement. But it is often where the real advantage is built.

Your operating system matters more than your tactics

I think about this a lot with Copilot and Microsoft 365.

Plenty of organisations are trying to use Copilot as a faster keyboard. Draft an email in Outlook. Summarise a meeting in Teams. Turn notes into a document in Word. All good uses.

But the more interesting use is not speed. It is reflection.

After a difficult client meeting, I can ask Copilot in Teams to help me identify the decisions, risks, and unresolved questions from the transcript. I can put the next actions into Planner, save the working document in SharePoint, and use that as the basis for a better follow-up in Outlook.

That is not just automation. That is a leadership loop.

The value is not that Copilot wrote some sentences for me. The value is that I forced myself to inspect the way I work. What did I miss? What did I delay? What needs to become a rule, not another heroic rescue?

A better business is usually built from better loops.

The real edge compounds quietly

The strongest operators I know are not chasing every new trick. They are harder to knock off balance.

They recover faster from mistakes. They make decisions with less drama. They communicate sooner. They document what matters. They know which work to refuse. They build teams that do not need constant rescue because the thinking has been made visible.

That is difficult to compete with because it is not one tactic someone can copy. It is a collection of habits, standards, judgement, and self-awareness built over time.

You can copy someone’s landing page. You can copy their pricing model. You can copy their tech stack.

You cannot easily copy the way they think under pressure.

That is where I believe owners should spend more attention. Not less work on the business, but better work on the person making the business decisions.

Use the tools. Use Copilot. Use Teams, Outlook, SharePoint, and Planner to create cleaner loops.

But do not confuse the tool with the transformation.

The business changes when the owner changes the way decisions are made, captured, reviewed, and improved.

That is the work most people avoid.

It is also the work that makes you hard to catch.

Comparing LLMs in Copilot services–Round 4 – Cowork (GPT)

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

This round the battle is between all the models available in Copilot Cowork GPT. Namely, these:

Screenshot 2026-08-12 082720

The other interesting factor here is that all these models as PAYG, so I’ll also give the costs for the same prompt.

I used Grok to evaluate all the results which produced:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-GPT-Grok-eval.md

In summary, the rankings are:

1. 5.6 Terra

2. 5.6 Sol

3. 5.5

Interestingly, the less powerful model (Terra) produced a better result here.

And the costs of each:


GPT 5.6 Sol = 2,800.50 credits
GPT 5.6 Terra = 260 credits
GPT 5.5 = 1,200 credits

5.6 Terra wins again big here, only costing 260 credits! This is how the ranked costs table so far looks, from most to least expensive:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

Given that Terra is so much cheaper I’ll need to spend some more time investigating and verifying that with other requests, given that these tests are simply a one shot prompt to result.

Drum roll. The clear winner for this round is:

GPT 5.6 Terra

With all the preliminaries done we can now get onto comparing the winners of each round together.

Why AI Charges More to Write Than to Read

image

I keep seeing people look at AI pricing pages and assume someone has made a typo.

Input tokens cost one amount. Output tokens cost another. Cached input may cost less again. At first glance that feels odd. A token is a token, surely?

Not really.

The easiest way to think about it is this: reading is cheaper than writing. AI can read a lot of your prompt in parallel. It can process the instructions, pasted document, previous conversation and system rules in large chunks. That is the input side.

Writing the answer is different. The model generates the response one token at a time. Each next word, fragment, line of code or JSON field depends on what came before it. The longer the response, the longer expensive compute is tied up producing it.

That is why output tokens usually cost more. You are not just paying for text. You are paying for generation.

This matters more than most people think

Once you move into agents, automation, Copilot Studio, Azure OpenAI, GitHub Copilot, or anything that runs repeatedly, the economics change quickly.

A short prompt that produces a long report can cost more than a large prompt that produces a tiny answer. That surprises people. They assume the big document is the expensive part. Sometimes it is not. The real cost can be the polished, verbose output they asked the model to create.

Ask an AI system to read a SharePoint policy library and return three risks, and you are probably dealing with an input-heavy, output-light workload. That can be relatively efficient.

Ask it to create a 25-page report, an executive summary, a remediation plan, a Teams post, a client email and a formatted table every time it runs, and you have created an output-heavy workload. That is where the bill starts to move.

The MSP lesson is simple: design the output

This is where MSPs need to stop treating AI as magic and start treating it as infrastructure.

When we built servers, we cared about CPU, RAM, disk and backup windows. With AI, we need to care about prompts, context, output length, caching and repeatability.

The bad habit is asking for everything every time. “Give me the full report.” “Include all the detail.” “Make it comprehensive.” That sounds harmless until the same workflow runs fifty times across fifty tenants.

A better approach is to be deliberate. Ask for the smallest useful output first. Use summaries where summaries are enough. Generate detailed reports only when there is a reason. Reuse stable instructions and context where caching is available. Put spending limits around anything consumption-based. Review what the agent writes, not just what it reads.

Inside Microsoft 365, this means being clear about the difference between ordinary Copilot use in Word, Excel, Outlook or Teams and consumption-based AI work that may be billed differently. A user drafting an email is one thing. An agent chewing through documents and producing long artefacts all day is another.

This is not a pricing trick

I do not see the input/output price split as some mysterious vendor tax. It reflects how the technology behaves.

The mistake is pretending it does not matter.

AI costs are not just about how many people have a licence. They are about what those people, agents and workflows ask the model to produce. The output is where the hidden weight often sits.

So the practical rule is this: do not just prompt for the result. Design the cost shape of the result.

That might be the difference between AI being a useful business tool and AI becoming the next cloud bill nobody wants to open.