The Latest AI Model Is Not Always the Best Tool

image

I see this conversation more and more. Someone opens an AI tool, sees a new model name at the top of the list, and assumes that must be the one to use for everything. Newer must be smarter. More expensive must mean better results.

I do not think it is that simple.

The better question is not, “Which model is the latest?” The better question is, “What decision am I trying to improve with this task?” That matters as AI becomes normal work inside Outlook, Teams, Word, Excel, and Microsoft 365 Copilot.

The expensive model is often wasted on cheap work

A lot of AI work is not deep thinking. It is sorting, summarising, rewriting, comparing, and turning messy notes into something usable.

If I ask Copilot in Outlook to tidy a reply, I do not necessarily need the most advanced reasoning model available. I need something that understands the email thread, preserves the intent, and helps me move the conversation forward. That does not require the biggest hammer in the toolbox.

The same applies to meeting recaps in Teams. If the job is to produce a clean summary, identify obvious actions, and help me catch up, then the real value is context and workflow. The model matters, but so does whether the result lands where I can use it.

This is where many businesses get the economics wrong. They confuse “best model” with “best outcome”. Those are not always the same thing.

Better models matter when the work is harder

Advanced models have a place.

If I am asking AI to compare strategies, review a complex proposal, reason across several documents, test assumptions, or help me structure an argument, I want the strongest model I can reasonably use. That is where better reasoning can show up: fewer shallow answers, better trade-offs, and more useful challenge.

For example, if I have a client planning an AI adoption project, I might ask Copilot to review Teams notes, pull themes from Word documents in SharePoint, and help me identify risks that have not been properly addressed. In that situation, a stronger model may produce a better result because the task requires judgement and careful comparison.

But even then, the model is only part of the answer. If the source material is poor, scattered, outdated, or inaccessible, the smartest model in the world is still working with a weak brief. AI does not fix bad information hygiene. It exposes it.

Pay for capability, not fashion

The practical approach is to tier your use.

Use the everyday model for everyday work. Draft the email. Summarise the meeting. Clean up the notes. Turn the rough spreadsheet explanation into something your client can understand.

Use the more advanced model when the cost of being wrong is higher or the thinking is genuinely harder. Strategy. Risk. Governance. Technical design. Client-facing recommendations. Anything where you need deeper reasoning rather than quicker wording.

That is also the advice I would give MSPs talking to SMB clients. Do not sell AI as a race to the newest model. Help clients build judgement about when AI is good enough, when it needs checking, and when the work deserves the better tool.

The latest model may be worth paying for. Sometimes. But always using it is not automatically clever. It can become another form of waste dressed up as sophistication.

The real maturity test is not whether you have access to the newest AI model. It is whether you know when to use it, when not to use it, and how to measure whether it actually improved the result.

That is the shift I am watching now. Not model chasing. Outcome choosing.


What Happens When The Screen Stops Being The Starting Point?

image

I keep coming back to a simple question: what if the future interface is not a grid of icons?

For years we have trained ourselves to use devices by hunting for the right app, opening it, finding the right menu, tapping the right button, then hoping we remembered where the command lives. That feels normal because we have done it for so long. But normal is not the same as permanent.

AI changes that. Not because it makes the existing interface prettier, but because it challenges whether that interface needs to be the starting point at all.

The app may become the background

I don’t think apps disappear overnight. That is too neat and too dramatic. Businesses still need systems of record. People still need Outlook, Teams, Word, Excel, SharePoint and the rest. The real shift is that the app may stop being where the user begins.

Instead of opening Outlook, finding the email thread, checking the calendar, then drafting a reply, I can see people simply asking Microsoft 365 Copilot: “What do I need to respond to before tomorrow morning, and can you draft the first three replies?”

That is not just a faster way to use Outlook. It is a different relationship with the machine.

The software is still there. The data is still in Microsoft 365. The security, permissions, retention and compliance boundaries still matter. But the user experience moves up a layer. The task becomes the centre, not the application.

That is a big change.

Mobile makes this more obvious

This shift feels especially likely on mobile devices.

A phone is powerful, but it is still a small piece of glass. The more we expect from it, the more ridiculous some workflows become. Open an app. Switch apps. Copy something. Paste it somewhere else. Tap through three screens. Accept a prompt. Go back because you missed something.

That is a lot of ceremony for a device that is usually in your hand while you are walking, travelling, waiting, or trying to get something done between other things.

Voice changes the equation. If I can pick up a phone and say, “Summarise the Teams discussion about that client issue, check whether there is anything in my inbox I need to act on, and create a Planner task for the follow-up,” then the icons become less important.

Not irrelevant. Just less central.

The interface becomes more like a conversation with context. That means the device needs to understand intent, identity, permissions and the work graph around me. This is where Microsoft 365 has an advantage, because so much of the business context already lives inside Outlook, Teams, SharePoint, OneDrive and the calendar.

The danger is sloppy delegation

There is a trap here though.

Asking AI to “do something” sounds simple, but business work is rarely simple. A useful assistant needs to know when to act, when to ask, when to show its reasoning, and when to stop. That matters even more when the interface becomes conversational.

If the AI becomes the front door to the device, then governance becomes part of the user experience. Not something hidden in the admin centre. Not something bolted on after the fact.

For MSPs and business owners, that is the real lesson. The future interface may be conversational, but the foundation still needs to be boringly practical: identity, conditional access, data classification, permissions, auditability and user training.

AI does not remove operational discipline. It exposes whether you had any.

I’m watching the starting point

I don’t think the question is whether AI replaces the operating system. The operating system will still exist. Something has to manage the device, the hardware, the identity and the applications.

The better question is whether people will feel like they are using an operating system at all.

My guess is that, over time, more people will start with the request rather than the app. They will ask, instruct, refine and approve. The screen will still matter, but it may become more of a confirmation surface than a navigation surface.

That is the shift I’m watching. Not AI as another icon on the device, but AI as the place where the work begins.

Stung by my own stupidity

If you aren’t aware, when you built or run anything new in Copilot Studio it will all be charged PAYG against Azure. To allow this you need to connect your Power Platform environment to an Azure subscription. You can find details on how to do that here:

https://blog.ciaops.com/2022/04/29/set-up-payg-for-power-platform/

Now with all that in place I built a new modern agent in Copilot Studio to answer M365 questions (called Sage) built using the new Github Copilot harness and using Claude Opus 5 as the LLM.

image

I then tested it a few times in ‘Preview’ and was happy that it was all working. I knew at this point, all that was going to cost me a few bucks because now even creation costs with the new Copilot Studio. All good so far and still in budget.

Next, I wired up a new Workflow in Copilot Studio to wait for a message to be posted into a Microsoft Teams channel, take that, post it to the newly created Sage agent, then take the reply from the agent and post it back into the same channel. Quick and easy to create. Job done, I thought.

Can you see the logic flaw yet? I certainly didn’t initially. In short, the workflow I created basically replies to every message posted into a channel. Ahem, those replies then trigger the agent to run again and post yet another message, which again triggers another message posting from the agent, and on and on. So, I had created an infinite loop.

My mistake was now running the workflow and calling the agent and posting into the Team every minute or so. I didn’t recognise my error for a few hours! Yes, hours. I estimate the loop I created with the workflow ran for about 3.5 hours in total. Ouch. When I finally realised upon checking back into the channel I immediately deleted the workflow to stop the race condition, however I knew I was going to pay for my mistake.

Fast forward a day or so when I have all the billing data available. Here’s what the results of my oversight were:

Screenshot 2026-08-21 073742

The error had cost me around AU$250. D’Oh!

All of this is always a learning experience, so now that I had understood the ‘bill shock’ amount I wanted to see what more billing information I could obtain about what had actually happened. I visited the Power Platform admin center | Licensing Copilot Studio, scrolled down to Top 5 agents and users, then selected View all agents which showed me this:

image

then when I drilled into my environments I can see:

image

and at the bottom you can see the autonomous consumption of 16,376.01 credits. If I divide that by the 3.5 hour run time I get 4,678.86 credits consumed per hour. If I then divide that by 60 to get the cost per minute I get 77.98 credits. Thus, each post to the channel in effect cost around US$0.78 which is about AU$1.20.

The detail also shows that creating the agent cost around 67.26 + 379.77 = 447.37 credit which is around US$4.50 and AU$6.95 to create.

I have now added notifications at 90% capacity like so in this admin console because they are not enabled by default:

image

I would expect, like the budget notifications from Azure, they are not immediate which makes avoiding costly mistakes harder when your error maybe racking up a few dollars per minute charges!

I accept full responsibility for my error and oversight and bill incurred, however I think there are some important learnings and observations here with the new PAYG billing for AI services. These in essence boil down to the fact that it very difficult to get a good understanding of exactly what your costs are in real time or prior. Typically, you need to wait a full 24 hours until all the billing data has been collected and by then you maybe up for thousands of dollars if you are not very careful.

Another observation is that if you make a logic error in your build you won’t find that until you look at your bill. I was lucky that I found mine after a few hours, imagine if it had run for more than 24 hours before the billing data alerted me? Ouch.

I believe this lack of immediacy an d visibility on costs is going to be a major barrier for adoption of PAYG agents in Microsoft 365, whether Cowork or the new Copilot Studio, especially in SMB.

image

Hopefully, we get to a point like we have with Github Copilot (above) where I can quickly and easily see my usage in the development environment (here Visual Studio Code). Without this type of spending certainty many business are simply not going to use what are fantastic AI tools to help their business. This risk of runaway costs is simply too great.

Another point that I want to reinforce here is that when you implement PAYG with agents you need to monitor your costs DAILY! This will be a big change for many MSPs who may occasionally go into a customers tenant to look at licensing monthly. If your customer has PAYG AI and you are responsible for managing these costs you need to keep an eye on this every single day to minimise what a single logic error could cost.

Ensure you enable all the alerting that you can when you use PAYG AI services, no matter where or whom they are from. Hopefully, doing this and my sorry tale here helps you better monitor your costs and avoid ‘AI usage bill shock’.

Your Skills Files Are Business IP

image

I keep seeing businesses get excited about building reusable prompts, agents, SOPs and skills.md files. Fair enough. That is where the real value starts to appear. Not in one clever prompt, but in documenting how the business works and making that repeatable.

But there is a quiet problem underneath it.

The better those files become, the more they stop being “documentation” and start becoming business intellectual property. A well-written skills.md file may contain how you scope jobs, respond to clients, handle exceptions, use Microsoft 365 Copilot, build proposals, or deliver a managed service consistently. That is not just a file. That is your operating model written down.

And yet, in many businesses, that content sits in a SharePoint library or Teams channel where almost everyone can read it, copy it, sync it, or forward it.

Convenient? Yes. Sensible? Not always.

The process library is now crown-jewel data

Businesses normally think about sensitive data as payroll files, contracts, financial spreadsheets and customer records. Those still matter. But AI changes the definition of what is valuable.

If your business has spent months refining reusable skills, prompts, checklists and delivery playbooks, then you have created a process asset. It explains how your business turns knowledge into outcomes. That deserves proper governance.

I am not saying every user should be locked out. The point of documenting processes is that people can use them. But there is a difference between “available to the team that needs it” and “available to anyone who inherited access from a Team created three years ago”.

That distinction matters when these files become grounding material for Copilot or custom agents. If Copilot can find the content, summarise it and reshape it quickly, then poor permissions become easier to exploit.

Copilot is not the leak. Oversharing is.

Access should match the work

The first control is boring and important: permissions.

Store these files in a dedicated SharePoint site or library. Do not scatter them across personal OneDrives, random Teams channels, email attachments and old project folders. Give the library an owner. Use Microsoft 365 groups or Entra ID security groups to control access. Review membership regularly.

The test I like is simple. If someone left tomorrow and joined a competitor, what could they still download today?

That question cuts through wishful thinking.

Some people need edit rights. Most only need read rights. Some only need access to the process area they work in. If everyone has everything, you do not have knowledge management. You have a shared filing cabinet with the front door open.

Labels and DLP are not decorations

This is where Microsoft Purview matters.

Apply sensitivity labels to the process library. A label like Confidential – Internal Process can make the handling expectation clear and drive protection such as encryption and access restrictions where appropriate.

Then add Data Loss Prevention policies around the same content. If someone tries to email a bundle of process files externally, copy them into chat with an outside party, or move them into unmanaged locations, you want friction, warning, logging, or blocking depending on the risk.

This is not about distrusting staff. It is about recognising that staff move on, mistakes happen, and valuable business knowledge should not be one drag-and-drop away from leaving the organisation.

Offboarding is too late

Many businesses only think about this when someone resigns. By then, the files may already be synced locally, copied into another tool, or forwarded elsewhere.

The right time to protect process IP is when the library is created. Classify it. Restrict it. Monitor it. Review it. Make access part of onboarding and removal part of offboarding.

If you are building AI skills for your business, treat them as assets, not notes.

The future advantage will not belong to the business with the most prompts. It will belong to the business that protects, improves and governs the way it works.

Add Microsoft Learn as a knowledge source to Copilot Chat

image

I need to refer to Microsoft Learn documentation a lot. Therefore, I want it included irectly as part of my data sources with Copilot Chat. To do that go to the settings in the top right hand corner as shown above.

image

Add Microsoft Learn to your sources from the Browse sources section as shown above.

That will now make Microsoft Learn a primary data source when prompting Copilot Chat.

Comparing LLMs in Copilot services–Round 7-Final

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

Round 5  – https://blog.ciaops.com/2026/08/14/comparing-llms-in-copilot-services-round-5-cowork/

Winner – Opus 5

Round 6 – https://blog.ciaops.com/2026/08/16/comparing-llms-in-copilot-services-round-6/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

For the final, we have Opus 5 (Cowork) and Opus (Chat). The results are:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-Copilot-Grok-eval.md

1. Opus 5 (Cowork)

2. Opus (Chat)

Therefore, the initial over all Copilot LLM winner is a clear win for:

Opus 5 (Cowork)

The trade off is that the cost for this US$22, while the runner up’s price was included in the cost of a Microsoft 365 Copilot license.

These test have revealed a few things, in my opinion:

A. Claude Opus is the superior model

B. Opus in chat is almost as good as Opus Cowork but much cheaper

C. For most work Opus in chat is probably you best option

The challenge with these report, as with any LLM’s is they are not definitive. Another round could reveal completely different results as could the method of evaluation, however as a best effort I think it does have validity and provide some general findings that do help understand the model choices in Copilot.

For comparison, I have create a document parameters page here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-File-paramters.md

so you can now compare all the various output document sizes, pages, etc to each other. I have also made available a copy of each model output document, without any changes to it, so you can look at each for yourself. You will find them all at:

https://github.com/directorcia/general/tree/master/Copilot/Comparisons/20260816

The idea is to wait until we see new models appear in Copilot and use the same methodology against these also when they appear. That will hopefully provide some sort of bench mark when it comes top model strength.

I hope this series of tests has been interesting and helpful to you and I’d love to hear your thoughts on the results.

Comparing LLMs in Copilot services–Round 6

MAI_cf947564d0f53b48

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

Round 5  – https://blog.ciaops.com/2026/08/14/comparing-llms-in-copilot-services-round-5-cowork/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

Next up, I’ve pitted Opus 5 and Researcher Critique. The results are:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260814-Deep-Grok-eval.md

1. Opus 5

2. Researcher Critique

the gap between these two was quite large again. The clear winner therefore is:

Opus 5

This mean the upcoming final is going to be Opus 5 (Coworok) vs Opus (Chat) for thr crown which will be very interesting given they are the same model, while one is an included LLM and the other is PAYG!

Stay tuned for the action of the first CIAOPS LLM final!

Does Your AI Environment Need a Reset?

image

I used to rebuild my computer every so often.

Not because it was broken. Not because the hardware had failed. It was because, over time, the machine collected rubbish. Old drivers. Forgotten utilities. Startup items I no longer needed. Temporary files. Half-removed applications. Things that made sense at the time, but slowly turned a clean system into something sluggish and unpredictable.

I think we are heading toward the same problem with AI.

Not with the model itself. With the environment around it.

AI inherits your mess

When people talk about AI performance, they usually talk about the model. Bigger model. Newer model. Faster model. Better reasoning. More context.

That matters, but it is only part of the story.

Inside a business, Microsoft 365 Copilot works with the information, permissions, meetings, files, chats, and workflows already sitting in your tenant. If those foundations are messy, Copilot does not magically clean them. It works with what it finds.

If SharePoint is full of abandoned sites, Copilot has more noise to search through. If Teams channels contain years of half-finished conversations, it has more stale context to interpret. If OneDrive is full of duplicate drafts named “final-final-really-final.docx”, the AI is not the problem. The house is.

This is why I think an AI reset is becoming a real operational practice.

The reset is not a rebuild

I am not suggesting businesses wipe everything and start again. That would be ridiculous.

The better comparison is a deliberate clean-up cycle. A review of what AI can see, what it should use, what it should ignore, and what needs to be retired.

That means checking who has access to what. It means looking at old SharePoint sites, inactive Teams, forgotten guest accounts, broad sharing links, stale files, unused prompts, and agents nobody owns anymore.

It also means reviewing the human habits around AI.

Are people still using the same prompts they wrote six months ago? Are those prompts pointing at the right source material? Are users asking Copilot to summarise everything because they cannot be bothered choosing the right file? Are teams saving useful prompt patterns somewhere reusable, or are the good ones buried in private chat history?

That is cruft as well.

Performance is not just speed

When an old PC slowed down, the answer was often to remove junk and reduce the background load. With AI, performance is broader than speed.

A good AI environment gives better answers because the data is cleaner. It gives safer answers because permissions are tighter. It gives more useful answers because users know which sources to reference. It gives more predictable outcomes because prompts and workflows are standardised.

For example, if a user asks Copilot in Outlook to draft a client follow-up based on a meeting, the quality depends on more than Copilot. Was the meeting transcribed? Were the notes clear? Are the relevant project files in the right SharePoint location? Is the client information current? Has the user given Copilot a clear task, or just thrown a vague request at it?

The model may be smart. The environment still has to be tidy.

The monthly AI reset

I can see this becoming a monthly rhythm for MSPs and internal IT teams.

Review Copilot usage. Check new agents. Inspect oversharing. Look at sensitivity labels. Clean up old guests. Archive dead Teams. Refresh prompt libraries. Remove obsolete source documents. Confirm DLP policies still make sense. Ask whether AI is helping real work or just creating more output to manage.

That is not busywork. That is how you stop AI from drifting into the same state as an old Windows install full of utilities nobody remembers installing.

AI does not remove the need for operational discipline. It raises the price of not having it.

The organisations that get the most from AI will not be the ones constantly chasing the newest model. They will be the ones that keep their environment clean enough for AI to work with confidence.

Sometimes the best AI upgrade is not a new feature.

It is a clean-up.