Care About The Work

image

I can usually tell when a team is disconnected from the work it produces.

Not because the work is terrible. Often it looks fine from the outside. The newsletter goes out. The video gets posted. The webinar happens. The support article lands in SharePoint. The client update gets sent from Outlook.

But inside the business, nobody is really paying attention.

The team does not watch the video. They do not read the newsletter. They do not know what was promised in the webinar. They do not see the questions clients are asking after the content goes live.

That is where the problem starts.

Your own team is the first audience

I think many businesses make a mistake here. They treat content, education, communication and client resources as something created by one person or one small group. Everyone else is expected to be “too busy” to engage with it.

I see it differently.

If you work in the business, you need to understand what the business is saying.

That does not mean everyone has to become a marketer. It does not mean every technician needs to write articles. It does not mean every account manager needs to appear on camera.

It means the people who serve the client should know the message the client is hearing.

If a Microsoft 365 security update goes out to clients, the service desk should know about it before the client rings. If a Copilot readiness article is published, the account manager should know the point of view before the next review meeting. If a training video explains why governance matters, the project team should have watched it before they start implementation work.

That is not bureaucracy. That is alignment.

Consumption creates better feedback

The other thing people miss is that internal consumption makes the work better.

A team that reads, watches and listens can challenge weak ideas before clients see them. They can say, “That example will confuse people.” They can say, “Clients are already asking about this.” They can say, “We need a simpler explanation for the business owner.”

That feedback is gold.

This is where Microsoft 365 can really help, but only if it is used deliberately. I like the idea of having a dedicated Teams channel where published content, draft ideas and client questions are visible. Drop the newsletter link in there. Pin the SharePoint page. Use Copilot in Teams to summarise the discussion after people have commented. Capture the useful follow-up items in Planner rather than letting them disappear into chat history.

The tool is not the culture. But the tool can make the culture easier to practise.

The important part is that everyone understands the loop.

Create something. Share it internally. Let the team react. Improve the next version. Listen to what clients say. Bring that back to the team. Repeat.

That loop builds judgement. It also builds ownership.

Make it expected, not optional

I do not have much time for the idea that “mandatory” automatically means heavy-handed.

Some things should be expected.

If the business publishes advice to clients, the team should know what that advice is. If the business takes a position on AI, security, productivity or service delivery, the team should be familiar with that position. If clients are being educated, the people supporting those clients should not be the last to know.

The standard is simple: if we produce it, we pay attention to it.

That does not need to become a long meeting or a corporate ritual. It can be a short weekly review. It can be a Teams post with three questions. It can be a Copilot-generated summary of what went out, what clients asked, and what the team noticed.

But it does need to happen.

Because care is visible. Not in a fluffy motivational sense, but in the final product. You can feel when a team has thought about the work. You can feel when people have challenged it, improved it, and connected it back to the client.

You can also feel when it was pushed out by someone working alone while everyone else walked past it.

If your own people are detached from what you create, do not be surprised when the outside world is detached as well.

The work has to matter internally first.

That is where better culture starts. That is where better feedback starts. And that is where better client outcomes start.

Add Microsoft Learn as a knowledge source to Copilot Chat

image

I need to refer to Microsoft Learn documentation a lot. Therefore, I want it included irectly as part of my data sources with Copilot Chat. To do that go to the settings in the top right hand corner as shown above.

image

Add Microsoft Learn to your sources from the Browse sources section as shown above.

That will now make Microsoft Learn a primary data source when prompting Copilot Chat.

Productivity Starts With Owning The Work

image

When someone on my team tells me they are overwhelmed, I usually do not start with the task itself. I start with where the work is captured.

Not in their head. Not spread across ten chats. Not buried in a half-read email thread. Captured somewhere they can see it, sort it, challenge it, and act on it.

That sounds basic, but I keep seeing good people come unstuck at this point. They are smart. They care. They are capable. Yet the work is scattered, and once that happens, every day becomes reactive.

The problem is not usually effort.

It is visibility.

Busy Is Not The Same As In Control

A lot of people mistake movement for productivity. They answer messages quickly. They jump into meetings. They fix whatever is currently making noise. From the outside, it can look like they are working hard.

But if you ask what matters most this week and the answer is vague, there is a management problem. Not always from above. Often it is a self-management problem.

I have learned not to treat that lightly.

If a person cannot show what they are working on, what is waiting, what is blocked, and what matters next, then they are running their day on memory and pressure. That might survive for a while, but it does not scale.

This is where Microsoft 365 can help, but only if the behaviour is already there. Planner will not make someone disciplined. Teams will not magically create priorities. Copilot will not fix a messy operating rhythm if there is no trusted place for the work to live.

The tool supports the habit.

It does not replace it.

Make The Work Visible

The best teams I see make work visible early.

That might mean capturing actions from a Teams meeting into Planner. It might mean asking Copilot in Outlook to summarise a long thread, then turning the actual next steps into tasks instead of leaving them as good intentions. It might mean keeping a shared Loop page for the current priorities so everyone can see what has changed.

None of that is fancy.

That is the point.

Productivity is not about building a complicated system that only one person understands. It is about creating enough structure so the next decision is easier. What needs doing? What can wait? What needs someone else? What should be deleted entirely?

Those questions matter far more than another app, another dashboard, or another clever prompt.

Lead The Standard Before You Demand It

If I want my team to operate with clarity, I have to model it first.

That means I need my own priorities written down. I need to show how I decide what gets attention and what does not. If everything is urgent when it reaches me, I am teaching the team that noise wins.

This is especially important with Copilot adoption. I can tell people to use Copilot all I like, but if my own workflow is chaotic, the message will not land. A better approach is to show the habit in action. Here is the meeting summary. Here are the actions. Here is what I moved into Planner. Here is what I am not doing this week.

That is how productivity becomes cultural rather than personal.

The Real Foundation

I do not see personal productivity as a soft skill.

I see it as operating discipline.

Before strategy, before automation, before AI, people need to manage their own commitments. If they cannot do that, every improvement you add sits on unstable ground.

So start small. Make the work visible. Decide what matters. Put it somewhere trusted. Review it regularly.

Then teach the team to do the same.

Because once people can manage themselves, everything else has a chance to work.

Comparing LLMs in Copilot services–Round 7-Final

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

Round 5  – https://blog.ciaops.com/2026/08/14/comparing-llms-in-copilot-services-round-5-cowork/

Winner – Opus 5

Round 6 – https://blog.ciaops.com/2026/08/16/comparing-llms-in-copilot-services-round-6/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

For the final, we have Opus 5 (Cowork) and Opus (Chat). The results are:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-Copilot-Grok-eval.md

1. Opus 5 (Cowork)

2. Opus (Chat)

Therefore, the initial over all Copilot LLM winner is a clear win for:

Opus 5 (Cowork)

The trade off is that the cost for this US$22, while the runner up’s price was included in the cost of a Microsoft 365 Copilot license.

These test have revealed a few things, in my opinion:

A. Claude Opus is the superior model

B. Opus in chat is almost as good as Opus Cowork but much cheaper

C. For most work Opus in chat is probably you best option

The challenge with these report, as with any LLM’s is they are not definitive. Another round could reveal completely different results as could the method of evaluation, however as a best effort I think it does have validity and provide some general findings that do help understand the model choices in Copilot.

For comparison, I have create a document parameters page here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-File-paramters.md

so you can now compare all the various output document sizes, pages, etc to each other. I have also made available a copy of each model output document, without any changes to it, so you can look at each for yourself. You will find them all at:

https://github.com/directorcia/general/tree/master/Copilot/Comparisons/20260816

The idea is to wait until we see new models appear in Copilot and use the same methodology against these also when they appear. That will hopefully provide some sort of bench mark when it comes top model strength.

I hope this series of tests has been interesting and helpful to you and I’d love to hear your thoughts on the results.

Comparing LLMs in Copilot services–Round 6

MAI_cf947564d0f53b48

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

Round 5  – https://blog.ciaops.com/2026/08/14/comparing-llms-in-copilot-services-round-5-cowork/

Winner – Opus 5

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

Next up, I’ve pitted Opus 5 and Researcher Critique. The results are:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260814-Deep-Grok-eval.md

1. Opus 5

2. Researcher Critique

the gap between these two was quite large again. The clear winner therefore is:

Opus 5

This mean the upcoming final is going to be Opus 5 (Coworok) vs Opus (Chat) for thr crown which will be very interesting given they are the same model, while one is an included LLM and the other is PAYG!

Stay tuned for the action of the first CIAOPS LLM final!

Does Your AI Environment Need a Reset?

image

I used to rebuild my computer every so often.

Not because it was broken. Not because the hardware had failed. It was because, over time, the machine collected rubbish. Old drivers. Forgotten utilities. Startup items I no longer needed. Temporary files. Half-removed applications. Things that made sense at the time, but slowly turned a clean system into something sluggish and unpredictable.

I think we are heading toward the same problem with AI.

Not with the model itself. With the environment around it.

AI inherits your mess

When people talk about AI performance, they usually talk about the model. Bigger model. Newer model. Faster model. Better reasoning. More context.

That matters, but it is only part of the story.

Inside a business, Microsoft 365 Copilot works with the information, permissions, meetings, files, chats, and workflows already sitting in your tenant. If those foundations are messy, Copilot does not magically clean them. It works with what it finds.

If SharePoint is full of abandoned sites, Copilot has more noise to search through. If Teams channels contain years of half-finished conversations, it has more stale context to interpret. If OneDrive is full of duplicate drafts named “final-final-really-final.docx”, the AI is not the problem. The house is.

This is why I think an AI reset is becoming a real operational practice.

The reset is not a rebuild

I am not suggesting businesses wipe everything and start again. That would be ridiculous.

The better comparison is a deliberate clean-up cycle. A review of what AI can see, what it should use, what it should ignore, and what needs to be retired.

That means checking who has access to what. It means looking at old SharePoint sites, inactive Teams, forgotten guest accounts, broad sharing links, stale files, unused prompts, and agents nobody owns anymore.

It also means reviewing the human habits around AI.

Are people still using the same prompts they wrote six months ago? Are those prompts pointing at the right source material? Are users asking Copilot to summarise everything because they cannot be bothered choosing the right file? Are teams saving useful prompt patterns somewhere reusable, or are the good ones buried in private chat history?

That is cruft as well.

Performance is not just speed

When an old PC slowed down, the answer was often to remove junk and reduce the background load. With AI, performance is broader than speed.

A good AI environment gives better answers because the data is cleaner. It gives safer answers because permissions are tighter. It gives more useful answers because users know which sources to reference. It gives more predictable outcomes because prompts and workflows are standardised.

For example, if a user asks Copilot in Outlook to draft a client follow-up based on a meeting, the quality depends on more than Copilot. Was the meeting transcribed? Were the notes clear? Are the relevant project files in the right SharePoint location? Is the client information current? Has the user given Copilot a clear task, or just thrown a vague request at it?

The model may be smart. The environment still has to be tidy.

The monthly AI reset

I can see this becoming a monthly rhythm for MSPs and internal IT teams.

Review Copilot usage. Check new agents. Inspect oversharing. Look at sensitivity labels. Clean up old guests. Archive dead Teams. Refresh prompt libraries. Remove obsolete source documents. Confirm DLP policies still make sense. Ask whether AI is helping real work or just creating more output to manage.

That is not busywork. That is how you stop AI from drifting into the same state as an old Windows install full of utilities nobody remembers installing.

AI does not remove the need for operational discipline. It raises the price of not having it.

The organisations that get the most from AI will not be the ones constantly chasing the newest model. They will be the ones that keep their environment clean enough for AI to work with confidence.

Sometimes the best AI upgrade is not a new feature.

It is a clean-up.

What Happens If All LLMs Become the Same?

image

I’ve been thinking about a question that sounds a bit strange at first: what happens if all the major language models start converging into one common capability?

Not literally one company. Not one product. Not one button.

I mean something more subtle. What if the difference between the big models becomes less obvious to the average business user? What if the answer from one model is good enough, the answer from another is also good enough, and the real distinction is no longer the model itself but where it lives, what it can access, what it can do, and how safely it can do it?

That is a very different world from the one a lot of people are still arguing about.

The model may become the least interesting part

Right now, there is still a lot of energy around model comparison. Which one writes better? Which one reasons better? Which one codes better? Which one is cheaper? Which one won the latest benchmark?

That matters, but I’m not convinced it will matter in the same way for most organisations.

For many business users, the model is already beginning to disappear into the workflow. They don’t want to pick between ten engines before replying to an email. They want to open Outlook, ask Copilot to summarise the thread, draft a response, and make sure it reflects the real conversation. They want to sit in Teams, catch up on a meeting, identify the unresolved decisions, and move on.

If every serious model gets broadly competent at writing, reasoning, summarising, analysing, and planning, then the contest shifts. The question becomes less “which LLM is smartest?” and more “which environment gives this model the right context, guardrails, and business action?”

That is where Microsoft 365 starts to matter.

A generic super LLM might know a lot about the world. But it does not automatically know your SharePoint structure, your Teams conversations, your Outlook history, your policies, your client files, your permissions, or your business rhythm. And if it does get access to those things, the real issue becomes governance.

Common intelligence makes business discipline more important

If AI capability becomes common, then competitive advantage moves somewhere else.

It moves to your data quality.
It moves to your process maturity.
It moves to your permission model.
It moves to your ability to describe the outcome you actually want.

That is uncomfortable for many businesses because it means AI does not magically fix operational mess. It exposes it.

If your documents are scattered across personal OneDrives, old Teams channels, duplicated SharePoint libraries, and mystery folders called “Final Final Real Final”, a better model may not save you. It may simply find the wrong thing faster.

This is why I keep coming back to the practical layer. Before worrying about whether the world ends up with one dominant super LLM, I’d rather ask whether your organisation has clean source material, sensible access controls, repeatable workflows, and people who know how to challenge the output.

Ask Copilot in Word to draft a client-ready explanation from a properly maintained policy document, and you start to see real value. Ask it to work from five conflicting policy drafts and a half-forgotten email thread, and you get a polished problem.

AI does not remove responsibility. It compresses the time between messy input and messy output.

The future may be less about models and more about orchestration

My guess is that we won’t care as much about individual model names over time. We’ll care about orchestration.

Which model should handle this task?
Which data should it use?
Which actions is it allowed to take?
Which human signs off?
Which audit trail remains?

That is the operating model businesses need to build. Not a fan club for a particular LLM.

If all roads eventually lead to a broadly common intelligence layer, the winners will not be the organisations that simply had access to it. Everyone will. The winners will be the ones that wrapped that intelligence in good process, clean data, sensible governance, and practical human judgement.

The super LLM, if it arrives, may not be the finish line.

It may just be the new baseline.

And once everyone has the same baseline, the old boring things start to matter again: discipline, clarity, trust, and execution.

Comparing LLMs in Copilot services–Round 5 – Cowork

image

Round 1 – https://blog.ciaops.com/2026/07/30/comparing-llms-in-copilot-services-round-1-chat/

Winner – Opus

Round 2 – https://blog.ciaops.com/2026/08/08/comparing-llms-in-copilot-services-round-2-researcher/

Winner – Critique

Round 3  – https://blog.ciaops.com/2026/08/11/comparing-llms-in-copilot-services-round-3-cowork-claude/

Winner – Opus 5

Round 4 – https://blog.ciaops.com/2026/08/12/comparing-llms-in-copilot-services-round-4-cowork-gpt/

Winner – GPT 5.6 Terra

I’ve been pitting different LLMs inside M365 Copilot against each other ina world cup style elimination to see which comes out on top. The process involves taking a standard prompt and running it against all options. This prompt creates a multi page document requiring deep research and is quite involved. The results are then compared against each other using SharePoint Copilot and Gemini. Conclusions are then drawn.

For this round I’ve pitted the Cowork Claude winner (Opus 5) vs the Cowork GPT winner (5.6 Terra). The result is:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260811-Cowork-Grok-eval.md

1. Opus 5.0

2. GPT 5.6 Terra

and this time the gap was much larger than before. Not unexpected given that Terra is designed as a lighter weight, less powerful model. However, don’t forget the cost factor which isn’t included in these evaluations:

Claude Fable 5 (Preview) = 3,387 credits
Claude Fable 5 (Copilot)(Preview) = 3,373.8 credits

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

If costs were indeed taken into account it would certain swing the pendulum back toward Terra.

However, this time the winner is clear

Opus 5

moves onto the next round.