GPT 6.1 Sol is our new champ

image

Another week, another new model. This time it is GPT Sol 6.1 arriving, so I’ve put it to the test against the current champion GPT Astra 6 and found:

GPT Sol 6.1 beats Astra 6 – https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260930-Cowork-GPT61Sol-Grok-eval.md

but it is more expensive than Astra 6:

Opus 5.5 = 23,494 credits

GPTSol 6.1 = 10,173 Credits

GPT Astra 6 = 6,067 credits

Fable 5.1 = 3,453 credits

Fable 5 (Preview) = 3,387 credits [Model no longer shown]
Fable 5 (Copilot)(Preview) = 3,373.8 credits [Model no longer shown]

GPT 5.6 Sol = 2,800.50 credits [Model no longer shown]
Sonnet 5 = 2,509.3 credits [Model no longer shown]
Opus 4.8 = 2,487 credits [Model no longer shown]

GPT Sol 6 = 2,243 credits [Model no longer shown]

Opus 5 = 2,240 credits [Model no longer shown]

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits [Model no longer shown]

Interestingly, Sol 6.1 generated a 73 page report while Astra 6 was only 35! That’s 2 x the output. Stop and think about that, 73 detailed pages with a single prompt! Amazing!

I also compared it to its predecessor, GPT Sol 6 and found – https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260930-Cowork-GPT6-61Sol-Grok-eval.md

GPT 6.1 Sol is also the clear winner here as the results show. Again, an amazing improvement jump by my comparison from only Sol 6.0 to Sol 6.1! Where is this going to end?

The summary of the all the outputs produced is here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-File-paramters.md

while the results, including outputs are here:

https://github.com/directorcia/general/tree/master/Copilot/Comparisons

How much better can these models get? We’ll know shortly, I’ll bet.

Astra still wins but not on cost

image

With GPT 6 Sol now also available in Copilot, it is time for the LLM battle to continue and the results are in.

Astra 6 beats Sol 6 – https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260923-Cowork-GPT6Sol-Grok-eval.md

However, GPT 6 Sol is very cheap at 2,243 credits and that ranks towards the bottom of costs in Cowork as seen here.

Opus 5.5 = 23,494 credits

Astra 6 = 6,067 credits

Fable 5.1 = 3,453 credits

Fable 5 (Preview) = 3,387 credits [Model no longer shown]
Fable 5 (Copilot)(Preview) = 3,373.8 credits [Model no longer shown]

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits [Model no longer shown]

GPT 6 Sol = 2,243 credits

Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

I also decided to pit the recent Claude models against each other and found:

Opus 5.5 beats Fable 5 – https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260923-Cowork-F51vO55-Grok-eval.md

and it did so quite significantly if you look at the report. Opus 5.5 is clearly superior to Fable 5 according to my rudimentary benchmark. So it is really good but also really (really) expensive.

Remember, all the reports and created documents are here so you can judge for yourself:

https://github.com/directorcia/general/tree/master/Copilot/Comparisons

Let me know what you find.

I think we need a count back

image

I have been comparing all the LLMs as they become available in Copilot inside the different services in Microsoft 365. The previous winner was GPT 6 Astra, with the details here:

https://blog.ciaops.com/2026/09/17/astra-takes-the-trophy/

As is the world AI, it isn’t long before another new model becomes available, this time Opus 5.5, so I put it through the same test as all the other models and it produced this out which you can download yourself:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816/20260923-Cowork-Opus55.docx

As always I pitted the newcomer against the incumbent with the analysis report here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260923-Cowork-Opus55-Grok-eval.md

and is where the controversy is going to start. According to the ‘standard’ overall rating Astra scored 8.65 overall beating the newcomer Opus 5.5 with a score of 8.40. Close. However, if you dig a little further you see that Opus 5.5 score all 9’s for

– Evidence

– Detail

– Presentation

and was let down by

– argument

The overall score a weighted average in favour of argument. This helped Astra to pip Opus 5.5, but honestly I would suggest on review that Opus 5.5 is pretty much the equal of Astra 6 but a win is a win.

Ok, next point of amazement is that cost of Opus 5.5

Opus 5.5 = 23,494 credits

Astra 6 = 6,067 credits

Fable 5.1 = 3,453 credits

Fable 5 (Preview) = 3,387 credits [Model no longer shown]
Fable 5 (Copilot)(Preview) = 3,373.8 credits [Model no longer shown]

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits [Model no longer shown]
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

So, the same prompt with Opus 5.5 cost a whopping US$235! That is roughly 5 x the cost of Astra 6 and almost 8 x the price of Fable 5.1, even though Opus 5.5 is supposed to be more ‘cost effective’ than Fable 5.1 according to Claude.

I’ll have to run a report and compare Opus 5.5 to Fable 5.1 and see what differences are evident but at 8 x the price they’d wanna be MASSIVE!

I have also updated the summary report for all the documents created by the different models here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-File-paramters.md

GPT 6 Sol has also just been made available in Copilot so I’ll be testing that next.

SharePoint Is Not a Digital Dumping Ground for AI

image

I keep seeing the same mistake with AI projects.

An organisation decides it wants better results from Microsoft 365 Copilot, so everyone starts moving more information into SharePoint. Old project files, meeting notes, policies, spreadsheets, PDFs and assorted documents are copied across as quickly as possible.

The thinking seems to be simple: more data must produce better AI.

It does not.

If the information arrives without structure, ownership or context, you have not created an AI knowledge base. You have created a larger pile for the AI to search through.

Storage is not understanding

SharePoint can store an enormous amount of information, but storage alone does not tell Copilot what that information means.

Consider a document called Pricing-FINAL-v2-new.xlsx.

Is it the current price list? Is it an old draft? Which client does it relate to? Who approved it? When should it be reviewed? Can sales staff rely on it when preparing a proposal?

A person might work out the answer by opening the file, checking its contents and asking the colleague who created it. Copilot has to make the same judgement from the signals available to it.

That is where naming, metadata, content types, permissions and information architecture matter. They provide the surrounding context that helps Microsoft Search and Copilot retrieve the right material rather than merely finding something that looks related. Your existing SharePoint guidance identifies file names, titles, headings, metadata columns and page structure as important semantic signals for Copilot retrieval. [Copilot_Re…tion_Guide | Word], [Optimizing…t_Improved | Word]

Random information creates confident confusion

Imagine a finance library containing three expense policies.

One is current. One was replaced last year. The third was drafted for a proposed change that never happened. All three are readable, all three contain similar language and none has a meaningful status field.

Someone asks Copilot, “What can I claim when travelling for work?”

Copilot may produce a polished answer, but the underlying source environment has made the task unnecessarily uncertain. The problem is not that the AI needs a cleverer prompt. The problem is that the organisation has failed to identify which policy is authoritative.

A well-structured library changes this. The current policy can have a clear document type, owner, approval status, effective date and review date. Superseded material can be archived. Permissions can reflect the real business audience. A descriptive SharePoint page can provide the agreed answer and link to the supporting procedure.

The AI now has better choices because the business has done the work of describing its own information.

Structure is an AI investment

This does not mean every organisation needs a huge records-management project before using Copilot.

Start with the information that matters most. Pick one recurring business process, such as proposals, client reporting, employee onboarding or policy enquiries. Give it a clear SharePoint site or library. Remove duplicates. Use consistent file names. Add a small number of useful columns such as client, document type, owner, status and review date.

Then test real questions through Copilot.

You can also use Copilot in SharePoint to help generate summaries and populate descriptive fields, reducing some of the manual effort involved in improving existing libraries. I demonstrated this during a recent webinar by having SharePoint Copilot create a document summary and save it into the library’s Summary field. [Need to Kn…ugust 2026 | Meeting], [Need to Kn…July 2026 | Meeting]

The lesson is straightforward.

AI does not reward organisations simply for collecting more data. It rewards organisations that make their information easier to identify, trust and use.

Before asking whether your business has enough data for AI, ask a better question:

Have we made it clear what our data actually means?

CIA Brief 20260919

image

Industry News — AI Models & Research

  • Early Gemini 4 Pro Test Stuns with 3D Models and Code — Arena.ai testers appear to be hitting what many think is Google’s Gemini 4 Pro (labelled gemini-3.8-flash), producing detailed 3D renders, vector art, and full monographic sites. Side-by-side checks reportedly beat GPT-6 Astra on visual fidelity, with rumours of a very large context window and a possible October release.
  • On the Navier–Stokes Millennium Prize Problem — OpenAI published a post on work related to the Navier–Stokes Millennium Prize Problem. It’s a research-facing item that sits at the intersection of advanced maths and AI capability claims. Worth a read if you follow frontier model research narratives.

Industry News — AI Safety & Policy

Announcements — Copilot & Microsoft 365

Industry News — Security

After hours

SpaceX – “Holy Grail Of Rocketry” Documentary 4K – https://www.youtube.com/watch?v=mQCtXRtRuBc

Editorial

If you found this valuable, the I’d appreciate a ‘like’ or perhaps a donation at https://ko-fi.com/ciaops. This helps me know that people enjoy what I have created and provides resources to allow me to create more content. If you have any feedback or suggestions around this, I’m all ears. You can also find me via email director@ciaops.com and on X (Twitter) at https://www.twitter.com/directorcia.

If you want to be part of a dedicated Microsoft Cloud community with information and interactions daily, then consider becoming a CIAOPS Patron – www.ciaopspatron.com.

Watch out for the next CIA Brief next week

Astra takes the trophy

image

In the last round of the Copilot LLM challenge Fable 5.1 came out on top. You can find that here:

https://blog.ciaops.com/2026/09/02/we-have-a-new-llm-in-copilot-winner/

That was then and this is now. Less than a month later a new model from OpenAI, GPT 6 Astra has become available in Copilot (Cowork specifically). I therefore pitted Astra 6 against the reigning champion Fable 5.1. The result is, unsurprisingly, that we have a new winner:

GPT 6 Astra

and you can see the results for yourself here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260917-Cowork-Grok-eval.md

with all the outputs here:

https://github.com/directorcia/general/tree/master/Copilot/Comparisons/20260816

Interestingly, there was only a small improvement last time when Fable 5.1 pipped Opus 5 by 8.7 to 8.5 in the overall score. This time however Astra won by a significant margin of 9.03 to 8.30.

Again, not unsurprisingly, where Astra lost out was on cost, almost doubling the cost of Fable 5.1 as you can see:

Astra 6 = 6,067 credits

Fable 5.1 = 3,453 credits

Fable 5 (Preview) = 3,387 credits [Model no longer shown]
Fable 5 (Copilot)(Preview) = 3,373.8 credits [Model no longer shown]

GPT 5.6 Sol = 2,800.50 credits
Sonnet 5 = 2,509.3 credits
Opus 4.8 = 2,487 credits [Model no longer shown]
Opus 5 = 2,240 credits

GPT 5.5 = 1,200 credits
GPT 5.6 Terra = 260 credits

So the latest model is the best (unsurprising). The latest model is also the most expensive (unsurprising).

I have also updated the summary report for all the documents created by the different models here:

https://github.com/directorcia/general/blob/master/Copilot/Comparisons/20260816-File-paramters.md

I’m sure there will be more model releases coming soon but my rudimentary testing certainly indicates they are improving, if somewhat more expensive each time.

 

You’re Using AI. Your Company Still Runs Through You.

I talk to founders every week who are already deep into AI. They’ve got ChatGPT open in another tab. Someone on the team has spun up agents. There’s a Power Automate flow humming in the background. On paper, they’re doing everything right.

Yet their week still looks the same. Hard calls still queue on their desk. Staff still ping them for answers that should live somewhere else. The calendar is still a wall of back-to-backs. The tools made them faster. The organisation barely moved.

That’s the gap I keep noticing. Personal speed is not the same as company capacity.

Faster you, same bottleneck

When a founder uses Copilot in Outlook to knock out replies, or asks Copilot in Word to reshape a proposal, they finish their own pile sooner. Good. Necessary, even. But if every material decision still waits for their nod, you’ve only compressed the work that was already on their plate. You haven’t changed who the work flows through.

I see it in Teams more than anywhere else. A channel fills with questions that only the founder can settle. Meeting after meeting exists so they can “just confirm” something. Copilot can summarise those meetings and draft the follow-ups. Useful. It still doesn’t stop the next wave of the same questions landing on the same person.

The quiet cost is attention. Every interruption that only they can clear is a tax on the thinking the business actually needs from them.

Put the answers where the work happens

The shift that matters is boring, and that’s why people skip it. Stop treating AI as a private accelerator and start treating Microsoft 365 as the place your team’s answers already live.

When policies, pricing notes, and “how we do this here” sit in SharePoint or a well-kept Teams channel, Copilot can pull them for the next person who asks — without another ping to the founder. When that knowledge only lives in someone’s head or a buried email thread, no amount of prompting helps. The model can only work with what the organisation has bothered to store.

I’ve watched founders paste the same explanation into ChatGPT three times in a week, then wonder why their week still vanished. The better move was one clear page in SharePoint, shared with the right people, and a habit of pointing the team at Copilot in Teams instead of at their inbox.

Same idea with decisions. If the pattern of a decision can be written down — thresholds, who owns what, what “good enough” looks like — Copilot in Word or Loop can help a lead draft the first pass. The founder reviews the exception, not every ordinary case. That’s not abdication. That’s designing the company so ordinary work doesn’t need a royal signature.

Free the calendar by changing the meeting

Packed calendars are often a symptom, not the disease. People book the founder because the founder is still the only reliable source of context. Copilot in Teams can capture notes and actions, but the lasting fix is fewer meetings that exist only to transfer what should already be findable.

I’m watching founders who treat Copilot as part of how the team works — not just how they clear email — start to get real hours back. Not because a chatbot is magic. Because decisions and answers stop routing through one desk by default.

If your AI stack has made you personally sharper but left the company just as dependent on you, the next move isn’t another tool. It’s putting the knowledge, the thresholds, and the follow-through into Microsoft 365 where Copilot — and your people — can use them without waiting for you.

The Future of MSP Support Is Not Better Search–Get my FREE publication for help

image

Get you copy here – From Search to Intelligence Just enter $0 as the price, although a donation or a subscription to my newsletter would be appreciated.

I have watched plenty of skilled technicians work a support ticket with a browser full of tabs, two admin portals open, an old ticket on one screen and Microsoft Learn on another.

They eventually find the answer.

The problem is that the next technician often starts again from the beginning.

That is the limitation of search. Search can help me locate information, but it does not automatically connect what I found with the customer’s history, the evidence in the ticket, the conversation in Teams and the lesson buried in somebody else’s inbox.

For an MSP, that repeated reconstruction is expensive.

The real shift is from retrieval to reasoning

I do not believe Microsoft 365 Copilot should replace search, technical documentation or engineering judgement. Those remain essential. The change is in what happens around them.

An engineer investigating an Exchange Online mail-flow issue may still need a message trace, an NDR and the current Microsoft documentation. Copilot can help organise that material into known facts, missing evidence, competing explanations and a safe diagnostic plan.

That is very different from asking, “How do I fix this?”

The first approach exposes uncertainty. The second invites a confident answer before the diagnosis exists.

I want Copilot helping an engineer compare evidence, challenge assumptions, draft a clean escalation and turn the verified outcome into a useful knowledge article. I do not want it making untested production changes or weakening a security control because that appears to be the quickest path to closing the ticket.

Faster guessing is not progress.

A closed ticket should improve the next ticket

One of the biggest missed opportunities inside many MSPs is what happens after a problem has been resolved.

The engineer understands what occurred. The customer receives an update. The ticket is closed with a few hurried lines. Most of the reasoning then disappears.

Microsoft 365 Copilot gives an MSP a practical way to change that. Once the engineer has verified the resolution, Copilot can help produce the internal note, create a customer-friendly explanation and draft a reusable article for a permissioned SharePoint knowledge library.

The human still confirms the facts. The human still approves the wording. The human still owns the outcome.

The difference is that valuable knowledge no longer has to remain trapped in personal memory.

An MSP that makes knowledge capture part of ticket closure starts building organisational capability from ordinary support work. The next engineer begins with more than a search box and a vague recollection that someone saw something similar last month.

This is an operating-model decision

Buying licences will not create this result by itself.

I would begin with a small service-desk group and a few support categories where research, handover and documentation consume real time. I would change the ticket template so evidence, assumptions, tests, negative results and validation are recorded. I would review permissions before allowing Copilot to surface information more easily. I would then measure the effect using the MSP’s own baseline rather than somebody else’s productivity claim.

That produces something far more valuable than a collection of clever prompts.

It produces a support method.

The MSP opportunity is not merely to answer tickets more quickly. It is to make investigations more consistent, handovers more complete, customer communication clearer and verified knowledge easier to reuse.

Search helps me find an answer.

Intelligence helps my organisation retain what it learned.

That is the transition MSPs should be preparing for, and it is the focus of my new guide, From Search to Intelligence, which is now available for free (a denotation would be welcome and maybe subscribe to my email newsletter as well.)