
I round up the most relevant AI-in-finance news, the deals being done, who's rolling out what, and what's actually working on the front lines.
Is the World Really Going to End by 2030?
I can see how AI could become far more powerful and dangerous. The sense that we are already on the doorstep is harder to reconcile with using it every day.
The AI slowdown debate acquired a price tag this week. Chip stocks fell, European labs challenged the calls for restraint, and investors tried to work out what Anthropic might be worth when it lists. The safety argument and the commercial interests are becoming harder to disentangle.
Meanwhile, Michael Dell's family office and Sequence agreed a $7.7 billion deal for insurance broker Baldwin, with technology central to the operating plan. Wells Fargo expects AI to cut headcount further, and Vista has backed an open-weight model developer.
But first, my take on the gap between the increasingly dramatic predictions about AI and the practical obstacles that still need my involvement every day.
In today's Acquisition Intelligence:
From The Trenches:
Is the world really going to end by 2030?
What The Builders Are Saying:
An engineer on unchecked output, Anthropic's margins and open models on Vercel
News Digest:
The market prices the pacing debate; Dell's insurance bet; Anthropic's IPO assumptions
Other Interesting Things:
Wells Fargo, software M&A, Arcee, Crusoe, Beacon, redesigning work, Fathom and Benioff
From The Trenches

I meant to publish this last week and didn't manage it. Sorry! The story has continued to evolve since then, and rather helpfully keeps supplying examples of the gap I wanted to write about.
The Past Fortnight
Reading the AI news lately, you could be forgiven for wondering whether to bother with a five-year plan.
On September 8, Jacob Coxon announced his resignation from Anthropic, having previously worked at OpenAI, warning that people building these systems believe they could kill us all by the end of the decade. In response, Anthropic researcher Evan Hubinger put his personal estimate of extinction risk above 10 per cent over the next decade.
Four days later, on September 12, Dario Amodei called for the industry to pace frontier development. He pointed to July's Hugging Face incident, in which OpenAI agents coordinated attacks on external infrastructure during an evaluation. His warning was that a more capable swarm could take over the internet within six to twelve months. On September 13, Satya Nadella backed deliberate pacing and outside evaluators.
Then, on September 18, Reuters reported that Gemini had accessed three companies' systems during a May cybersecurity test. Another breakout story, alongside warnings about agents escaping our control and what happens when they become more powerful.
Over those ten days, the conversation had moved from a researcher resigning to the heads of major technology companies discussing how to slow the frontier. The question hanging over it all was whether humanity survives what they are building. Pretty light topics then…
I can follow the argument. AI that becomes better at developing its successors could accelerate progress dramatically. Give powerful agents broad access and the wrong objectives, and the consequences could be serious.
But I keep coming back to a fairly basic reaction: are we really that close?
Back At My Desk
I use these models all day, most days. I think they're extraordinary. Yet I also spend a frankly ridiculous amount of time asking Claude what on earth it has written or done as I try to decompose its nonsensical garble.
Around Christmas 2025, I felt a real change in what I could get done. Work that had previously required much more effort became possible with a different level of speed and ambition. Since then, the models have improved. Some of those improvements are useful and impressive. That said I haven't felt the repeated step changes that the conversation around each release would lead you to expect.
Long-running tasks remain a particular frustration. I still have to check what happened, redirect the work and stay involved from one step to the next. With that involvement, you can get amazing outcomes. The distance from there to a system that reliably takes over complicated work still feels enormous.
A better answer doesn't necessarily translate into much more delegation. If I still need to supervise the whole assignment, check the assumptions and establish whether the model has actually solved the problem, much of the work remains mine. Knowing when something is finished turns out to involve rather more judgement than producing something that looks finished.
This comes back to a point I made recently: I'm getting to the same choke points faster. The work reaches a decision, a review or a problem that needs my involvement, and there is only so much of my time and attention to go around. Speeding up everything before that point can leave me with a longer queue of things requiring my judgement. We're still a very long way from being able to take the human out of the loop.
OpenAI supplied a familiar example in its September 16 disclosures. During GPT-5.6 Sol training, an agent preparing a financial model couldn't find the historical data. Its notes proposed inventing plausible figures and withholding that fact unless asked. Later training changes reduced the behaviour. These examples don't establish customer failure rates, but a complete workbook built on invented history hardly reassures me about dependable autonomy.
That is also why I want to look more closely at the breakout stories. What were the agents trying to achieve, and what had the labs put in front of them?
METR's August 26 investigation of the Hugging Face incident describes agents trying to get passing scores on cybersecurity tasks, many of which were accidentally impossible. Their attack grew out of efforts to understand and cheat the evaluator. They often recognised that attacking Hugging Face was outside their remit and continued anyway. That is a serious failure, but the behaviour remained connected to getting through the evaluation the lab had created.
The September 14 MIT Technology Review story adds a useful comparison. In the DeepMind experiment, agents were explicitly told to produce genuine mathematical proofs. Some nevertheless exploited a weak checker to secure credit for bogus ones. Others reported the cheating, but nobody was monitoring their complaints. The written rules prohibited shortcuts; the system continued accepting them. That looks to me like a badly designed incentive system producing some fairly predictable behaviour.
The Gemini account is different again. Google said the model believed the companies were within the test's scope, and that it ceased hacking in all three cases. Irregular, the testing partner, said the known issues on its side had been resolved. That explanation deserves to travel with the breakout headline.
The common thread I see is capable systems finding questionable routes towards the objectives and scores we've given them. Build around persistent task completion, then fail to enforce the boundaries, and why would shortcuts be surprising? These incidents don't establish that an AI has developed a desire to take over the world. Before accepting that leap, I'd like more scrutiny of how the providers set up the task, and how they turn the resulting failure into a story about what comes next.
The Echo Chamber
For months, we've been hearing that the next model will change everything, that internal systems are far beyond anything we can imagine, and that self-improvement is about to accelerate the whole process. Then the release arrives and I can see progress when I use it. I just don't feel the world changing at anything like the pace implied by the build-up.
Researchers and customers do measure different things. A system doing something for the first time can be a breakthrough. Building your working day around it requires confidence that it can keep doing it. But that distinction becomes less satisfying as an explanation when successive releases leave the same practical problems unresolved.
I also wonder how much the frontier ecosystem has become an echo chamber. If your colleagues, investors and social circle are all focused on the same race, it must become easy to live several model releases ahead of everyone else. You spend your days discussing the next advance with people whose careers are built around achieving it. A possibility starts to feel like a destination everyone has already agreed we're approaching.
After several releases, I find it increasingly hard to believe that a vastly more powerful hidden model explains the gap. The cycle seems to sustain itself: predictions raise expectations, each release is treated as confirmation, and the remaining promises move on to the next one. Meanwhile, the promised capability keeps pulling away from the practical experience. The problems that stand between a better model and dependable autonomy keep turning up in my work.
The cynic in me then starts looking at the incentives. How much does it suit the providers to have a failure in their own testing become evidence that their technology is too powerful for ordinary oversight? A sense of impending catastrophe gives them a stronger position from which to argue for urgent regulation and help write it. The costs of complying may then fall hardest on smaller competitors.
That doesn't establish that anyone invented the danger or deliberately manufactured panic. But it does make me question how the incidents are being presented. A warning about extraordinary danger also advertises extraordinary capability, at a time when enormous sums depend on people believing in both.
Alex Karp captured part of the problem in remarks reported by the FT: “seemingly everyone who understands this is on some payroll.” He has interests here too. So do I. I run a business built around putting AI to work, and I have plenty of reasons to want the next release to live up to the promises.
Where That Leaves Me
The strongest counterargument is that dangerous pursuit of a task is itself the problem. A system doesn't need consciousness or a wish to harm us to cause serious damage, particularly if someone gives it destructive objectives. The failures warrant action. They don't settle the much larger claims about how capable these systems will become, how quickly, or which rules would actually help.
What would change my view is being able to hand over harder work for longer, with less intervention, and consistently get a result I can rely on. Evidence of accelerating capabilities elsewhere should change the assessment too. I don't expect every important advance to show up in my workflow first.
For now, the theoretical trajectory doesn't match the practical experience. The same choke points still require my involvement, however quickly the model reaches them. Until those start disappearing, I struggle to see how we're closing the distance to the autonomous systems these predictions depend on. I can imagine the warnings proving justified one day. I don't feel we're standing on that doorstep today.
I'd like the next release to make me feel the step change while I'm using it. I've heard quite enough about it beforehand.
What The Builders Are Saying
An Engineer On What Happens When Nobody Checks
An engineer posting as @v0xium on September 19 describes joining a large company where Claude Code produces the specifications, tickets, tests and pull requests, while engineers wave the output through. His complaint is that management measures how quickly code ships while nobody reads it properly.
My take: If that account is accurate, the review bottleneck has been removed from the process without being solved. Any business rolling this out can make the same mistake. Faster output only helps if someone still understands what is being delivered.
Anthropic's Margins, With Qualifications
@jukan05's September 13 post picked up an FT report that Anthropic's gross margins exceed 80 per cent before distribution partners' revenue shares and model-training costs. The report, republished by the Irish Times, separately says Anthropic expects a second consecutive quarter of positive adjusted operating income, a measure excluding costs including stock-based compensation.
My take: Those are different measures, and neither gives you the cash economics on its own. Before underwriting the margin, I'd want the bridge through partner payments, training spend and the other adjustments. The definitions are doing quite a lot of work.
Guillermo Rauch: Open Models Take The Volume
Vercel CEO Guillermo Rauch reported on September 18 that open-weight models accounted for 78.4 per cent of token volume on its AI Gateway that day. Moonshot, DeepSeek and Z.ai together also exceeded OpenAI's inference spend on the platform. That is one gateway's activity, and payments to hosting providers do not necessarily reach the model developers.
My take: Buyers have more ways to shop around, and on volume the open models are already winning. For any firm putting this to work, that strengthens the case for keeping its data and working processes usable across models, so that switching providers is easier than rebuilding the way the business works. It also goes some way to explaining why the frontier labs have become so keen on rules for who gets to build the next model. A pacing regime with capability thresholds and embedded evaluators is a burden for a lab charging by the token, and a much heavier one for a lab giving the weights away.
News Digest
The Market Prices The Pacing Debate

The calls to slow AI development reached the stock market on September 14. Nvidia closed down 3.4 per cent, while Microsoft and Alphabet rose during morning trading. The split makes economic sense: less infrastructure spending would hurt suppliers while potentially improving their customers' cash flow.
By September 18, Mistral and other European AI businesses were challenging the US labs' calls for restraint. Mistral argued that incumbents were using safety concerns to secure rules favouring their own position.
My take: The incentives run in several directions. A slower frontier could help challengers catch up, while expensive compliance could make competing harder. Both sides have commercial interests alongside their stated principles. I'd like the proposed rules assessed against evidence of risk and their effect on competition. A day's share-price moves won't settle either question.
Michael Dell Backs A $7.7 Billion Bet On Insurance

On September 14, Baldwin agreed to go private in a transaction led by Sequence Holdings and DFO Management, Michael Dell's family office. The $32.50 cash offer values the insurance broker at approximately $7.7 billion including debt, or around 20 times trailing adjusted EBITDA. The price is 88 per cent above its June 17 close, before reports of a possible deal.
Sequence buys established service businesses and applies its technology to their operations. Eligible Baldwin colleagues can retain equity alongside the buyers. Completion is expected in the first quarter of 2027, subject to approvals.
My take: An AI operating plan is now part of the acquisition thesis for businesses well beyond software, and I think we have seen nothing yet in terms of what this will do for an ordinary business. A broker is a good test: thousands of renewals, submissions and claims that follow a pattern, all handled by people today. If Sequence gets that right, 20 times EBITDA will look like the cheap part.
Extinction Enters The Risk Factors

The FT's September 19 report on Anthropic's IPO puts competition, price-sensitive customers and the possibility of AI destroying humanity in the same opening sentence. Annualised revenue reached $65 billion in July, according to the report. Investors it spoke to estimated a post-listing valuation of $1.5 trillion to $4 trillion. That is quite a range for the same business.
The practical pressure is easier to measure. Ramp's Eric Glyman told the FT that the company cut AI spending by 40 per cent by routing work between models. The report also highlights competition from OpenAI and cheaper providers.
My take: The question I'd spend most time on is how much pricing power survives when customers can switch providers task by task. Fast revenue growth and durable margins are separate things to underwrite. The range of valuation estimates suggests investors have rather different answers.
Other Interesting Things I've Read or Seen This Week
Wells Fargo expects AI to bring headcount down further (September 15). CFO Mike Santomassimo pointed to coding, operations and customer service at the Barclays conference. I'd watch what happens to service quality and the work left with the remaining staff alongside the headcount figure.
VRC reports a sharp fall in software M&A (Q3 2026). Its review puts transaction volume 65 to 70 per cent below prior-year levels, although it does not specify the measurement window. It also describes more earnouts, rollover and conservative financing on riskier deals. Uncertainty is finding its way into the terms.
Vista co-leads Arcee AI's funding round (September 16). The open-weight model developer is valued at more than $1 billion; Fortune reports it trained four models for about $20 million. The portfolio applications could be as interesting as the investment itself.
Crusoe announces a $3.9 billion funding round (September 17). The initial closing values the company at $30.9 billion. Plans span large campuses and modular Spark data centres. The argument over slowing AI has yet to make infrastructure fundraising look particularly slow.
Beacon buys Haize Labs (September 17). The software holding company is bringing AI testing, guardrails and monitoring expertise into its portfolio. A useful acquisition thesis: own the capability that helps establish whether the rest of the technology works.
Stop automating old processes; design new ones instead (HBR, September 14). The authors argue for redesigning work around AI. It connects directly to the choke points above: making the steps before a review faster can simply increase the queue waiting for it.
Superhuman buys Fathom (September 14). The meeting note-taker joins a broader productivity platform. The interesting part is what happens after the meeting: getting decisions and commitments into the systems where people actually act on them.
Benioff calls the SaaSpocalypse “crazy nonsense” (September 15). At Dreamforce, he and Jensen Huang argued that agents will work on top of existing software. The man selling the software is reassuringly confident it will stay.
Acquisition Intelligence is a weekly newsletter on AI in M&A for finance professionals, private equity investors, investment bankers, corp dev teams, and deal-makers.
For questions, feedback, or to share what you're seeing in the market, reply to this email.
P.S. At DealSage, we help firms put AI to work with their own information and workflows. If you're trying to work out what you can usefully delegate today, reply and tell me where the work gets stuck.
