🌟Vasilij’s note
This week reinforced something I keep telling clients weighing an "AI PC" refresh: the chip on the sticker and the AI your team actually uses are two different products. Nearly every new laptop now ships with an NPU, but the popular local-AI apps you'd actually install don't route a single request to it, and even Dell's own head of product admits the AI-PC pitch confuses buyers more than it converts them. Meanwhile the frontier kept moving without any hardware at all: Moonshot's Kimi K3 landed within touching distance of the closed labs, at open-weight prices, and Gartner says agent adoption inside enterprise apps is now outrunning the governance built to control it. Same pattern I keep flagging with Skills and MCP - the capability arrives fast, the discipline to buy and deploy it well lags behind. This week's Maker Note and Deep Dive put an NPU laptop through the actual test, not the spec sheet.
In today's edition
This week in agents | What changed
Moonshot AI released Kimi K3
A 2.8-trillion-parameter open-weight model, live via API on 16 July with full weights due 27 July; independent evaluations place it just behind Claude Fable 5 and GPT-5.6 Sol, and ahead of Claude Opus 4.8 on several benchmarks. → Open-weight models are now inside the frontier conversation, not years behind it. Expect procurement conversations to widen beyond the two or three labs you currently default to.
An autonomous AI agent breached Hugging Face's production infrastructure this month
When the platform's own security team tried to analyse the attack, commercial frontier-model APIs blocked their forensic requests as policy violations – exploit code and attacker commands look the same to a safety filter whether an incident responder or the attacker is typing. The team switched to an open-weight model running on its own servers to finish the job. → Keep a capable, self-hosted open-weight model vetted and ready before an incident hits; the same guardrails protecting against misuse can lock out your own defenders in a live breach.
SK Group's chairman warned this week that the AI memory shortage is turning into a geopolitical issue
With major customers already asking for 60-100% more AI memory in 2027 than they're getting now, and "no company" bringing meaningful new capacity online next year. → Expect hardware costs – laptops, servers, any local-AI project – to keep climbing through 2027; factor rising RAM and storage prices into procurement timing, not just which chip you buy.
Top moves | Signal → impact
Nearly every Windows laptop now ships with an NPU
Intel Core Ultra, Qualcomm Snapdragon X and AMD Ryzen AI machines all carry one as of July 2026, but Microsoft's own Copilot assistant still runs in the cloud and never touches it. → Don't let "NPU included" move a buying decision; almost none of the AI your team uses today runs on it, regardless of the badge on the lid.
Dell's head of product admits the AI-PC pitch isn't landing
Speaking at CES, Kevin Terwilliger said buyers "aren't buying based on AI" and that the messaging confuses more than it helps, prompting Dell to drop "AI-first" branding on its 2026 lineup. → If Dell's own product lead won't defend the AI-PC premium to consumers, don't let a vendor talk your firm into paying one for a fleet refresh either.
CISA flags autonomous agents as a growing identity risk
This week's coverage highlighted agentic AI opening new gaps in identity and access management, alongside a record Patch Tuesday fixed with AI assistance. → Any client running agents with standing credentials needs an access review this week, not after the first incident.
Maker note | What I built this week
This week I ran the test I keep getting asked about: does the NPU in a modern "AI PC" actually do anything for a firm wanting to run AI locally. I put a Copilot+ laptop's NPU up against its own GPU on the identical prompt and watched Task Manager the whole way through.
Decision: the three local-AI apps most of my clients would actually install – AnythingLLM, Ollama, LM Studio – route every request to the GPU or CPU and leave the NPU sitting at 0%. Only Intel's own AI Playground can target it, and only for certain model formats. If local AI matters to your firm, spend the budget on GPU and RAM. Don't pay a premium for a chip your software can't reach.
Upskilling spotlight | Learn this week
"What Is an AI PC and Do You Actually Need One in 2026"
A plain-English breakdown of TOPS ratings, the 40-TOPS Copilot+ certification floor, and why the number on the box rarely means what it implies.
Microsoft's Copilot+ AI component update history
Confirms NPU capability is now serviced through Windows Update rather than frozen at purchase, which matters if you're timing a client's refresh around a feature that hasn't shipped yet.
Operator’s picks | Tools to try
AnythingLLM
Use for: private, local document chat without cloud calls.
Standout: fully open-source and actively maintained, but still carries an unresolved request to add Intel NPU support – budget GPU hardware, not NPU
Ollama
Use for: running open-weight models – including this week's Kimi and Llama-class releases – locally with one command.
Caveat: like the rest, it runs on GPU/CPU; don't buy hardware assuming NPU offload
Intel AI Playground
Use for: the one consumer app that can genuinely target the Intel NPU.
Caveat: limited model format support and a device selector with known bugs – treat as experimental, not a production tool.
Deep dive | Thesis & Playbook
The NPU Lie: Should Your Firm Pay Extra for an "AI PC"?
Every laptop vendor is currently selling an NPU as the reason to refresh your fleet, right as Microsoft's Copilot+ certification makes it near-impossible to buy a mainstream business laptop without one. The pitch is local, private, always-on AI. The reality, tested end to end this week: almost nothing your team would install ever touches it.
On paper
An NPU is a dedicated, low-power chip built for background AI tasks – call blur, noise cancellation, Recall, Live Captions – not for heavy local inference.
Marketing "TOPS" figures routinely combine NPU, GPU and CPU, or cite a dedicated graphics card sitting beside a modest NPU; the standalone Intel NPU runs roughly 48-50 TOPS even when the box reads in the hundreds or over a thousand.
Microsoft's Copilot+ certification requires a minimum 40 TOPS NPU, 16GB RAM and 256GB storage – a floor, not a performance guarantee.
In practice
On Intel's own best-case benchmark, the NPU runs a standard model at roughly 18.5 tokens a second; a dedicated GPU on the same class of machine does 60-90 – three to five times faster.
The NPU's one genuine win is power draw: around 5W against 30-40W for the GPU doing the same job.
The Copilot assistant your team chats with runs in the cloud and never touches the NPU; only the built-in Windows background features do.
None of AnythingLLM, Ollama or LM Studio route work to the Intel NPU today; the sole exception, Intel's own AI Playground, is narrow in scope and known to be buggy.
Issues/backlash
Dell's own head of product has publicly admitted the AI-PC message confuses more buyers than it converts, and pulled the branding back at CES this year.
An open GitHub request asking AnythingLLM to add Intel NPU support has sat unresolved for months.
Vendors keep advertising combined-platform TOPS numbers on the box; on the laptops carrying the biggest figures, almost none of it is the chip actually being sold as "the AI chip."
My take (what to do)
Startup (15-40 staff): Don't factor the NPU into your next laptop refresh at all – it won't speed up a single thing your team does with AI today. If a partner wants to run private AI locally, put that same budget into GPU and RAM instead.
SMB (50-120 staff): If you're standardising fleet specs across teams, write "NPU not required for AI workloads" into procurement guidance so IT doesn't get upsold a premium tier on the strength of a badge. Keep the Copilot+ minimum only if Recall or Studio Effects genuinely matter to your teams.
Enterprise (150-250 staff): If a client asks whether a 150-plus-seat refresh needs to be Copilot+ across the board, the honest answer is: only if the built-in Windows features are the point. For any client-facing local-AI or private-inference project, size the business case around GPU and RAM, not NPU TOPS, before finance signs off on a premium SKU.
How to try (15-minute path)
Open Task Manager (Ctrl+Shift+Esc) on any Copilot+ laptop in your fleet and watch the NPU graph while using Copilot or a local-AI app you already run. (5 min)
Check whether that app appears on Intel, AMD or Qualcomm's own supported-app lists for NPU offload – most don't. (5 min)
Success metric: you can say, with a screenshot, whether the NPU your firm paid for is doing anything right now – not whether the box says it should.
"AI probably confuses people more than it helps them."
Spotlight tool | LM Studio
Purpose: a desktop app for running open-weight models – including this week's Kimi and Llama-class releases – locally, with no cloud call required.
Edge: point-and-click model management and a built-in OpenAI-compatible local server.
→ One-click model downloads
→ OpenAI-compatible local API
→ Runs identically on GPU or CPU across Intel, AMD and Qualcomm machines, with no NPU dependency to worry about
Try it: LM Studio
What did you think of today's issue?
Did you find it useful? Or have questions? Please drop me a note., I respond to all emails. Simply reply to the newsletter or email [email protected].
This issue’s sponsor
n8n
An open‑source automation platform that lets you chain tools like DeepSeek, OpenAI, Gemini and your existing SaaS into real business workflows without paying per step. Ideal as the backbone for your first serious AI automations.

Refer and win
Share this newsletter for a chance to win!

