The AI investment story has been built around a simple assumption: as artificial intelligence becomes more capable, demand for computing power will soar, driving enormous spending on Nvidia (NASDAQA:NVDA) GPUs, datacentres, networking equipment and electricity.
But new Stanford research raises a potentially disruptive alternative.
Instead of every AI request travelling to a hyperscale datacentre, increasingly capable small language models (SLMs) could run directly on PCs, smartphones and edge devices. The attraction is obvious: dramatically lower latency, less bandwidth, potentially lower energy consumption and greater privacy.
Stanford’s Intelligence per Watt study found that local models with 20 billion or fewer active parameters could accurately answer 88.7% of a representative sample of one million real-world single-turn chat and reasoning queries. Local intelligence efficiency improved 5.3-fold between 2023 and 2025, while the proportion of queries local models could handle rose from 23.2% to 71.3%.
Joachim Klement, head of strategy, economics, and ESG at Panmure Liberum, has taken the argument further, suggesting that if this trend continues, investors may be substantially over-estimating future demand for datacentre infrastructure.
If this is true, the hyperscalers are toast – Klement on Investing
That is an important warning — but not yet a reason to abandon Nvidia or the hyperscalers. It is, after all, just one informed opinion among many. There is a credible and increasingly explicit counter-argument to the thesis that SLMs, more efficient open-weight models and cheaper inference will eventually reduce the need for large-scale LLM training and hyperscale datacentre capacity.
The strongest is that lower inference costs increase the number of AI workloads enough to offset — or even exceed — the compute saved per task. This is essentially the Jevons paradox/rebound effect
At this point in time, the more likely outcome is a hybrid AI economy, in which simple and latency-sensitive inference moves towards devices while frontier models remain in the cloud for complex reasoning, agentic AI and large-scale workloads.
AI expert views snapshot
| Analyst / expert | Organisation | What they argue | Implication for AI infrastructure |
| Stephen Byrd | Morgan Stanley | Cheaper/more efficient models should stimulate more AI use rather than reduce compute demand | Bullish for GPUs, data centres, power |
| Dave Chen | Morgan Stanley | AI efficiency gains will increase the industry’s total addressable market through the Jevons effect | Bullish for infrastructure |
| George Lee | Goldman Sachs | Cheaper models could trigger more pre-training because more organisations can afford to build models | Bullish for compute |
| UBS CIO team | UBS | Lower token costs should drive greater demand for agentic AI, more than offsetting lower compute per task | Bullish for datacentres |
| Jensen Huang | Nvidia | AI efficiency will expand usage dramatically rather than eliminate infrastructure demand | Very bullish |
| Satya Nadella | Microsoft | Lower AI costs make AI more economically useful and expand demand | Bullish |
| Christophe Fouquet | ASML | Greater AI efficiency ultimately enables greater AI consumption | Bullish for semiconductor equipment |
| Demis Hassabis | Google DeepMind | Compute remains a critical constraint as models become more capable | Bullish for hyperscale infrastructure |
| Sam Altman | OpenAI | AI compute requirements will continue to grow enormously as capabilities and applications expand | Very bullish |
| Dario Amodei | Anthropic | Demand for compute will continue rising as AI capabilities scale | Bullish |
| Elon Musk | xAI/SpaceX | AI compute requirements are rising sufficiently rapidly to justify enormous new capacity | Very bullish |
You might argue that hyperscaler bosses will inevitably back the billions of dollars being invested in datacentres today, which perhaps makes Morgan Stanley’s the strongest Wall Street evidence. Morgan Stanley’s Stephen Byrd and Dave Chen have both done a lot of work in this area and it is particularly useful because it directly addresses the ‘efficient models = less compute’ argument.
The Jevons paradox
Morgan Stanley’s Stephen Byrd recently said that some investors worry better efficiency means less computing demand but argued the opposite: lower costs encourage more users, more frequent use and more sophisticated applications, explicitly characterised as the Jevons paradox.
Morgan Stanley’s economics are particularly compelling for investors. It estimates that an enterprise AI task can generate around $55 of economic value versus only $2-$5 of token cost. If that relationship holds, a substantial reduction in inference costs doesn’t necessarily cause customers to spend less: it gives them an incentive to deploy AI in many more workflows.
The bank also estimates that hyperscalers could increase available power capacity from roughly 30GW in 2025 to around 120GW by 2028.
That’s a particularly powerful statistic from an SLM-vs-LLM investment thesis standpoint.
Dave Chen’s research suggests efficiency could actually expand the AI TAM, or total addressable market. Morgan Stanley’s Chen made the argument even more explicitly in 2025:
‘Recent AI advancements will harness the power of the Jevons Paradox’
Why SLMs could be a big deal
The first advantage is ultra-low latency.
A local AI assistant does not need to send a request to a remote datacentre, wait for processing and receive the response. For applications such as robotics, industrial automation, autonomous systems, gaming and real-time voice assistants, eliminating that network round trip can be highly valuable.
The second is bandwidth.
AI is increasingly generating and processing large amounts of data. Analysing a video stream, sensor feed or personal document locally can avoid continuously uploading that information to the cloud.
The third is energy efficiency.
Stanford’s research suggests local inference is becoming substantially more efficient, while the research team’s work on hybrid local/cloud inference has demonstrated major potential cost and energy savings.
The fourth is privacy.
Sensitive information can remain on the device rather than being transmitted to a third-party cloud.
The SLM economic proposition
Device → local model → answer
rather than:
Device → network → hyperscaler → network → answer
That does not eliminate computing demand. It potentially redistributes it.
But LLMs still have a major advantage
The strongest counterargument is capability.
Frontier LLMs remain better suited to complex reasoning, long-context analysis, sophisticated coding and agentic applications involving multiple tools and steps.
Stanford’s research does not demonstrate that a small model running on a laptop can replace the most powerful frontier model in every circumstance.
Indeed, the researchers’ own work points towards cooperation between the two.
Their Minions research showed that local models can handle subtasks while a frontier model performs higher-level planning and aggregation. One configuration achieved 97.9% of GPT-4o’s performance at 5.7 times lower cost, illustrating how hybrid architectures can reduce the amount of work sent to the cloud.
That is potentially the most important investment conclusion.
SLMs may not kill LLMs. They may reduce how much work LLMs have to do.
What happens to Nvidia?
Nvidia remains the most obvious test case.
The bull case for Nvidia is that AI compute continues expanding so rapidly that both datacentre and edge demand grow. Nvidia is already pushing into local AI with its RTX Spark platform.
Its CEO Jensen Huang has argued that inference efficiency is increasingly critical because hyperscalers are power constrained: more performance per watt effectively means more revenue-producing tokens from each unit of electricity and infrastructure.
But there is a SLM risk.
If routine inference migrates from expensive datacentre GPUs to cheaper consumer GPUs, NPUs (Neural Processing Unit – specialist chips built specifically to run machine learning and AI inference tasks with high energy efficiency on local devices) and dedicated edge accelerators, the number of AI queries could rise while the dollars of datacentre compute required per query fall.
That would be a problem for the current investment model.
Reuters described Nvidia’s AI-PC push as a ‘high-stakes bet’ on a market where demand beyond developers and specialist users remained unproven.
For Nvidia investors, therefore, the crucial question is not whether local AI succeeds.
It is whether Nvidia can capture the local-AI opportunity without sacrificing too much of its exceptionally lucrative datacentre economics.
The stocks: who wins if SLMs take share?
The table below ranks major AI-related stocks according to their potential positioning in a world where SLMs take a much larger share of inference.
| Rank* | Stock | SLM positioning | LLM/datacentre positioning | SLM scenario |
| 1 | Apple | Excellent | Moderate | Major winner |
| 2 | Qualcomm | Excellent | Low/moderate | Major winner |
| 3 | Dell Technologies | Strong | Strong | Winner |
| 4 | TSMC | Excellent | Excellent | Winner either way |
| 5 | Nvidia | Strong but changing | Exceptional | Mixed / high upside |
| 6 | AMD | Strong | Strong | Potential winner |
| 7 | Broadcom | Moderate | Exceptional | Mixed |
| 8 | Microsoft | Strong | Exceptional | Hybrid winner |
| 9 | Alphabet | Strong | Exceptional | Hybrid winner |
| 10 | Amazon | Moderate | Exceptional | More exposed to cloud |
| 11 | Meta | Strong | Strong | Mixed-positive |
| 12 | Arm Holdings | Exceptional | Moderate | Potential SLM beneficiary |
| 13 | Super Micro Computer | Moderate | Strong | More exposed to datacentre cycle |
| 14 | CoreWeave | Low | Exceptional | Highest SLM downside risk |
*Ranking is Sharesify’s scenario analysis, not an investment recommendation.
Why Apple and Qualcomm stand out
Apple (NASDAQ:AAPL) could be one of the most interesting beneficiaries because it already controls the entire device stack: silicon, operating system and hardware.
The company can optimise models around its own chips and deploy AI directly across iPhones, Macs and other devices.
Qualcomm (NASDAQ:QCOM) offers a purer SLM exposure. Its Snapdragon X platform is explicitly targeting on-device AI, and Qualcomm says local agentic AI applications are already being demonstrated on Snapdragon-powered PCs.
IDC research cited by Qualcomm suggests 75% of US organisations surveyed expected to deploy next-generation AI PCs by 2026, while a significant proportion expected more than half of their PC fleets eventually to become AI-enabled.
Dell (NYSE:DELL) is another interesting beneficiary because it can monetise the physical transition from conventional PCs to AI PCs and workstations. Dell itself describes enterprise AI as moving towards a more distributed model, with intelligence moving closer to users.
TSMC may be the ultimate pick-and-shovel winner
For investors who do not want to decide whether SLMs or LLMs ultimately win, TSMC (NYSE:TSM) may be the most interesting compromise.
Advanced AI processors increasingly require leading-edge manufacturing whether they are designed for datacentres, PCs or smartphones.
That means a shift towards local AI could change the mix of semiconductor demand without necessarily destroying semiconductor demand itself.
TSMC therefore potentially benefits from both scenarios.
Arm Holdings (NASDAQ:ARM) is also worth watching because edge AI increases the importance of highly efficient CPU architectures across smartphones, PCs, vehicles and embedded devices.
SLM vs LLM/datacentre revolution
| 🐂 SLM bull case | 🐂 LLM/datacentre bull case |
| Local models rapidly close capability gap | Frontier models remain substantially better |
| 70–80%+ of routine inference migrates locally | AI workloads become increasingly complex |
| Ultra-low latency becomes essential | Network latency becomes less important |
| Privacy drives on-device adoption | Enterprises prefer centralised security |
| AI PCs become mainstream | Consumers resist expensive hardware upgrades |
| NPUs become standard silicon | Cloud GPUs remain more capable |
| Bandwidth costs encourage local processing | Connectivity becomes faster and cheaper |
| Edge AI cuts cloud inference costs | AI usage grows faster than local substitution |
| Apple/Qualcomm/Arm benefit | Nvidia/hyperscalers retain pricing power |
| Datacentre capex expectations fall | AI capex continues accelerating |
The hidden problem with edge AI: security
Investors should not assume that local AI automatically means safer AI.
The advantage is that sensitive information can remain on a device. The disadvantage is that organisations must secure potentially millions of endpoints rather than a relatively concentrated number of datacentres.
That creates new risks:
- stolen or compromised devices
- malicious model modification
- insecure software updates
- model extraction
- prompt injection
- local data leakage
- compromised edge agents
- inconsistent security patching
- supply-chain attacks
In other words, local AI potentially improves data sovereignty while expanding the attack surface.
That could create a new market for endpoint security, identity management and AI-specific cybersecurity.
The datacentre threat may be overstated
There is another reason not to rush into an ‘SLMs kill Nvidia’ thesis.
The hyperscalers are still spending extraordinary sums.
Microsoft (NASDAQ:MSFT), Alphabet (NASDAQ:GOOG), Amazon (NASDAQ:AMZN) and Meta (NASDAQ:META) are collectively expected to spend hundreds of billions of dollars on infrastructure in 2026, with estimates around $725bn depending on guidance assumptions.
And recent results have eased some investor concerns. Reuters reported on 17 August that strong Microsoft and Amazon results had helped reassure investors that cloud growth remained robust and AI infrastructure demand was still strong.
Microsoft Q4 FY2026: AI investment finally starts paying off
Amazon Q2 2026: AI finally delivers the proof investors wanted
The question is therefore one of incremental economics.
If AI usage grows 10-fold but local models absorb 70% of inference, hyperscalers could still experience enormous growth — but potentially not enough to justify the most aggressive assumptions embedded in today’s datacentre investment cycle.
That distinction is critical.
Investor verdict
SLMs are a genuine threat to the assumption that every additional AI workload requires additional hyperscale computing. But the evidence does not yet support abandoning LLMs, Nvidia or the datacentre trade.
The most credible scenario is a three-layer AI economy:
1. Local SLMs — fast, cheap, private, everyday inference.
2. Edge AI — real-time industrial, automotive, robotics and enterprise applications.
3. Frontier LLMs — complex reasoning, agentic AI, research and large-scale computation.
For UK retail investors, the implication is that the AI opportunity may be broader than Nvidia and the hyperscalers.
Apple, Qualcomm, Dell and Arm could benefit from the migration of intelligence towards devices. TSMC provides exposure to both sides. Nvidia and AMD (NASDAQ:AMD) could participate in both datacentre and edge computing, although investors must watch whether local inference eventually cannibalises their highest-margin workloads. Broadcom (NASDAQ:AVGO) remains strongly positioned if hyperscaler custom silicon continues to expand.
Meanwhile, Microsoft, Alphabet, Amazon and Meta have an important advantage: they can potentially participate in both architectures, using local AI to reduce inference costs while retaining cloud models for demanding workloads.
The biggest SLM loser would therefore not necessarily be a particular chipmaker.
It could be the assumption that AI’s future requires unlimited growth in centralised compute.
Stanford’s research suggests local AI is becoming substantially more capable and efficient. Its local models accurately handled 88.7% of the study’s single-turn queries, while local intelligence efficiency improved 5.3x in two years.
That is not evidence that datacentres are obsolete.
It is evidence that the AI infrastructure map is changing.
For investors and the AI investment story, that may be the more important story.
Disclaimer: The author Steven Frazer has a personal interest in Nvidia and Broadcom.
You might also like:







