top of page
  • Black Facebook Icon
  • Black Instagram Icon
  • LinkedIn

The most interesting AI story this week isn't about a chatbot. It's about your laptop.

a scatter plot data chart
Are hyperscalers doomed?

There's a Stanford and Together AI paper making the rounds, "Intelligence per Watt" (Saad-Falcon et al.). Investors have seized on it as a "the hyperscalers are doomed" story.


I read the actual paper, and the real finding is quieter, more careful, and to me more interesting than the headline and gave me genuine hope for a more eco-friendly future.


The researchers introduce a metric they call intelligence per watt: task accuracy divided by unit of power. Then they test how much AI work can move off the cloud and onto the machine in front of you. I've gotta tell ya, folks. This makes my nerdy little heart so happy. Their definition of "small" is specific: local models of 20 billion active parameters or fewer, running on local accelerators like an Apple M4 Max.


What they found:


Local models already handle 88.7% of single-turn chat and reasoning queries, the everyday stuff, which is most of what any of us actually asks. Not uniformly, though. Coverage runs above 90% in creative work but around 68% in technical fields. And it's moving fast. The share of queries that could be serviced locally rose from 23.2% in 2023 to 71.3% in 2025, and intelligence per watt improved 5.3 times in two years.


Here's the part the "cloud is dead" crowd skips. The paper isn't arguing for replacement. It's arguing for redistribution. The authors draw the analogy themselves: the way computing work moved from data-center mainframes to personal computers once efficiency, not raw power, met people's needs. And the biggest savings come from a hybrid of local and cloud, not from ditching the cloud. Smart routing between the two cut energy by about 80%, compute by 77%, and cost by 74% versus cloud-only. They're also honest about the catch. On identical models, local hardware still trails cloud accelerators by at least 1.4 times on efficiency. The win isn't that your laptop is a better engine. It's that a right-sized model, run close to home, with the ability to govern tightly, and whith less environmental impact, is often enough for your needs.


And this is the part that gives me actual hope. The environmental cost of the current path isn't abstract. Every new hyperscale data center is a draw on a real power grid, and in a lot of places, real drinking water to keep it cool. The communities next door are already pushing back, and, in my opinion, they're right to. What this paper points to is a future where we don't have to keep building our way out of demand with more concrete, more megawatts, and more strain on the people who happen to live nearby. If most of the everyday work can run on hardware we already own, then the greenest data center is the one nobody has to build. For someone who cares about both good data and the world it runs in, that's a genuinely different path forward.


I don't lead an investment portfolio. I lead an analytics team in a regulated environment, where "just send the data to the cloud" is never a casual sentence. So the number that stops me isn't the cost curve. It's where the work happens.


Three things this actually changes, if it holds:


The trust problem gets smaller. Governing data in a regulated setting is, in large part, governing where it goes. A capable model that runs locally and never phones home isn't just cheaper. It's a fundamentally easier thing to trust. In my world that's not a footnote. It's close to the whole game.


The "bigger is always better" story gets complicated. Two years of being told the answer is more: more parameters, more compute, more data center, more water, more community outrage. This says the answer is often enough, close to home. That matches what I already believe about analytics teams. The win is rarely the biggest model. It's the right-sized one, run where the work lives.


And the humans matter more, not less. When a genuinely capable model runs on everyone's laptop, the differentiator stops being access to the tool. It becomes knowing which question to ask and whether to trust the answer. That's judgment. That's a human skill. It always was.


The honest caveat: this is one paper, and the version I read has been revised several times since November. The responsible move is to watch whether the results replicate, not to treat a single preprint as settled. But the direction is credible, and it points somewhere hopeful. Not a handful of giants owning all the intelligence in a few enormous buildings, but capable tools running quietly on the machines regular people already own.


The future of AI might be less spectacular than the headlines promised, and a lot more human.


What's your read? Does small, local, and boring win? I'd rather hear from people running this in the real world than from people modeling it from the outside.

 
 
 

Comments


CONTACT 

ADDRESS

North Haven bb

The 06473 for life!

CONTACT ME

OPENING HOURS

I go to bed early but text or email anytime

Feel free to call prior to like 8:30 p.m. - use your best judgement and make good choices!

Thanks for submitting! I'll reach out soon :)

Powered by me :) 

  • Facebook
  • Instagram
  • Linkedin
bottom of page