The Local Agentic Future is Already Here — Just Not All at Once
Hi friend,
First, thanks for all the feedback on my hand-crafted newsletter 🤌🧑🍳 Why you should be excited about local LLMs too. I see you liked it, so I’m continuing in the same fashion, until someone gets tired of my typos.
For example, I wrote Composter 2.5, instead of Composer 2.5. I’m really into camping these days. We visit our cabin quite often during the summer and besides best local models, portable, composting toilets are also something that appear in my search history. Sorry for wasting your time with silly typos, I’ll do my best!
Back to the main topic, the local agentic future.
I’ve been seeing random signals on X lately and can’t stop connecting the dots.
Better consumer hardware, smaller models
About a month ago, Simon offered to build an H200 cluster and rent out the compute to a small group for a monthly fee. A practical way for regular people to share serious hardware without feeding OpenAI or Anthropic every month. A private, crowd-sourced cluster. Amazing idea.
But I didn’t give it too much attention a month ago for two reasons:
at that time the only local model I tried was some early version of ollama that sucked
since I was on a Claude Max this year, I knew what SOTA models can do but haven’t experimented much with Kimi K2.6, except on side projects, which isn’t the real thing
But if you read Why you should be excited about local LLMs too you know that I got stuck with Kimi K2.6 at work for a week, and suddenly I understood why Simon was so excited.
Then I saw this tweet and drew the conclusion:
We are moving towards local models.
Not all at once, we still need better hardware, more model optimisations… but the technology to create a smooth transition is already here.
And it was one of Cursor’s first selling points.
Model Routing
Their Auto mode looks at the task complexity and picks the right model.
In the near future we could have a smart-enough proxy locally that picks:
fast, local model for simple tasks: rename, code edits, grep
reasoning model for brainstorming, architecture work, etc
That’s the transition.
Dipankar Sarkar mentioned this in the comments on my previous article and I think this might actually be where we are headed.
And we don’t have to wait for consumer hardware to run massive reasoning models. We just need the router to get seamless and private.
I don’t think I’m the only one thinking in this direction.
In Andrei Jikh’s recent video China Is About To Pop The AI Bubble, around the 8:06 mark, Alex Karp, CEO and founder of Palantir highlights exactly this shift happening among big players:
“But what is happening among the most technical players is they’re saying, ‘I want something I own. This is my business. I want to own the GPUs. I want to own my data. I want to own the model. I want to control the alpha.’ ... instead of renting an AI from someone, instead you just download one, you run it on your own computers using your own data where no one can see it and no one can learn from it and also no one can shut it off.”
He notes that most businesses don’t even need the absolute cutting-edge models.
They need ones tuned for their specific use cases, running privately and reliably.
I think the shift from “everything in the cloud” to “mostly local” is starting.
What signals have you been seeing?
Hit reply, I’m reading everything.
- Akos





