Why you should be excited about local LLMs too
Hi friend,
it’s summer, if you are in the EU you probably noticed by the heat wave lol, so first things first: please make sure you are drinking enough water.
I’m back from my vacation and I have two thoughts to share with you. But before a little bit of context.
Half a year ago when Agentic coding was on fire I wrote How much are you really paying for AI engineering? It was way before big writers picked up the topic.
By mid March everyone started to question, why I’m getting so much usage for so little money? This can’t last forever right?
Avoiding the Lock-in
The generous token subsidies made people addicted. We did the simplest tasks with AI because it was so easy to just dump our thoughts into an input box. Then tokenmaxxing became a thing, and pricing increases followed. I saw this coming, and did a few things:
I allocated a monthly budget to play with different tools instead of digging myself deep into a specific one
I created skills and recreated agentic workflows that were previously Claude Code specific in other environment
Basically I tried to safeguard myself against a lock-in that has eventually arrived.
This strategy paid off well since none of my workflows are harness-specific, and I’m working with my $20 Cursor sub, Grok Build or Claude Code enterprise at work.
Cursor is an insane offer especially with Composer 2.5 which is really cheap and efficient. I’ve been using it to build two apps I’m going to announce soon. Here’s a referral link from me where you get 50% off at Stripe checkout for Pro, Pro+, or Ultra. I get API credits to build more. They recently rolled out an iOS app too, so now there’s no excuse to not build all day.
But avoiding the lock-in is beneficial both short and long term:
it forces you to build and learn to use universal tools that work across harnesses
it’ll work well with whatever model you end up using in the future
Local LLMs
Now, full disclosure, the last model I used on my machine a year ago was some version of ollama and it was terrible.
But let me tell you something.
When I returned from vacation three things happened:
my company rolled out Claude Code enterprise
I didn’t complete the mandatory training (I was on vacation) and was left out from the rollout that week
I used up all my copilot credits before I left
Circumstances forced me to connect to our LiteLLM setup, where we run Kimi K2.6 on our own cloud machines.
A local Kimi K2.6 is crazy good in an enterprise setting for one reason: in our huge codebase with dozens of microservices, legacy stuff lying around, Claude just hasn’t been much of a help when I came to brainstorming. I wrote about this in I Tried to One-Shot My Entire Sprint With AI and Failed.
So where SOTA models would shine, the brainstorming is done by us engineers, and even though our plans still have some rough edges, Kimi K2.6 is not totally dumb to help figure those out.
The Composer 2.5 I mentioned above that I use for all my side-work? It’s built on Kimi K2.6.
Where I’m going with all this?
We went from:
ChatGPT to
proprietary LLMs to proprietary harnesses to
some proprietary tools now base work on open LLMs to
running a pretty good open model in our own cloud
A decade ago our laptops struggled to compile Sass into CSS.
Today we are running a local S3, multiple dev servers, with hot reload and developing them just fine.
My big prediction, and the future I’m looking forward to is, consumer hardware being able to run a model like Kimi K2.6.
I don’t know when that is coming, but it’s also inevitable.
What do you think? Can we reach the point where we run great, local LLMs on consumer, not fully-specced out hardware?
If you made it this far, thanks for reading. Here’s a glass of water on me 🚰.
I spent 0 tokens to write this to you. Hope it feels refreshing to read something a human would write.
- Akos


