Conversations-32: Owning your stack
The ChatGPT moment in 2023 was the first technological shift I experienced as a genuine paradigm change, the kind that is a little scary. Aware it would reshape how we work, I gave myself thirty minutes a day to experiment with it.
Three years and many “aha” moments later (Claude Code, Cowork, Browser, OpenClaw, omp.sh…), I feel I have earned an informed enough take to share it.
Despite the clamor, LLMs will not take jobs at scale. Someone has to operate the agents, judge their outputs, teach the machine what is right and what is wrong. Work as a whole is shifting: administrative tasks fade, and the more I integrate AI into my workflow, the more I can focus on high-value work like negotiating key BD deals, designing products.
That shift is exactly why the rest of this matters. If the human’s role becomes operating and directing a fleet of agents, then those agents move to the center of how we work and think, and the question of who controls them, and who sees what flows through them, stops being academic.
Integrate AI deeply enough, professionally and personally, and you eventually realize you share more information with the model than with any other person or piece of technology. That data is, in theory, accessible to and exploitable by the model providers. Three consequences follow, in rough order of how confident I am about each.
First, I would bet an AI lab gets breached at some point, and when it does it could be one of the largest personal-data leaks we have seen, simply because of how intimate and concentrated that data is.
Second, these labs are burning more cash than almost any company in history. That pressure pushes them toward adjacent business models (robotics, pharma, crypto…), and it is plausible they will use what they learn from user prompts as a lever to bootstrap those new activities.
Third, states will eventually (if they have not already) require identification to use these models and seek some form of surveillance access (in Europe, a modest extension of Chat Control would suffice). It will start with universally condemned behavior (terrorism, child abuse…) and can drift toward the more authoritarian (tax data, say). How far the line moves will depend on the democratic health of each country, and there are real counter-forces too: GDPR, end-to-end encryption, competition between labs.
The sole possibility of any of these scenarios happening is likely enough to plan around. So, in an age where trust in institutions is at a low, and like any self-respecting tech and crypto person, I decided I wanted to own my AI stack.
This is where it gets complicated.
1. First, the good news, the models finally exist locally. Frontier models have historically been far ahead of everything else, and have only really worked at scale since Opus 4.5 (December 2025). The consensus was that open models would trail by one to two years. That gap collapsed faster than anyone expected: GLM 5.2 shipped in June 2026 at roughly Opus 4.5 level, and weeks later Kimi K3 landed above Opus 4.8. High-quality local models already exist, this is no longer the blocker.
2. But these models are enormous. These models run between 100B and 1T parameters, which means 500–2000 GB of fast memory (GPU VRAM, or unified memory) to run them well. A rig that does this smoothly costs on the order of $250,000. This is an amount far beyond what any individual will spend. You can go cheaper (a 512 GB unified-memory machine runs large MoE models, slowly), but “slow and roughly equivalent” is not the same product as a snappy subscription.
3. Hardware is scarce and expensive. It turns out every tech company, and soon every legacy company, wants to run AI. Suppliers cannot produce a fraction of the demand which leads to prices exploding.
4. And the grid is not ready. Energy infrastructure across much of the West has grown little in two decades while demand is about to surge; it needs serious investment to absorb the coming decade. France has an edge with its nuclear fleet, but complacency, plus obligations to export power to neighbors, risk landing us where the US already is.
5. And even then, the economics do not work yet. Labs have subsidized their models so heavily to drive adoption that switching to ownership is not marginally more expensive; it is dramatically so. The honest way to size the gap: Claude Max 20x costs $2,400 a year. A comparable local setup runs about $250,000 in hardware. Even amortized over the five years such a rig might last (~$50,000 a year), you are still paying roughly 20x more, before electricity, maintenance, and the model quality you give up. This one-plus order of magnitude, compared to current subscription offers.
Businesses will likely move first. As individuals, most of us cannot go fully local for the foreseeable future. That leaves four options:
Keep subscribing: cheap, but you are exposed on privacy and hostage to whatever the labs decide caps and prices should be.
Go fully local: sovereign, but $250k and years away for most people.
Use a private aggregator:
A private aggregator pools prompts across many users, runs as many open models locally as it can, and encrypts the data in transit and at rest. Here is what it does and does not protect: if you insist on a frontier (SoTA) model, your prompt still has to reach whoever hosts that model, so aggregation cannot fully eliminate that exposure. But for the growing set of tasks that open models now handle, an aggregator means no single provider builds a profile of you, your prompts are not training fodder, and the attack surface shrinks dramatically. It will not beat a $250k sovereign rig on privacy, but it gets most of the way there for a tiny fraction of the cost.
4. The path we built
The catch is that you now pay per token instead of a flat subscription. For a heavy subscription user, metered usage of these private models can run 10–20x what your flat plan effectively costs you per token today. For most people, that is a non-starter. Which is why we built cheaptokens.ai: access to these private models at 40–90% off standard metered token prices, closing most of that gap.
Full disclosure: I sell this. So weigh the diagnosis on its own merits; if you are a light user, honestly, a subscription may be all you need. This piece is for power users.
Once I got hooked to AI assisted work, dependence on token consumption became the glaring trap. As the labs mature and approach their IPOs, they will have to rationalize costs. That is when subscription caps get cut hard, and you are left choosing between lower-quality models, paying multiples more, or simply doing less. Privacy and cost turn out to be the same problem wearing two masks. Both issues come down to not owning the thing you now depend on.
cheaptokens.ai is the fourth and most comfortable path for my situation. It is the one we chose, and that we are now opening to everyone.
Happy vibe coding, folks. See you next time.
PS: one side effect of the AI hype I am grateful for: everyone’s attention moved to it and forgot about crypto. The promise of algorithmic coordination is still very much alive, and there is so much left to build. And yes, crypto is still my main thing.




