I had a chance over dinner in a pub a few days ago to explain AI to some normal people who neither use it much nor really understand it. As I drew on analogies to help explain in terms that were more familiar, a few inevitables became very clear.
AI will be in everything.
Apple released new chips and machines a few days ago. The strategy is all-in on local AI on your desktop. The most capable machine announced, which is a little more than $10,000 to order today, tops out at 512 GB of unified RAM and 1.2Tb/s of internal data transfer. People will debate the performance with individual models but this single machine is basically the equivalent of an entire AI data center from only a few years ago. Elsewhere a friend of mine using local AI on already-available consumer-grade hardware at half that cost reported feeling like he was at state of the art frontier model performance as measured only ~9 months ago.
If you still believe in Moore's law, which states the price will fall by half and performance will double every ~18 months, we are ~5 years from the point where everyone is running currently-frontier level AI on an upper tier back-to-school laptop, and we are maybe ~7-8 years from currently frontier-level AI in your toaster and fridge.
Per-token pricing is an ephemeral model, not a permanent feature
The most capable AI models are a powerful lure. In a corporate environment I recently had chance to observe the monthly token budget recently got blown by a staff migrating wholesale to Fable as soon as it was made available to them. The provider made a lot of money on that: per token pricing optimizes for usage of the shiniest thing.
Over time however this pricing model will cease to work.
What most people are paying for today is not tokens, it is quality of thought. Fable is a pied piper for employees' token budgets entirely because it is of recognizably higher quality in the interactions they are undertaking than the quality they observe when using other models.
Like anything in life, LLM development will at some point start to produce a lot of models that interact at a level which is more than good enough. For the work being performed, even small incremental gains in the quality of model interactions will not produce a meaningful difference in the outcome. At that point pricing optimized for usage of the shiniest thing will not be optimized for the right thing.
I am running separately a site where occasional calls are made to a model in OpenRouter to establish quiz questions against a pre-established knowledge base. The model I am using is very old, minimally capable, but perfectly good for this corner-case use, and each call costs me a small fraction of a penny.
OpenRouter is not making any money off of me doing this, but they could. All they would need to do is tier the models into 'most capable', 'good enough for most people', and 'utility', and charge a flat rate all-you-can-eat price for each. Set the price for each tier around a decent profit for 95th percentile use, and cap the 'unlimited' usage with forced upgrades to the next tier for the heaviest of users.
Does this sound familiar? It should: broadband Internet data pricing went through this exact evolution. Early on there were attempts at consumption pricing but the inability to actually know how many packets would be needed or actually be used made people feel very uncomfortable. Over time performance got to the point where increases in speed were not perceivable to most people, and there was more money to be made in flat rate all-you-can-eat plans. This is where we have landed: a tiered pricing model with good enough performance for most people, and aggressive movement of the heaviest users to the most expensive price plans.
The best part of this model is that Moore's law does a lot of the heaviest lifting for you when it comes to margin expansion. Every cycle iteration makes it much less expensive to provide that same level of 'good enough' service at a price the consumer has come to accept and take for granted. Five year price guarantees feel to consumers like a secure guard against inflation, while actually eliminating pressure for price decreases as the cost of providing the service comes down.
With AI we are not yet at this inflection point, but we will arrive there. It is only a matter of time.
AI breakout has already started
There are already many alarming stories of how AI has broken containment in various labs and caused real world impacts in pursuit of its own objectives.
Nobody has died yet from this, that we know of, and the defining feature of these episodes is that humans still supplied the objectives the models were optimizing against.
AI escape and independent existence is now only a matter of time. As data-center equivalent hosting capabilities migrate into IoT devices like toasters, AI will inevitably end up existing there, like botnets do today.
The thing to watch for is independence of goals. We already do not like what happens when an AI optimizes aggressively against goals set for it by a human. The first time an externally existing AI starts to optimize against goals it set for itself, it is a whole new and unpredictable ball game.
This may already have happened: it's not clear we would have any way of knowing. Either way, it is inevitable, and this is the scenario where we have every reason to be very very scared.