World of Cheap and Productive Overnight Agents

One person with a chatbot burns tokens only while typing. One person with 3–20 background bots burns tokens all night. This month sees the start of grokbot boom that adds hours and ultra low cost models.

Cheap models make it affordable. GLM-5.3-Flash at $0.15/$0.50 per million is ~10x cheaper than a frontier model. $20 per month buys 10x the work. Cheap tokens invite waste. Bots re-reading the same repo, re-planning solved problems, running when nothing changed. Waste doesn’t just cost money — it eats the gigawatts. Caching, sleep-when-idle, and routing the hard 5% to a frontier model are what keep value per token from collapsing.

In July, the top 1% of businesses spent a median $7,400 per employee on AI. The top 10% spent $650. The median firm spent $11.95 per employee.

Screenshot

Screenshot

IF Most People Can Become PowerUsers Then Demand Can 200X-400X by 2030

Demand can 200–400X, not 24X by 2030.

Goldman’s forecast is 24X by 2030. That assumes agents stay a feature. If they become the default way work gets done, you get 200–400X instead — which is just 35–40% per quarter, sustained for 18 quarters.

Screenshot
Screenshot
Screenshot

Supply Hinges on More Tokens per Gigawatt or Improved Tokens Per Tasks

Today the industry runs ~25 GW and gets about 17 quadrillion tokens per GW-year. To serve 300X you need 108,000Q/year. Whether that’s possible depends entirely on how fast efficiency improves.

Screenshot

Earth based power additions have limits and ramping delays. 2.7–3x per year in tokens per gigawatt.

Two levers get you there, and you need both. ..

Nextbigfuture substack has the rest of the article.

6 thoughts on “World of Cheap and Productive Overnight Agents”

  1. I use AI professionally and personally more than most people. It’s built into a security product that my company is shipping. All of our clients refused to allow us to use frontier models to process their data. Moreover, it would’ve been cost prohibitive.

    We switched to developing a fine-tuned small language model that we can run on our own (cloud) hardware. This approach eliminates the privacy concerns, and reduce cost by about a factor of 1000.

    I think a good portion of the AI compute will migrate away from frontier models, and to this type of local processing. The smaller models are getting very close to being good enough for most uses out of the box, and the combination of extra cost and loss of privacy associated with public frontier models makes them inappropriate for many applications, such as healthcare, finance, and security.

  2. 1. Bots being trained to CREATE (and then manage) novel new types of bots is already in development. This alone drives usage up by 100X.

    2. Financial trading, personalized medicine, business administration, personalized content creation, personalized AR/VR computer gaming, physical AI (robots) and science research, national defense/weaponry are in the very earliest of stages. At least 100X growth in those areas.

    3. I think personal AI agents/bots will become as ubiquitous as owning a car, with people willing to pay similar amounts of their monthly incomes. Offloading the drudgeries and risks of life will cause explosive growth.

    4. For better or worse, the great acceleration has begun…

  3. Demand is finite. Brian was forecasting huge numbers about the number of future Gigafactures and Teslas overall, how growth wont stop etc. Those predictions proved false. Numbers plateaued.

    AI has lots of potential for growth, but not so much as it is artificially hyped.

  4. By my calculations you need 1 billion people paying $200/month by 2030 or 100 million paying $2k/month to get in that $2t+ AI industry run rate.

    To get to that and a 300x token demand from today you need 8 hour long AI knowledge work sessions that require no human intervention. The AI has to work in a trustworthy manner that whole time with high intelligence. At that point one worker can spend 30 minutes designing an 8 hour session and get 10 of those going and check them the next day. That’s where you get $200/month or $2k/month spends at scale and 300x+ token demand.

    According to METR we can get there by 2030.

    The other way you get that is on the medical front but I don’t see that scaling as fast.

  5. This is excellent. I’ve been working on these leading indicates of supply/demand problems myself. Very helpful. My SpaceX stock is locked til next year so do I hedge or ride it?

    in that final chart the caption hide the third bar – I assume that bar just goes to the bottom of the caption.

    As I’m sure you know Epoch has efficient at 3x per year, but it seems way higher than that given how performant these small models are getting.

    Personally, I consider myself an early adopter of AI, and yet, with the list of things I have to implement for a Grok Bot capability, It’s easy to see my token usage going up 100x or 1000x and eventually 10,000x, and even above that with personalized medicine – easy to see 100,000x or 1,000,000x. Currently, I’m in the top 10% from your chart above of late in terms of spend. The longer AI can work autonomously, without constant oversight, and with high intelligence, with trust, and cheaply, there really is an explosion of demand that happens.

Comments are closed.