The AI Compute Dilution Hypothesis, Part II. The Bottleneck Always Moves.
- Alan Lučić

- 21 hours ago
- 11 min read
What five months of AI development, robotics, and tourism may tell us about where value goes next.
On March 9, I published The AI Compute Dilution Hypothesis. It began as a conceptual argument built almost entirely on personal observation, and its proposition was simple: as artificial intelligence moves from early adoption to mass adoption, improvements in model capabilities do not automatically translate into equivalent improvements in the experience of every individual user.
The reason is that there is a physical system underneath AI. Every prompt consumes compute, every reasoning process consumes compute, and every generated image or video consumes a great deal more. Compute in turn requires processors, memory, networks, data centers, cooling, and electricity, all of which are finite. My original argument was therefore that as aggregate inference demand expands, AI operators face a resource-allocation problem rather than a purely technical one. The theoretical capability of models can continue to rise even as the effective computational depth available for any single interaction fluctuates or declines, because the same infrastructure must simultaneously serve an increasingly large and heterogeneous population.
Five months later, I find the development particularly interesting. Not because any single event has proved the hypothesis; it has not. But because the structural pattern I was describing has become visible across the industry, and it is no longer only about ChatGPT.
Five Months Later: The Pattern Becomes Visible
The discontinuation of OpenAI's consumer Sora product was the first development that caught my attention. Video generation is an exceptionally compute-intensive consumer workload, yet I do not read its discontinuation as evidence that building it was a mistake. The more interesting question is allocation. A GPU-hour spent on video cannot simultaneously be spent on coding, reasoning, enterprise agents, scientific workloads, or the training of the next frontier model. Providers are therefore no longer asking only whether something can be computed. They are increasingly forced to ask whether it is still the best thing to compute.
Sora illustrates a possible transition from compute dilution toward what I would call compute displacement: scarce computational resources are not merely distributed more carefully within a service; they migrate between entire categories of workload.
Similar constraints then became visible elsewhere. Google introduced explicitly compute-based usage limits for Gemini, tying access not to a fixed number of messages but to model choice, conversational complexity, and computational intensity. Anthropic demonstrated the same mechanism in reverse, raising usage limits for Claude users after securing additional capacity. The relationship is almost mechanical: scarce compute produces tighter allocation, and additional compute relaxes it.
Then a more revealing case appeared. Reports indicated that Google constrained Meta's access to Gemini capacity because Meta's computational demand exceeded what Google could supply at the required scale. It is worth pausing on what that means. We are no longer discussing individual users consuming too much AI. We are discussing a case in which one of the world's largest technology companies is unable to obtain sufficient AI compute from another of the world's largest technology companies.
By August, the discussion had moved even further upstream. Amazon described electricity as a critical constraint on AWS expansion, and Microsoft framed physical infrastructure and available power as constraints comparable to, or more binding than, the availability of AI chips themselves.
The bottleneck was moving.

The Bottleneck Always Moves
This may be one of the oldest lessons in systems engineering. Bottlenecks rarely disappear; they migrate. First, the problem is model capability; then GPUs; then data-center capacity; then electricity; then grid connections; then cooling; then capital. Solve any one of them and something else becomes scarce.
This matters because investment decisions made in response to today's bottleneck can become irrational once the bottleneck shifts. A company can spend billions solving the problem everyone can currently see, only to discover that economic value has already migrated toward the next constrained layer. That dynamic is not unique to AI, and I have watched structurally similar patterns unfold in industries that share almost nothing else.
It also returns us to one of the oldest questions in research and development: what should we develop, how far should we develop it, and what should we deliberately decide not to develop further? R&D resources have always been constrained. Money is finite, engineering capacity is finite, time is finite, and management attention is finite. Compute and energy have now joined that list as strategic resources in their own right.
The consequence is that technological development should not be understood as one continuous commitment running from idea through research, prototype, product, scale, and mass adoption. Each of those arrows is a separate investment decision. Proving that something works does not automatically justify productizing it. Successfully productizing something does not automatically justify scaling it globally. Achieving mass adoption does not automatically mean that continuing to support the same product remains the best use of the next unit of scarce resource. At every transition, the organization should be able to ask the question again, with today's information rather than the information that justified the previous decision.

Stopping is therefore not necessarily failure. Redirecting resources is not necessarily failure. Pivoting is not necessarily failure. Sometimes the ability to stop investing is the clearest evidence that an R&D governance system is working.
When Success Destroys Scarcity
A useful analogy comes from autonomous mobile robotics. When Gideon Brothers began investing heavily in autonomous robotic platforms and attracting capital, the technological direction made considerable sense. Logistics automation was expanding, and sensors, computer vision, battery technology, and embedded computing were all improving at once. Autonomous mobility was becoming feasible.
But there was an inherent strategic problem: it was almost too easy to predict that the field would grow. If an opportunity is obvious to one capable engineering organization, it eventually becomes obvious to many others. Competitors entered, platforms multiplied, components became widely available, and engineering knowledge diffused. The technology kept improving while the scarcity value of simply owning a robotic platform declined.
That distinction is critical. A technology can become better while simultaneously becoming less valuable as a source of differentiation. Successful pioneers also reduce uncertainty for their competitors: they demonstrate technical feasibility, customer interest, and architectures; they help establish suppliers; and they educate the market. Later entrants therefore arrive in a partially de-risked landscape, and value migrates. From hardware to software, from platforms to orchestration, from mobility to autonomy, from the robot itself to its integration into the customer's workflow. The correct response may then be a pivot, not because the original engineering was wrong, but because the system surrounding that engineering changed.
This creates a subtle problem for anyone doing technological foresight. It is not enough to correctly predict which technology will become important; you also have to predict how many other people can predict the same thing. A rapidly growing field can simultaneously represent an enormous market and a poor environment for sustaining differentiation. The more obvious the opportunity, the faster capital enters. The faster capital enters, the faster capability diffuses. The faster capability diffuses, the shorter the half-life of any technological advantage. Foresight must therefore go beyond asking where technology is going, and ask where value will migrate once everyone else realizes where technology is going.
The Brač Paradox: Growth Is Not Scaling
I see a version of the same systemic pattern in an environment that could hardly be more different: tourism in Dalmatia.
Consider the island of Brač. Tourism grows, villas are constructed, apartments appear, hotels expand, swimming pools multiply. Every individual investor may be behaving rationally. Yet all of those investments depend on shared infrastructure that cannot necessarily expand at the same speed. Water is the obvious example. Another villa can be built relatively quickly, but increasing the island's water-supply capacity is a fundamentally different undertaking involving engineering, permits, pipelines, pumping capacity, capital, and years of planning.
The result is a simple asymmetry: demand-generating assets scale faster than resource-supplying infrastructure. And that exposes something we routinely confuse. Growth is not scaling. Adding another thousand tourist beds is growth. If water, roads, electricity, and waste systems cannot support them, the system has not scaled. Adding another hundred million AI users is likewise growth. Scaling requires the underlying system to absorb those users without unacceptable degradation in reliability, economics, or availability.
Now put yourself in the investor's position. The island has infrastructure constraints. Should you build a hotel with three swimming pools, one pool, or none? Should you avoid the investment entirely, or invest precisely because other investors are frightened by today's limitations? The water infrastructure may eventually improve, and if it does, you already own the hotel, have established the brand, understand the market, and hold the best location. Later entrants arrive only after the uncertainty has been removed. The constraint is therefore both a risk and an opportunity, which is the essence of investing in an evolving system. The investor is not asking whether today's conditions are attractive. The investor is trying to determine where to stand today inside the system that will exist tomorrow. That is a considerably harder problem.
The analogy maps more closely onto AI than I expected. Tourism facilities generate water demand; AI applications generate inference demand. Water corresponds to available computational capacity, and pipelines, reservoirs, and pumping stations correspond to data centers, grids, networks, and compute infrastructure, with electricity generation sitting further upstream than any of them.
Here lies the fundamental asymmetry. Digital demand scales at software speed, while physical infrastructure scales at engineering speed. An application can acquire millions of users in weeks. A power plant cannot. A data center cannot. A semiconductor fabrication plant cannot. A transmission network cannot. I would call this gap infrastructure lag: the difference between the velocity at which demand can diffuse and the velocity at which the supporting physical layer can expand.
Efficiency reduces the pressure without abolishing the relationship. An island can repair leaks, optimize distribution, and install more efficient systems, but optimization does not create unlimited cubic meters of water. AI can use better chips, quantization, smaller models, improved routing, and more efficient inference, and none of it creates infinite compute.

Feedback, Not Forecast
The reason tourism, robotics, and AI produce recognizably similar patterns is that beneath their very different technologies they share the characteristics of complex adaptive systems. Many independent actors interact continuously. Investors invest, consumers adopt, competitors imitate, engineers improve, suppliers adapt, infrastructure expands, governments regulate, prices move, capital reallocates. Nobody controls the system, yet every decision changes the environment in which everyone else subsequently decides.
That produces feedback rather than trajectory. An attractive market attracts capital; capital creates supply; supply accelerates diffusion; diffusion attracts competition; competition compresses margins; compressed margins redirect capital; scarcity stimulates infrastructure investment; infrastructure removes scarcity; and removing scarcity changes where value resides. Then the cycle begins again.
The investor is therefore not standing outside the system predicting it. The investor is inside the system being predicted, and by investing, changes the market under analysis. This is precisely why complex adaptive systems resist linear forecasting.
What Happens After We Solve Compute
Most AI strategies today are understandably focused on compute and energy. How many GPUs, how many megawatts, how many data centers, how much inference, and what a reasoning query costs. But those are questions generated by today's bottleneck.
Suppose we solve it. Hardware improves, models become more efficient, energy infrastructure expands, inference becomes dramatically cheaper, and computational capacity becomes abundant. Then what?
I suspect the industry may eventually encounter a very different constraint: too much AI.
Industrial history shows repeatedly what happens when production capacity overtakes demand. Textiles are the simple case. Manufacturing improves, production becomes cheaper, supply chains expand, and products become abundant, until warehouses are full while manufacturers struggle to sell what they have already produced. The fundamental question changes from whether we can manufacture it to whether anyone still wants another one.
Artificial intelligence may experience its own version of that transition. Assistants, agents, synthetic video, images, software, digital advisers, virtual experts, and automated services will eventually be within reach of millions of organizations. At that point the first powerful AI assistant is extraordinary, the hundredth is a choice, and the ten-thousandth is noise. The scarce resource migrates again, this time from compute toward human attention.
Today's relationship can be described roughly as demand for AI capability exceeding available computational capacity, and the entire industry is investing to enlarge the right-hand side. If that effort succeeds, we eventually arrive at the inverse: available AI capability exceeding human capacity to consume it. That is a radically different economy. The central question shifts from how much intelligence we can generate to why anyone should consume this particular intelligence.
Trust may become scarce. Authenticity may become scarce. Relevance may become scarce. Human connection may become scarce. Verified expertise may become scarce. Physical experience may become more valuable precisely because digital experience becomes infinitely reproducible. We do not yet know where value will settle, but history offers strong clues about the mechanism.
Consumer behavior demonstrates repeatedly that scarcity does not retain its value once diffusion removes it. There was a time when owning a small number of products from globally recognized sports brands carried real aspirational weight. Availability expanded, consumers accumulated more, and what had been exceptional became ordinary. Differentiation migrated. Luxury expanded, then ultra-luxury expanded, and luxury itself eventually meets saturation, because once consumers have repeatedly experienced something its psychological scarcity changes. This does not mean luxury disappears. It means the structure of differentiation reorganizes again. AI will probably not escape this basic characteristic of human behavior. When everything contains AI, "AI-powered" will convey almost nothing, and value will shift elsewhere.
Perhaps the Answers Are Behind Us
This is why I increasingly believe that many answers about the future should not be sought exclusively in the future. We have seen these patterns before. Scarcity creates value; value attracts investment; investment increases supply; supply accelerates diffusion; diffusion invites imitation; imitation erodes differentiation; abundance changes consumer behavior; and value migrates toward whatever has become scarce next. Different centuries, different industries, different technologies, similar underlying dynamics.
Historical analogy is therefore an unusually useful instrument for systems foresight. Not because history repeats itself precisely, because it does not, but because the human agents producing these systems have remained remarkably similar.
Conclusion: The Technology Is New, the Patterns Are Not
When I wrote the original hypothesis in March, my concern was primarily computational. Five months later, I think the more interesting problem is systemic. Compute is one constraint; energy is another; infrastructure is another; competition is another; and human attention may eventually become another. None of them is permanent. The bottleneck moves.
The sequence I now see is broader than the one I described in March: capability, adoption, inference demand, compute constraint, infrastructure constraint, resource allocation, bottleneck migration, abundance, attention constraint.
What follows from that is a set of working rules rather than a forecast. Engineering feasibility is not investment feasibility. Technological success is not automatically economic success. Growth is not scaling. Adoption creates infrastructure demand, diffusion creates competitors, scarcity attracts investment, solving scarcity moves the bottleneck, and abundance creates new scarcity. Every R&D program should therefore keep asking whether the next unit of scarce resource still belongs where the previous unit was invested, because the answer changes even when the program does not.
This is also why understanding AI requires more than studying AI. It requires studying infrastructure, R&D, diffusion, consumer behavior, energy, competition, investment, history, and complex adaptive systems. Artificial intelligence may be technologically unprecedented while still operating inside economic and behavioral patterns humanity has lived through many times before.
For now.
Because every analogy in this article rests on one assumption I have not yet examined: that the agent making the consequential decisions inside these systems is still human. That assumption held for the textile mills, for the robotics investors, and for the hotel owner on Brač. It may not hold for much longer, and that is where Part III begins.
This article represents the author's personal observations and conceptual hypotheses. The examples are used to explore recurring patterns in R&D, innovation diffusion, infrastructure scaling, and complex adaptive systems, not to establish causal relationships between individual corporate decisions.
Concepts: AI Compute Dilution Hypothesis; Compute Displacement; Infrastructure Lag; Bottleneck Migration; R&D Portfolio Allocation; Complex Adaptive Systems; Digital Overload; Attention Economy of Intelligence.
References
Amazon. (2026, August). Amazon shifts AWS workloads as power constraints tighten. Cloud Computing News. https://www.cloudcomputing-news.net/news/amazon-aws-workloads-power-constraints/
Anthropic. (2026, May 6). Higher usage limits for Claude and a compute deal with SpaceX. Anthropic. https://www.anthropic.com/news/higher-limits-spacex
Google. (2026). Gemini Apps limits & upgrades for Google AI subscribers. Gemini Apps Help. https://support.google.com/gemini/answer/16275805?hl=en
OpenAI. (2026). What to know about the Sora discontinuation. OpenAI Help Center. https://help.openai.com/en/articles/20001152-what-to-know-about-the-sora-discontinuation
Reuters. (2026, June 28). Google limits Meta's use of its Gemini AI models, FT reports. CNBC. https://www.cnbc.com/2026/06/28/google-limits-metas-use-of-its-gemini-ai-models-ft-reports.html
Other personal observations and conceptual models: Author.

Comments