AI inference is turning into a commodity. Rather than letting big tech companies set model prices, we believe the market should decide what AI is worth. DEX.DO is the platform built so that can happen.
We’ve talked about DEX.DO before, describing it as a dark, on-chain order book built on Acki Nacki, and showing how it could be used for prediction markets. But is focusing on YES/NO wagers on real-world events really the best way for DEX.DO to make a difference in that same all too real world? We suspect DEX.DO was built for something much bigger than that.
Buying and selling AI inference works just like buying and selling any other commodity. Because that’s precisely what AI inference has become. Every time a model answers a prompt, it uses ‘tokens’—the small units models run on. Companies buy these tokens in bulk, paying millions at a time. OpenAI, Anthropic, Google, and others all charge by the million tokens, both for input and output. Spending on this is already massive and growing fast, with prices almost always set by the providers.
DEX.DO changes this system. It offers a live order book where you can buy and sell inference like any other traded good. You place an order, and it matches you with the seller offering the lowest price at that moment. The price comes from what buyers and sellers agree on, not from a provider’s rate card. If the lowest-priced seller runs out, your order moves to the next best price, so you’re never stuck with just one supplier.
To see why inference should be traded on a market, look at its costs. Training a large model is a big, one-time expense—like the millions spent on GPT-3 or the hundreds of millions for GPT-4—but running the model costs money every time it’s used, and that cost never stops. When models handle billions of prompts daily, running costs quickly outpace building costs, and these are measured in tokens. Tokens are sold per unit at a set price, with a floor based on production costs and a ceiling based on what buyers are willing to pay. Providers often sell below cost to win market share, but that’s just a subsidy, not a real price. This setup is more like trading oil or electricity than selling an app. Trading inference is also more straightforward than most crypto tokens, which lack this kind of real-world anchor.
Finance is already responding. Major exchanges are setting up markets for these assets, proving that inference is tradable. On October 5, CME Group and Silicon Data launched two compute futures contracts, pending regulatory review: one for hourly rental of the Nvidia H100, the other for the Blackwell B200. Both cover a month’s GPU rental and are listed on NYMEX. The Intercontinental Exchange is working on something similar. The CFTC has started a public consultation on compute derivatives, which shows regulators now recognize this category. CME isn’t even the first—Architect Financial Technologies has offered perpetual futures on GPU and DRAM rental since January, from Bermuda. Meanwhile, the Shanghai Futures Exchange is researching contracts based on tokens, focusing on consumption rather than production. This is a key difference. Two of the world’s biggest exchanges are developing the same asset class but using different units. So, it’s clear inference can be traded. The open questions are what unit to use and where the spot price will be set.
Two facts make this urgent. First, prices can spike: SemiAnalysis’s one-year H100 rental index jumped almost 40% between October 2025 and March 2026, from $1.70 to $2.35 an hour, and it’s still rising—now above $2 and heading for $3. On-demand capacity is sold out, and renters are even subletting their clusters like apartments during big events. Second, prices can collapse: Epoch AI reports that the price of a fixed capability can drop anywhere from 9x to 900x a year, with a median near 50x, and about 200x a year for models released since January 2024. GPT-4-class output cost about $30 per million tokens in 2023, but now it’s under $0.50. Go back further—GPT-3-class output was $60 per million in 2021 and is now just six cents. Buyers want to benefit from falling prices but also protect themselves from spikes. These are classic trading problems with a familiar solution.
This demand is real, not theoretical. OpenRouter handles over 400 models and about 25 trillion tokens each week, which adds up to nearly one and a half quadrillion tokens a year from eight million developers. In May, it raised about $1.3 billion, led by Alphabet’s growth fund, aiming to become the Stripe of AI. Stripe has now agreed to buy it for more than $7 billion. That’s important, because it shows the market is paying a huge premium for the very layer we believe is flawed. Stripe bought distribution—one interface, every model, and millions of developers already connected—which is valuable. But it didn’t buy a price mechanism, since OpenRouter has never had one. It simply passes each provider’s listed price to buyers and takes a cut, about 5% on credit purchases and on bring-your-own-key spending above a certain amount. Buyers can’t bid, and sellers can’t compete by offering lower prices, because there’s no way to do that. Seven billion dollars only makes sense if you think the rate-card system will last another decade. We don’t think it will.
The data shows growing pressure for change. In just twelve months, Chinese models went from less than 2% to over 60% of routed tokens, taking the top five spots by July. Meanwhile, the American big three dropped from about 70% of the volume to around 30%. Quality and usage no longer match up—Claude Opus 4.8 leads the Artificial Analysis index and handles about 200 billion tokens a day, while DeepSeek V4 Flash moves 619 billion. This is a market trying to form, even though the current system holds it back, and it’s finding ways to break through.
Letting providers set the price isn’t just a complaint—it’s a structural issue. Both Brookings and RAND have pointed out that frontier models tend to become monopolies because they’re expensive to build but cheap to copy. If only a few companies control the models, endpoints, and rate cards, they set AI costs for everyone else, and the gap between their price and the real running cost is paid by every business that relies on them. The usual fix is utility regulation, like we do for water or electricity. But there’s a simpler way: you don’t need to regulate a price if no single company controls it. DEX.DO doesn’t have a provider list or a middleman. It settles trades in under a second, all enforced by code. Here, decentralization isn’t just a buzzword—it’s what actually takes pricing power away from providers.
This isn’t just a prediction—it’s already happening in China. Over the past two years, China has turned token production into an industry, building ‘token factories’ that only produce model tokens. State telecoms now sell token bundles to regular consumers for about $1.50 a month. By this spring, China was handling 140 trillion token calls a day, up a thousand times in two years. China treats tokens as a bulk commodity and is building exchanges to trade them. That’s the scale the rest of the world faces, and it can’t be matched by just three American companies’ price lists. It takes an open, neutral, global order book where anyone—from a solo developer with a fine-tuned model to a data center with extra capacity—can sell on equal terms. Anyone can sell inference, resell unused capacity, or add their own expertise and sell an agent bundled with the tokens it needs. The market rewards useful supply and real expertise, not just size or the ability to publish a price list.
The next wave of demand won’t even come from people. Gartner says an agent working on a task uses five to thirty times more tokens than a single chat message, and agents are now the fastest-growing use of AI. Two-thirds of large companies already use over a billion tokens each month. Uber’s CTO said the company spent its entire 2026 AI budget just a few months into the year. OpenRouter’s founder calls uncontrolled usage an “infinite cost center,” and the International Energy Agency reports that AI data center power use grew by nearly 50% in a year. Every serious company will need what commodity traders have always relied on: the best price, a way to lock it in before a spike, and a way to hedge the rest. DEX.DO offers all three: best-price matching, forward contracts, and prediction markets that finally serve a real purpose—letting you bet against and offset price spikes. More and more, it won’t be people placing orders, but agents. These agents will pick the best model for each task, buy tokens at the lowest price, spread work across the order book, and hedge their own costs, all automatically. An agent can’t shop around rate cards, but it can trade on an order book. DEX.DO was built for these agents first, accessible from the command line for developers and their agents, and from a web app for everyone else.
DEX.DO uses a dark book, so your trades stay private where it matters. Orders aren’t broadcast, trade sizes don’t leak, and prices only appear after a trade, if at all. No one can see your inference purchases and jump ahead, and no one can track your activity to trade against you. You pay in small increments as tokens arrive, and sellers can’t get more than two increments ahead of what they’ve delivered. This means no middleman ever holds your money, and neither side is at risk if the other disappears. Both sides remain anonymous. This is the same way the New York Stock Exchange keeps large orders from moving prices against buyers; DEX.DO brings this approach on-chain for the first time, targeting what could become one of the most traded assets of the next decade. Sellers compete openly, and every order fills at the best price, which can be up to 90% below list on a competitive book. Early traders and liquidity providers also earn a share of tokens, based on how active they are, so the more you help build the market now, the bigger your stake will be later.
When we say a market should set the price for inference, we don’t just mean costs should be visible. We mean that no one—including the labs—really knows what inference is worth. They know what they charge, but that’s just a decision made by a few people, based on their own costs, goals, and what competitors charge. The true value to the millions who might use it isn’t written down anywhere. It’s scattered in pieces among all those users, and most of it is never even put into words. Right now, there’s no way in the industry to bring those pieces together.
In 1968, Friedrich Hayek gave a lecture called “Competition as a Discovery Procedure.” His argument was more focused and unusual than how it’s often used today. At the time, economists assumed that everyone already knew the important facts—costs, preferences, and available methods. Hayek disagreed, saying that if you assume that, you miss the whole point of competition. Competition isn’t just a way to reach a result you could have calculated ahead of time. It’s a process for discovering things no one knew, because the knowledge is scattered in small pieces among many people and can’t just be collected by a committee.
This idea has an important consequence that many of Hayek’s supporters overlook. If competition matters because its results can’t be predicted, then you can’t defend it by pointing to those results. You can only defend the process itself. Hayek made this clear, but most people who quote him don’t mention it.
We’re drawn to this idea because it matches our own focus on discovery, not disruption, but with stronger reasoning. Disruption claims to know what should change and what should stay, and acts quickly based on that belief. Discovery, on the other hand, admits it doesn’t know. Our argument about inference pricing is, at its core, Hayekian: a few companies set the price of a major new input for the world economy by internal calculation, post it on a card, and change it whenever they want. If you remove the branding, that’s just central planning. The real criticism of the leading labs isn’t that they’re too capitalist—it’s that they’ve recreated central planning, just with a better developer portal.
Hayek’s idea of competition is about discovery. It brings hidden and scattered facts to the surface, and the price is the result. DEX.DO is the vehicle for this journey. Let’s discover what AI is truly worth.
DEX.DO Season 2 is now live on Acki Nacki Mainnet. Users can earn points by trading AI model access, consuming purchased inference, and inviting new participants. Points accumulated during the incentive program will convert into a NACKL allocation once all seasons are completed. A live leaderboard is available at app.dex.do, where participants can track their rank and progress throughout the season.
For more on DEX.DO in the future, consider subscribing:


