I’ve spent much of the last few years on the other side of the OpenRouter API.
At Together AI and FriendliAI, part of our business was doing the decidedly unabstracted work behind AI inference: taking expensive GPUs, increasingly sophisticated serving software, and open-weight models and turning them into tokens developers would actually pay for.
My job was to help answer an equally difficult question:
Why should they buy those tokens from us?
That’s why Stripe’s reported acquisition of OpenRouter is so interesting to me.
I don’t primarily see a payments company buying a model marketplace. I see something potentially much more consequential for the inference companies underneath it.
The winner in inference may not be the company that serves the most tokens. It may be the company that decides where the tokens are served.
The strange GTM problem of selling the same model
Inference is one of the more unusual markets I’ve worked in.
At Google Cloud, DigitalOcean and Vultr, there were plenty of infrastructure competitors. But the products themselves were ours.
Open-weight AI changes that.
If a developer wants an open model— DeepSeek, Qwen, Kimi, GLM, Nemotron, or whatever becomes popular next—there may be many providers willing to serve exactly the same weights.
That creates a fascinating product marketing problem.
I’ve spent a lot of time trying to solve it.
Maybe our inference engine is faster. Maybe we get more useful work out of the same GPUs. Maybe we have better batching, scheduling, caching, quantization or speculative decoding. Maybe we can deliver better throughput or lower latency. Maybe our economics allow us to offer a better price.
These aren't imaginary differences. They can represent extraordinary engineering.
But then comes the hard part:
Getting a customer to care.
From the developer's perspective, the output can still be tokens from the same model.
That gap—between enormous technical differentiation underneath the API and apparent similarity above it—is one of the defining GTM challenges of open-model inference.
OpenRouter turns that problem on its head.
OpenRouter makes the provider a variable
A developer using OpenRouter doesn't necessarily need to decide where inference happens.
They can decide what they want.
A particular model. A certain price. Sufficient performance. Maybe particular privacy or availability requirements.
The infrastructure underneath can increasingly be someone else's problem.
That sounds like a convenience feature.
I think it's much more important than that.
Because once the intermediary can choose among providers, the provider itself becomes a variable.
Who is cheapest?
Who is fastest?
Who has capacity?
Who is reliable?
Who can serve this model in the geography I need?
Who should get this particular request, right now?
I've worked for companies trying to win those requests.
OpenRouter gets to decide who wins them.
That is a very different position in the value chain.
Owning GPUs and selling inference are not the same business
I've become convinced of something else after spending the last several years around AI infrastructure:
Owning GPUs does not mean you have an inference business.
The AI industry has understandably spent enormous energy talking about supply.
GPUs. Clusters. Megawatts. Data centers. Blackwell. Who has capacity and who doesn't.
But a GPU is a long way from an inference product.
Someone has to operate the serving stack. Optimize the hardware. Support the models. Manage scheduling and batching. Hit the latency and throughput targets. Build the APIs. Price the service.
And then someone has to create demand for it.
That last part matters.
There are a lot of companies in the world that can own GPUs.
There are fewer that can turn those GPUs into a great inference service.
And there are fewer still that can convince a large number of developers to buy inference from them.
I have spent years thinking about this problem from the supply side:
How do you turn compute into inference, and inference into a product?
OpenRouter starts from the opposite direction.
It aggregates the demand first.
Then it can decide where the supply comes from.
The more I think about that distinction, the more important it seems.
Inference is becoming a market
This is the bigger change I think we're watching.
Historically, cloud infrastructure has largely been destination-based.
You chose AWS.
You chose Google Cloud.
You chose Azure.
The first generation of GPU clouds largely worked the same way. Pick a provider, provision infrastructure, deploy the workload.
Open-model inference has the potential to behave very differently.
The same model can be offered by many suppliers.
Prices are measurable.
Performance is measurable.
Capacity changes.
New providers can enter.
Traffic can move.
Customers don't necessarily care where the underlying compute lives.
Those aren't just characteristics of a cloud service.
They're characteristics of a market.
And markets create value for companies that organize them.
Discovery.
Comparison.
Measurement.
Routing.
Billing.
Settlement.
The machinery connecting supply and demand.
Which is why Stripe acquiring OpenRouter starts to make a lot more sense.
Stripe isn't buying GPUs. It's buying the market layer.
Stripe became one of the world's most important infrastructure companies without manufacturing the thing being sold.
It built infrastructure between buyers and sellers.
A developer doesn't want to think about acquiring banks, card networks, currencies, fraud systems and settlement every time someone clicks Buy.
Stripe absorbs that complexity and exposes an API.
OpenRouter is beginning to do something analogous for intelligence.
A developer doesn't necessarily want to track the constantly changing matrix of models, providers, prices, latency, throughput and availability every time an application needs an answer.
OpenRouter can absorb that complexity.
And there's another connection that I think makes the combination particularly interesting:
Every token is also a transaction.
Someone consumes compute.
Someone supplies it.
Someone meters it.
Someone pays for it.
And increasingly, the buyer may not even be human.
An agent might decide that one step requires the best reasoning model available, while the next ten can be handled by something dramatically cheaper. It might make that decision thousands of times a day.
At that point, routing intelligence and economic infrastructure begin to converge.
What intelligence should I buy?
From whom?
At what price?
For this request?
Right now?
Stripe has spent its life building infrastructure for transactions.
OpenRouter is building infrastructure for transactions in intelligence.
Put that way, the acquisition doesn't seem strange at all.
There's an uncomfortable GTM lesson for inference providers
For years, I've thought about the central GTM question for an inference company as:
Why should a developer choose us?
I think OpenRouter introduces another:
Why should the router choose us?
That distinction has profound implications.
If a meaningful share of demand is mediated by an intelligent routing layer, some things we've traditionally categorized as engineering suddenly become distribution.
Your latency is GTM.
Your throughput is GTM.
Your uptime is GTM.
Your price is GTM.
Your model availability is GTM.
How quickly you can support the model that went viral yesterday is GTM.
In a routed market, infrastructure performance doesn't merely give marketing something to talk about.
Infrastructure performance determines whether you get the traffic.
There is something wonderfully meritocratic about that.
A technically excellent inference provider may be able to win substantial volume without first building the world's biggest developer brand.
But there is also a catch.
You can win the request without winning the customer.
Who owns the customer?
This may be the question I would worry about most if I were running an inference company today.
OpenRouter can be an extraordinary distribution channel for inference providers.
But distribution and customer ownership aren't the same thing.
The company aggregating demand sees something individual suppliers cannot.
It sees which models are gaining share.
It sees which providers actually perform.
It sees how elastic customers are to price.
It sees where capacity is constrained.
It sees when demand shifts.
It sees which workloads tolerate cheaper models and which don't.
And most importantly:
It can move the demand.
I've spent enough of my career in technology platforms to recognize the significance of that position.
Suppliers compete inside the market.
The market maker learns from all of them.
That doesn't mean inference providers should resist OpenRouter. Quite the opposite. If that's where developers are, that's where providers need to compete.
But it does mean providers should be brutally clear about what they uniquely own.
Is it proprietary serving technology?
Superior economics?
Hardware access?
Enterprise relationships?
A specialized workload?
Operational reliability?
The customer?
Because "we have GPUs and serve these models" becomes a much weaker differentiation if someone else owns the interface through which customers can instantly choose among twenty companies that do the same thing.
I've seen value move up the stack before
There is a reason this feels familiar to me.
I worked on Google Cloud during the early rise of Kubernetes.
Kubernetes didn't make compute less important.
It made the infrastructure underneath it easier to abstract.
Instead of thinking primarily about individual machines, developers could increasingly think about workloads and let another layer determine how the underlying resources should be organized.
The compute remained essential.
But some of the strategic value moved upward toward the orchestration layer.
I think we may be watching a version of that happen again.
The GPU isn't becoming irrelevant.
The inference engine isn't becoming irrelevant.
The inference provider isn't becoming irrelevant.
There will be enormous companies built at all three layers.
But as the number of models, providers and workloads explodes, the complexity of deciding among them explodes too.
Someone has to make sense of that complexity.
And abstractions have a habit of becoming strategically important precisely because the things underneath them become abundant and complicated.
The battle is moving from supplying inference to organizing it
For the last few years, AI infrastructure has been dominated by a race to create inference supply.
Acquire GPUs.
Build clusters.
Optimize serving.
Support more models.
Drive down cost per token.
All of that continues.
But I think OpenRouter's rise—and now its reported acquisition by Stripe—signals that another race has begun.
Who organizes all of that supply?
Who aggregates the demand?
Who measures the market?
Who determines where workloads go?
Who handles the transaction when they get there?
Those may ultimately be as important as the infrastructure questions underneath them.
I've spent much of the last few years thinking about how an inference provider wins.
How do you make the GPUs faster?
How do you make inference cheaper?
How do you differentiate when competitors can serve the same models?
How do you convince developers to send you their tokens?
The OpenRouter acquisition has me thinking harder about a different question:
What if the most powerful company in the inference market isn't ultimately an inference provider at all?
Maybe it owns no GPUs.
Maybe it trains no models.
Maybe it doesn't need to predict which inference cloud wins.
It simply sits between all of them and the customer.
It sees the whole market.
It moves the demand.
It facilitates the transaction.
And it decides where the next token goes.
If inference really is becoming a market, that may be the most valuable place in the stack to be.
—
Ryan Pollock is the founder of FrontierGTM. He has spent the past two decades bringing emerging infrastructure and developer technologies to market, including AI inference at Together AI and FriendliAI, cloud infrastructure at DigitalOcean and Vultr, and Kubernetes at Google Cloud.
