We didn't want tokens.
We didn't want hardware.
We wanted both.
FlatCompute is what happens when a team that actually runs inference gets tired of the two extremes.
FlatCompute was born out of necessity. At Wetopi / Inte, we've been running real infrastructure since 1998. We needed inference compute that was close — not a region 2,000ms away — and we didn't want to pay per token. We wanted to run our own models, on our own terms, with predictable costs.
So we bought GPUs. We deployed vLLM. We measured throughput, tuned schedulers, and squeezed every last token per second out of our hardware. And somewhere along the way, we realized: we're not the only ones who need this.
There's a whole world of AI agent builders, autonomous pipelines, and developers with sustained, heavy inference workloads who are stuck between two bad options:
Pay per request. Watch costs spiral. Never own anything. Great for experiments, brutal for production agents that run 24/7.
Drop €40k on GPUs. Learn to operate them. Lose weekends to driver updates and thermals. Great if you're a datacenter, not if you're a team of five.
FlatCompute is the middle path. FlatCompute is the middle path. You rent a seat on real GPU infrastructure — not a virtual credit, not a rate limit, but an actual Active Worker running vLLM on bare metal. You get sustained throughput, fixed monthly pricing, and the freedom to run whatever model you want.
It's compute by the people who needed it, for the people who need it. No token meter ticking in the background. No surprise bills. Just your agents, running fast, on hardware someone else worries about.
Our goal is to deliver an average of 30 tokens per second per Active Worker, sustained. We achieve this by not overfilling a zone beyond the threshold where performance degrades below that target. Fewer seats per GPU means each seat runs faster — and the price stays as competitive as possible for everyone.
To be clear: this is not a Service Level Agreement and not a contractual commitment. It's the internal target we set for ourselves — the benchmark we design capacity around, the number we refuse to compromise on when deciding how many seats to sell per zone. If throughput drops, we stop selling seats before we let it fall further. That's the principle. Not a guarantee on paper, but a promise in practice.
Made with ♥ in BCN · 41.39°N 2.17°E
FlatCompute is a trademark operated by Inte, an organization incorporated in 1998 in Barcelona under the legal name: Inte, implantación de nuevas técnicas empresariales S.L.
N.I.F. B61386777
Business Official Registration Bureau at Barcelona, Book 29858, Page 0073, Sheet 163237, Entry 1
Dun & Bradstreet – D-U-N-S® Number: 564329605