Architect Financial Technologies has launched Liquid Inference, a system that auctions off every LLM request in real time. Providers bid to serve prompts, and the buyer pays the lowest offer that satisfies their rules. Developers simply swap a base URL, keep their existing code, and let the market decide the price.
In this article
What is Liquid Inference
Liquid Inference operates as an exchange-style router for LLM inference. According to Architect, providers post offers to serve specific models. Each request is auctioned across every provider quoting the named model. The lowest-priced offer that meets the buyer’s rules wins.
The product comes from a trading firm, not an AI lab. Architect runs the AX perpetual futures exchange. In May 2026 it acquired a US Designated Contract Market to list GPU compute futures, pending regulatory review. The team used its experience building financial exchanges to create two-sided price discovery for inference.
How does the inference auction work
The flow has four steps:
- Request: A client sends a standard OpenAI or Anthropic API call.
- Rules: The buyer’s constraints filter eligible offers.
- Auction: Providers quoting that model compete. The lowest qualifying offer wins.
- Receipt: The max price is locked before generation. Billing covers metered usage only.
Buyers can set per-job cost caps, time to first token limits and minimum throughput. They can also require approved regions, zero data retention and provider or model allow lists. An Auto mode can pick the model for a given unit of work. Harrison’s LinkedIn post adds routing rule presets and full multi-modal support.
Account holders can view live order books, per-provider and per-model quotes, and cleared transactions. That level of market data is unusual for an LLM API.
What do buyers get
- Drop-in compatibility with agentic coding tools. The post lists Claude Code, Codex, OpenCode, Cursor, Pi and Cline.
- Free email signup. The first 500 users get $20 of free inference, per Harrison.
- A referral program: 20% of referred fees as free inference, plus 10% on second-level referrals.
What do inference providers get
Providers onboard through the Liquid Inference app. Harrison says new providers are verified in minutes, not weeks. All prompts use the OpenAI API standard. A REST and WebSocket API registers models and quotes.
Providers can update quotes based on their own costs. That lets them sell spare GPU capacity only when they want. Payouts run through Stripe, with itemized records of every job.
How does Liquid Inference compare with OpenRouter and Hugging Face
| Feature | Liquid Inference | OpenRouter | Hugging Face Inference Providers |
|---|---|---|---|
| Routing model | Per-request auction across quoting providers | Price-weighted load balancing, inverse square of price | Fastest provider by default; :cheapest suffix optional |
| Model / provider count | “Hundreds” of models; providers not disclosed | 500+ models, 80+ providers | 18 listed partners |
| API compatibility | OpenAI and Anthropic-compatible | OpenAI-compatible | OpenAI-compatible, chat only |
| Price cap | Max price locked before first token | max_price parameter | Not disclosed |
| Data controls | ZDR, regions, allow lists | zdr, data_collection, only/ignore | Provider preference order |
| Platform fee | Not disclosed | 5.5% card credit fee, $0.80 minimum | No markup |
| Free credits | $20 for first 500 users | Not disclosed | $0.10/month free, $2.00 PRO |
| Public market data | Live order books and cleared trades | Not disclosed | Not disclosed |
What it means
Developers gain a way to bypass fixed pricing tiers. Tools like Cursor or Claude Code can now route work to the cheapest available GPU without rewriting their logic. Users see exactly what they pay for each job, rather than a blended rate. This setup treats inference like a commodity, letting market forces determine cost instead of a single vendor.



