Skip to content
Public beta

One API.Every model.

Point your app at one endpoint. Open models, models tuned for a domain and your own: every request goes to the cheapest one that clears your quality bar.

Route.

Every request scored before inference, not after a failed attempt.

Predictive routing.

Capability, cost and residency are scored before the call. The cheapest model that clears your floor is the one that runs.

POST /v1/chat/completionsscored
qwen-3-32bAlibaba0.91$0.0018routed
kimi-k2Moonshot AI0.88$0.0031
mistral-smallMistral AI0.84$0.0011
quality floor 0.86

Automatic failover.

If a provider degrades, rate-limits or goes dark, traffic moves down the chain in real time and your code never notices.

one request200 OK
PrimaryFireworksdegraded
SecondaryAmazon Bedrockcarrying
Fallbackyour endpointstandby

Primary degraded. Rerouted, and your code never noticed.

Rules you can read.

The routing table is a list, not a black box: the condition, the decision it forces, and what it carried in the last day.

Routing rules

Scored before inference. The first rule that matches decides.

4 rules
RuleWhenThen24h
summarise · costtask = summarisecheapest clearing 0.86641
code · qualitytask = codehighest score, cost ceiling $0.004292
private-datadata = restrictedyour endpoints only149
bulk-embedtype = embedcheapest available0

Account.

What every call actually cost, computed against a versioned price table.

Every request, logged.

Status, tier, latency, tokens and cost on every call, the failures included, because a gateway that hides them is not accounting.

Logs

Every request, the tier it was served at, and what it cost.

Live
TimeTypeTierLatencyTokensCost
14:22:08chatbalanced4121,896$0.0018
14:22:07chatfast308944$0.0011
14:22:05chatbalanced--$0.0000
14:22:05chatfast2661,204$0.0009
14:22:03embedembed88512$0.0001
14:22:01chatdeep4662,310$0.0027

Estate.

Every model you are approved to use, arriving at one endpoint.

The approved estate.

Open weights, models tuned for a domain, and your own private endpoints, each admitted on evidence and a signature.

Total Providers7
Approved Models120
Total Requests (7d)1,429
Total Cost (7d)$3.19

Model Catalogue

Overview of all configured providers, models and usage.

All Providers
ProviderModelsTraffic (7d)Cost (7d)
Moonshot AIkimi-k2, kimi-k2-turbo403$0.96
Alibabaqwen-3-32b, qwen-3-coder335$0.70
Mistral AImistral-small, mistral-nemo236$0.54
MiniMaxminimax-m2172$0.39
Zhipu AIglm-4.6146$0.29
Neural Arcfr-lex-1.7b, fr-forge-1.7b, fr-blaze-9b96$0.25
Your endpointsyour-model-1, your-model-241$0.06

One line to adopt.

Drop-in compatible with the OpenAI API. Point the official SDK at our base URL and send auto as the model.

client.py

OpenAI SDK
- base_url="https://api.openai.com/v1"
+ base_url="https://api.modelbeat.ai/v1"
api_key=MB_KEY

Point the official SDK at our base URL and send auto as the model.

For developers and builders.