One API.Every model.
Point your app at one endpoint. Open models, models tuned for a domain and your own: every request goes to the cheapest one that clears your quality bar.
Route.
Every request scored before inference, not after a failed attempt.
Predictive routing.
Capability, cost and residency are scored before the call. The cheapest model that clears your floor is the one that runs.
Automatic failover.
If a provider degrades, rate-limits or goes dark, traffic moves down the chain in real time and your code never notices.
Primary degraded. Rerouted, and your code never noticed.
Rules you can read.
The routing table is a list, not a black box: the condition, the decision it forces, and what it carried in the last day.
Routing rules
Scored before inference. The first rule that matches decides.
Account.
What every call actually cost, computed against a versioned price table.
Every request, logged.
Status, tier, latency, tokens and cost on every call, the failures included, because a gateway that hides them is not accounting.
Logs
Every request, the tier it was served at, and what it cost.
Estate.
Every model you are approved to use, arriving at one endpoint.
The approved estate.
Open weights, models tuned for a domain, and your own private endpoints, each admitted on evidence and a signature.
Model Catalogue
Overview of all configured providers, models and usage.
One line to adopt.
Drop-in compatible with the OpenAI API. Point the official SDK at our base URL and send auto as the model.
client.py
OpenAI SDKPoint the official SDK at our base URL and send auto as the model.
For developers and builders.
