Every model call.
One governed front door.
Point every app, agent, and developer at a single OpenAI-compatible endpoint. The gateway decides who may call what, on whose budget, in which region, and whether the payload deserves a public model at all.
- Routes across every major provider and your own sovereign appliances
- Per-team virtual keys with budgets, quotas, and spend attribution
- Data-loss policy federated to the edge, evaluated before dispatch
- A receipt for every request: who, what, where it ran, what it cost
/ How it works
Four decisions,
before a single token leaves.
Identify
Virtual keys resolve to a team, an org, and a budget. Unattributed traffic does not pass.
Evaluate
Budgets, quotas, model allow-lists, and data-loss policies run before dispatch, not after.
Route
Cloud, private, or sovereign: the destination is a policy decision, including reroutes for sensitive payloads.
Receipt
Every request lands in the ledger with caller, model, region, latency, and cost attributed.
/ Capabilities
Everything this layer holds.
OpenAI-compatible surface
Apps integrate once. Behind the endpoint, models can be swapped, pinned, or rerouted without touching application code.
Model allow-lists and pinning
Approve models per team, pin versions for workloads that must not drift, and retire providers without a migration project.
FinOps built in
Budgets, quotas, forecasts, and per-team attribution as gateway primitives. A budget breach degrades to a designated fallback, not an outage.
DLP policy federation
Data-loss rules authored once and enforced at the gateway, shadow-first so you observe impact before you enforce.
Residency-aware routing
Route by data class and geography. Payloads that must stay in country never leave it, by policy rather than promise.
Guardrail hooks
Runtime Security inspects traffic in line with the gateway. An unverified scan is treated as unverified, never as an allow.
/ The difference
A proxy forwards.
A gateway governs.
Plenty of products sit between you and a model. The difference is what they are allowed to decide while they are there.
Inspection happens in the vendor's cloud, in a region you do not choose.
Policy and inspection run where you deploy the gateway. The prompt's path is your decision, region included.
A blocked request is the end of the story. The user finds another way.
Sensitive-but-legitimate traffic reroutes to your sovereign appliance. The user still gets an answer.
Spend shows up as one undifferentiated bill at the end of the month.
Every request is attributed to a team and a budget at call time, with forecasts that refuse to extrapolate from thin data.
Model versions drift underneath production workloads.
Pinned model routes. What ran yesterday is what runs today, and the receipt proves which weights answered.
/ The rest of the stack
One layer is a feature. Six is a fabric.

Ready to run on WIT OS?
Talk to the team about a managed deployment, a pilot, or a custom agent. We typically respond within an hour.
