Your models. Your metal.
Your borders.
A GPU appliance in your rack or your region, serving curated open-weight models the stack can route to at any moment. It is where sensitive prompts go instead of a public model, and the reason a decision made today can be re-run years from now.
- Single appliance to in-region HA pairs, sized to your workload
- Curated model catalog served on pinned, versioned weights
- Private-first routing with controlled cloud failover
- No per-token meter. Inference on your metal has no toll booth
/ How it works
Private first.
Cloud only by choice.
Classify
Workloads are classed by audience and sensitivity: internal extraction and triage differ from customer-facing prose.
Route private-first
Classes you designate run sovereign by default, with an explicit, logged fallback if the appliance is unreachable.
Serve on pinned weights
The appliance answers from a curated catalog with versioned weights. What answered is recorded, not assumed.
Anchor the record
Every sovereign response carries its engine fingerprint into the decision ledger, making later replay meaningful.
/ Capabilities
Everything this layer holds.
Appliance form factor
A hardened GPU box, not a Kubernetes project. Compose-level operations your team can run, from one node to an HA pair.
Data that never leaves
Prompts routed sovereign are processed inside your walls or your jurisdiction. Residency by architecture, not by contract clause.
Curated model catalog
Current open-weight models for extraction, triage, summarization, code, and chat, tuned and pre-warmed for the appliance.
Engineered for throughput
Serving stack tuned per GPU generation. Our lab appliance sustains 130+ tokens per second single-stream on a 14B-class model, and over 1,000 aggregate at ten concurrent streams.
The Sovereign Reflex
When Runtime Security flags a prompt as sensitive but legitimate, this is where it lands. Users keep working, data stays home.
Foundation for replay
Pinned weights on hardware you control are what make Decision Records replayable. Frontier vendors cannot offer this: they retire the weights.
/ Numbers we can defend
Lab numbers,
labeled as lab numbers.
Measured on our reference lab appliance (single current-generation GPU, quantized 14B-class model, tuned serving kernel), August 2026. Your throughput depends on model size, quantization, context length, and GPU generation; sizing is part of every engagement. The A/B figure compares end-to-end latency on internal extraction and triage tasks against a frontier cloud model on identical prompts.
/ The difference
The cloud rents you intelligence.
This is the part you own.
Every token in and out is metered, and the meter only runs up.
Inference on your appliance has no marginal cost. Heavy internal workloads stop being a budget conversation.
Model versions are retired on the vendor's schedule. Yesterday's model is gone.
Weights are pinned on hardware you control. The model that made a decision stays available to re-run it.
Data residency is a compliance document.
Data residency is a routing rule enforced by the gateway and a box that physically sits where you put it.
Sovereign offerings mean a private region of someone else's cloud.
An appliance in your rack, or an HA pair in a facility you choose, operated with your keys.
/ The rest of the stack
One layer is a feature. Six is a fabric.

Ready to run on WIT OS?
Talk to the team about a managed deployment, a pilot, or a custom agent. We typically respond within an hour.
