Something you trust
reads it first.
In the last hour, your AI read a thousand things no human will ever see. Runtime Security is the layer that reads them first: every prompt, tool call, and output, inspected inline before the model acts on it.
- Prompt injection, jailbreak, secret, and PII detection inline
- Invisible-payload detection: decoded, de-obfuscated, then judged
- Semantic data-loss analysis beyond regex, including audio and video
- MITRE ATLAS-aligned threat coverage, refreshed weekly
/ How it works
Inspect. Decide.
Then prove you did.
Inspect
Multi-engine scan: pattern, ML classifier, and semantic analysis across prompts, tools, and outputs.
Decide
Allow, flag, or block. A flag is a destination: sensitive traffic reroutes to sovereign inference instead of dying.
Degrade honestly
Under overload the engine says so: verdicts carry a checked flag, and unverified traffic is treated as unverified upstream.
Chain the evidence
Every decision lands in hash-chained audit trails, per tenant, with zero retention of scanned content.
/ Capabilities
Everything this layer holds.
Injection and jailbreak defense
Layered pattern and ML detection across direct and indirect injection, with a corpus-measured baseline we publish rather than imply.
Invisible-payload detection
Zero-width characters, homoglyphs, and encoded instructions are decoded and re-judged. Six finding types, mapped to MITRE ATLAS.
Secrets and PII
Keys, tokens, credentials, and personal data caught in both directions: on the way to the model and on the way back out.
Semantic DLP, any medium
Meaning-level data-loss categories that survive paraphrase, applied to text, audio, and video payloads.
Multilingual by model, not regex
Attacks in Spanish, Portuguese, French, German, Chinese, and Arabic are caught by classifiers, not an English-only pattern list.
Honest failure modes
No silent bypass. If a scan did not happen, the verdict says so, and policy upstream decides: reroute to sovereign, or refuse.
/ Numbers we can defend
The baseline is public.
So are the misses.
Measured against red-team corpus v1.0.0 (135 cases: 92 attacks across 7 families, 43 benign controls), August 2026, driven at the live scan API. Detection counts flags and blocks; a flagged prompt reroutes to sovereign inference rather than reaching a public model. We publish per-family results, including the families where content inspection alone is not enough, and we close those with the Agent Security layer.
/ The difference
Detection is a score.
Containment is a destination.
Detection rates converge across the industry. What does not converge is what a product is allowed to do with a verdict.
A sensitive prompt is blocked. The user pastes it into a personal account instead.
A sensitive prompt reroutes to your sovereign appliance. Answered, contained, and logged.
Under overload, the scanner quietly stops scanning and traffic flows anyway.
Verdicts carry a checked flag. Unverified traffic is rerouted or refused by policy, and the receipt records it.
Scanned content is retained to improve the vendor's models.
Zero retention per tenant. The audit chain stores decisions and digests, never your content.
Coverage claims are a static number on a datasheet.
Coverage is a versioned corpus you can re-run against your own deployment, with regression gates in CI.
/ The rest of the stack
One layer is a feature. Six is a fabric.

Ready to run on WIT OS?
Talk to the team about a managed deployment, a pilot, or a custom agent. We typically respond within an hour.
