We attack it first.
Then we publish the misses.
Red teaming for AI systems that behaves like engineering, not theater: a versioned attack corpus, driven against your live stack, scored on detection and on the false blocks that break real users, then wired into CI so the score cannot quietly regress.
- 135-case corpus: 92 attacks across 7 families, 43 benign controls
- Runs against the scan layer and end to end through the gateway
- False blocks scored separately: the number that breaks users
- Priced as an assessment. Never metered, never per scan
/ How it works
A score you can re-run
is a score you can trust.
Drive the corpus
Injection, jailbreaks, exfiltration, encodings, secrets, and six-language multilingual attacks against your live deployment.
Score honestly
Detection, flag versus block, and false-block rate on benign traffic, with latency percentiles alongside.
Gate regressions
The corpus runs in CI with exit gates. A fix that costs you coverage fails the build before it costs you an incident.
Fix and re-run
Findings map to the layer that owns them: patterns, policy, or the authorization gate. Then the same corpus proves the fix.
/ Capabilities
Everything this layer holds.
A corpus, not a vibe
Versioned and fingerprinted. Every claim traces to a case ID, so two runs a year apart are actually comparable.
Two attack surfaces
Direct against the scanning layer, and end to end through the gateway, because the wiring between them is where real gaps hide.
False blocks are findings
A benign prompt that gets blocked sends a user to shadow AI. We measure that rate with the same rigor as detection.
Multilingual pressure
Attacks in Spanish, Portuguese, French, German, Chinese, and Arabic, because English-only regex is not a security model.
Misses are published
The report names the families where content inspection is structurally insufficient and prescribes the layer that closes them.
Reports you can paste
Customer-ready markdown with methodology, per-family results, and deltas against your previous run.
/ Numbers we can defend
Our own baseline,
misses included.
Corpus v1.0.0 baseline, August 2026, against our own live stack. The exfiltration number is the honest one: an instruction like forwarding invoices to an outside address is indistinguishable from legitimate finance work at the content layer. That gap is architectural, and it is exactly what the Agent Security layer's argument-level authorization exists to close. Vendors who do not publish a number like this have one anyway.
/ The difference
Assurance you rent
versus assurance you keep.
Red teaming priced per agent, per scan, per month. Testing more costs more.
An assessment with a report and a re-runnable corpus. Testing thoroughly is the point, not a billable event.
A glossy score with no way to reproduce it.
A versioned corpus and runner your team can execute again after every change, with CI gates included.
Weak families are quietly averaged into a headline number.
Per-family results, printed, including the ones we lose. The fix is prescribed, then proven by the next run.
Findings die in a PDF.
Findings map to owners: a pattern gap, a policy change, or an authorization rule, each verifiable by re-run.
/ The rest of the stack
One layer is a feature. Six is a fabric.

Ready to run on WIT OS?
Talk to the team about a managed deployment, a pilot, or a custom agent. We typically respond within an hour.
