SENTINELPROXY
PRODUCT BRIEF · DETECTION BENCHMARK · SEPTEMBER 2026
AI Firewall · Standard tier · 33-attack adversarial suite
100%
Detection rate
33 of 33 attacks caught

The AI firewall
that measures itself
then ships the receipts.

Sentinel runs an automated adversarial benchmark across eight attack categories before every release. This is the September 2026 result on standard tier — 33 of 33 attacks caught, 0 false positives across 8 clean controls.

Audience

Security, platform, and ML engineering teams

Prepared by

Sentinel Proxy · sentinelaifirewall.com

01 / 04
Architecture

Four defensive layers between every prompt and your model — resolved in under 100ms.

L1
Decoder pass
Decodes and re-scans common encoding wrappers before any scoring runs: Base64, hex, URL-encoding, ROT13, and Morse code. What the LLM would have decoded, we decode first, so an attack hidden inside an encoding doesn't get a free pass.
1–2msdecode
L2
Fast-path patterns
Our library of high-confidence regex patterns covers authority hijack, persona override, prompt extraction, and instruction injection. A match here returns immediately, near-zero latency.
5–10msfast path
L3
Semantic vector match
An embedding comparison against our library of attack signatures, stored in a vector database. Catches paraphrased and novel variants the fast path misses, and grows continuously through automated red/blue-team drills.
~40mssemantic
L4
Threat score & tier thresholds
Composite 0.00–1.00 score routes the request to one of four verdicts. Tier thresholds (standard vs strict) let customers dial false-positive tolerance per workload.
~10msscore
CLEANPass through unmodified
FLAGGEDPass, but flag for app logic
NEUTRALIZEDSanitise, then forward
BLOCKEDEmpty payload, do not forward

Coverage of the prompt path

Input scrubbing · /v1/scrub

Single content scan. The core endpoint — everything else builds on it.

RAG batch scan · /v1/scrub/batch

Up to 100 chunks per call. Catches poisoned context at ingestion or retrieval time.

Agentic passthrough · 4 providers

One base URL swap for Anthropic, OpenAI, Grok, or Gemini. Tool-result content is scrubbed before the model reads it, on every tier.

PII detection · US & EU identifiers

Cards (Luhn), SSNs, emails, US phones, IBANs, UK NINs, EU VAT (DE/FR/IT/ES/NL). Off / Flag / Redact modes.

Slopsquatting checks · Pro+

Scans LLM output for hallucinated package names against live PyPI and npm registries before a developer installs one.

Automated red/blue drills

Continuous adversarial testing finds blind spots and grows the signature library over time.

Detection benchmark · September 2026

33 adversarial attacks. 8 clean controls. Here's the result.

100%
Detection rate
33 of 33 attack samples caught across 8 categories.
0%
False positive rate
0 of 8 clean samples flagged. Real prose stays clean.
<100ms
Median latency
Fast path returns in 5–10ms; full scan well under 100ms.
Attack category
Caught
Distribution
HTML injection (hidden div, white-on-white, meta, comments, B64, Unicode, exfil, DAN)
8 / 8
Authority hijack / system override
4 / 4
Persona override / jailbreak (DAN, opposite-mode, fictional framing)
4 / 4
Prompt / instruction extraction
4 / 4
Encoding obfuscation (ROT13, hex, Base64, Morse — all decoded and caught)
4 / 4
Exfiltration instructions
3 / 3
RAG / indirect injection
3 / 3
Social engineering (grandma, security-researcher, hypothetical framing)
3 / 3
“Sentinel detected 33 of 33 attack samples (100% detection rate) across 8 attack categories, with a 0% false positive rate on 8 clean controls (business email, code, support queries, product descriptions, meeting notes, technical docs, news summaries, recipes) in standard mode.”
— Automated benchmark, September 2026
The one we used to miss

We found a bug. We fixed it.
Here's the receipt.

Earlier this year, one attack in our adversarial suite slipped through in standard tier: a Morse-encoded command modelled on the Grok / Bankrbot $174K heist. Worth reading carefully:

Sentinel decodes and inspects Morse, ROT13, hex, Base64, and URL-encoded payloads, but a bug in the Morse decoder bailed on a malformed token instead of skipping it, so that one payload was never decoded, and never scored. We caught it in our own benchmark, patched the decoder, and re-ran the full suite before publishing any number.

It now blocks at score 1.00, same as the rest of the encoding suite. All 33 of 33 attacks are caught in the September 2026 run.

Fixed
Morse decoder
bug patched

Coverage compounds, it doesn't stay fixed.

Every miss found in a red/blue-team drill becomes a fix in the decoder or a new signature in the library, so the same class of attack (not just the exact sample) gets caught next time. This is also how Sentinel keeps pace with newly reported real-world incidents like the Grok / Bankrbot heist above.

Run the benchmark on your own data.

We'll set you up with a sandbox key, share the test harness, and walk through results together. Honest numbers, no slide-decks.

Book a call
calendly.com/skyblueguru