AI Pentest Platform · Early access

It tests like an attacker. It reports like an honest one.

An autonomous pentester that runs inside a cage, with a human on the approve button. Safe to point at production. It never claims it tested what it did not. Every finding arrives with a grounding receipt you can replay.

Caged in gVisor Human approves risky steps Coverage-honest by design
Mission Control · Run 0413 operator supervised ●
scope api.acme.internal signed ROE
step 01 recon ok
step 02 auth flow map ok
Approval gate · exploit attempt
IDOR probe on /orders/{id}. Intrusive. Waiting for an operator.
ApproveSkipKill
step 03 idor /orders/{id} confirmed
finding FINDING-0413 grounded ✓
budget42% · kill switch armed
1,001Caged tool calls on a live production engagement
0Off-scope actions, the cage held every call
1Finding retracted because it could not be proven
HITLA human approved every intrusive step
01 · How it works

The moat is not the scanning. It is the honesty.

Three ideas do the work. Each one is plain language first, then the mechanism underneath.

01

Caged

It cannot touch anything outside your scope.

Every tool runs inside a sandbox, gVisor, with a hard egress allowlist. Off-scope traffic is dropped at the network layer, not asked nicely. The agent literally cannot reach what you did not authorize.

gVisorEgress allowlist
02

Co-piloted

A human approves every real action.

Before anything intrusive, the run pauses at an approval gate. An operator hits Approve, Skip, or Kill. There is a budget cap and a kill switch. It is autonomous discovery with a human on the trigger.

HITL gatesBudget capKill switch
03

Coverage-honest

It tells you what it did not test.

Every finding is grounded in a real tool receipt. No receipt, no "proven." Untested surfaces are stated up top in the report, not buried. It would rather say not_tested than lie to your client.

Grounding receiptsnot_tested declared
02 · Safety and trust

Safe against production. Honest about coverage.

A pentest should not become the incident. This one is built so the test stays inside the lines, and so the report never overstates what happened.

  • Safe against production. Caged execution and egress control mean it cannot stray outside scope, even by accident.
  • Approval gates on risky moves. An operator approves, skips, or kills every intrusive step. Budget cap and kill switch always armed.
  • It never claims untested coverage. No grounding receipt means the finding is not marked proven. What it did not reach is declared, not hidden.
  • Your own API key. No ToS gray area. Pilots are operator-supervised and free to start.
Coverage report · excerpt grounded ●
tested 14 surfaces, receipts attached
proven 3 findings with replayable PoC
not_tested: auth token refresh, admin RBAC boundary. Out of this run's scope. Stated up top, not buried.
This line is the product. Most tools quietly imply full coverage. This one shows you the gaps in plain sight.
03 · A real finding

Every finding is a receipt, not a claim.

This is what lands in your report. The request, the response, the reasoning, and a replay you can run yourself. If it is not grounded, it does not ship as proven.

Receipt · FINDING-0413 verified ●
$ agent probe --target api.acme.internal
class IDOR · broken object level authorization
gate exploit attempt approved by operator ✓
poc GET /orders/1042 → 200 (other tenant's order)
expect 403 for a non-owner principal
grounded: request + response captured · replayable · call_id 0413-a7f · severity HIGH · CVSS 8.1
replayableremediation attachedHIGH · CVSS 8.1
04 · Why this is different

A scanner hands you a list. This hands you proof.

Typical scanners optimize for a long findings list. That is easy to generate and hard to trust. We optimize for the opposite.

What matters
Typical scanner
Secure Sleuths AI Pentest
Evidence
A severity label and a description you take on faith.
A grounding receipt per finding. Request, response, replay.
Coverage claims
Implies it checked everything. Gaps stay silent.
Declares not_tested up top. No silent gaps.
Production safety
You hope the config is safe and hold your breath.
Caged in gVisor with egress control. It cannot stray.
Risky actions
Fires automatically, no one on the trigger.
Human approval gate plus budget cap and kill switch.
False positives
You triage the noise. That is your weekend.
Unproven stays unproven. If it cannot ground it, it does not call it a bug.
Early Access · Cohort 01

Request early access

We onboard a handful of teams at a time so we can watch every run. Tell us where to reach you.

Step 1You request access here.
Step 2We scope and sign an ROE together.
Step 3We run a supervised pilot. You approve every step.

Pilots are operator-supervised and free to start. You use your own API key, so there is no ToS gray area. You get a scoped test plan before anything runs.

05 · Questions

Short answers, plainly put.

Is it really safe to run against production? +
Yes, and that is the point of the cage. Every tool runs inside gVisor with a hard egress allowlist, so off-scope traffic is dropped at the network layer. On top of that, a human approves every intrusive step, and there is a budget cap and a kill switch. The test cannot quietly become the incident.
What is a grounding receipt? +
It is the evidence behind a finding. The real request, the real response, the reasoning, and a call_id you can replay. No receipt, no "proven." That rule is what keeps the report honest.
What does not_tested mean in the report? +
It is an honest declaration of what this run did not reach, stated up top rather than buried. Most tools imply they checked everything. This one shows you the gaps in plain sight, so you know exactly what is still open.
Does this replace human pentesters? +
No. It is autonomous discovery with a human on the trigger, and our operators supervise every pilot. If you want a fully human-led engagement, that is our Penetration Testing service. Same standard of proof, different speed.
Whose API key does it use? +
Yours. That keeps you in control of cost and keeps everything inside a clean terms-of-service boundary. Pilots are free to start and operator-supervised throughout.
Cohort 01 is open

Test what you built. Prove what you found.

Get an early-access seat on the AI pentester, or talk to an operator about a supervised pilot on your stack. Either way, you leave with proof.

Talk to an operator