It tests like an attacker. It reports like an honest one.
An autonomous pentester that runs inside a cage, with a human on the approve button. Safe to point at production. It never claims it tested what it did not. Every finding arrives with a grounding receipt you can replay.
The moat is not the scanning. It is the honesty.
Three ideas do the work. Each one is plain language first, then the mechanism underneath.
Caged
It cannot touch anything outside your scope.
Every tool runs inside a sandbox, gVisor, with a hard egress allowlist. Off-scope traffic is dropped at the network layer, not asked nicely. The agent literally cannot reach what you did not authorize.
Co-piloted
A human approves every real action.
Before anything intrusive, the run pauses at an approval gate. An operator hits Approve, Skip, or Kill. There is a budget cap and a kill switch. It is autonomous discovery with a human on the trigger.
Coverage-honest
It tells you what it did not test.
Every finding is grounded in a real tool receipt. No receipt, no "proven." Untested surfaces are stated up top in the report, not buried. It would rather say not_tested than lie to your client.
Safe against production. Honest about coverage.
A pentest should not become the incident. This one is built so the test stays inside the lines, and so the report never overstates what happened.
- Safe against production. Caged execution and egress control mean it cannot stray outside scope, even by accident.
- Approval gates on risky moves. An operator approves, skips, or kills every intrusive step. Budget cap and kill switch always armed.
- It never claims untested coverage. No grounding receipt means the finding is not marked proven. What it did not reach is declared, not hidden.
- Your own API key. No ToS gray area. Pilots are operator-supervised and free to start.
Every finding is a receipt, not a claim.
This is what lands in your report. The request, the response, the reasoning, and a replay you can run yourself. If it is not grounded, it does not ship as proven.
A scanner hands you a list. This hands you proof.
Typical scanners optimize for a long findings list. That is easy to generate and hard to trust. We optimize for the opposite.
Request early access
We onboard a handful of teams at a time so we can watch every run. Tell us where to reach you.
Pilots are operator-supervised and free to start. You use your own API key, so there is no ToS gray area. You get a scoped test plan before anything runs.
Short answers, plainly put.
Is it really safe to run against production? +
What is a grounding receipt? +
What does not_tested mean in the report? +
Does this replace human pentesters? +
Whose API key does it use? +
Test what you built. Prove what you found.
Get an early-access seat on the AI pentester, or talk to an operator about a supervised pilot on your stack. Either way, you leave with proof.
Secure Sleuths