Back to Fieldnotes

Running Google Mantis on My Own Site: Field Notes from an Agentic Security Scanner

The premise

Mantis is Google’s answer to a problem AI itself created: point an LLM at a codebase, ask for vulnerabilities, and get back a triage queue that is roughly 93% hallucination (industry true-positive rates under 7%). Its design answer is an evidence gate — a finding only survives if an agent reproduces it in a sandbox.

What shipped in September 2026 is not a product. It is a bundle of ~20 skill files (structured playbooks an agent reads) plus an ADK reference harness. Google’s internal version — the one protecting Cloud — is a fuller build on the same ideas. Think pattern release, not tool release.

The run

Target: this site’s own repo (22 files, Hugo static site). Setup: reference harness, --sandbox static-only, a flash-tier model via OpenRouter. Duration: ~45 minutes, 102 tool calls.

Stage one set the tone. The structural index classified the repo, flagged markup.goldmark.renderer.unsafe = true as the one config risk worth chasing, and mapped the diagram viewers’ JavaScript surfaces — all before any expensive model reasoning.

The finding was honest. One VALID, LOW-severity stored XSS (CWE-79): unsafe goldmark rendering plus bare {{ .Content }} emission over raw-HTML content files. LOW because a fully static repo has no external attacker path — exploitation requires write access to the content pipeline (a compromised authoring agent or CI token). The severity cap was reasoned, not defaulted.

The negative results were the better part. URL parameters whitelisted before setAttribute. No innerHTML/eval/postMessage sinks fed by untrusted input. Iframes sandboxed correctly, without allow-same-origin. Four things it checked and chose not to report, with reasons. That is the difference between a scanner and a hallucination filter.

The failure worth more than the finding

The researcher completed its analysis — then passed its report to the database as a JSON string instead of a dict. Pydantic rejected it. Seven retries. The finding never persisted; it survived only as text in the log.

Two lessons:

  1. Weak models cost more, not less. A model that cannot reliably emit tool-call schemas turns every structured stage into a retry loop. The schema-reliable model is the cheap one, whatever its sticker price.
  2. The --resume flag is the cost-saver. Completed stages checkpoint; re-run only what failed. Re-scanning from scratch because a reporting stage died would double the bill for zero added analysis.

Where it belongs

Google runs Mantis pre-submit on every code check-in, with localized threat models fed by live codebase metadata — false positives down to 3% in some cases. That is affordable per-change because a small diff needs far less context than a monolithic scan.

The rest of us do not have that machinery in the open-source harness. The pragmatic translation: risk-gated PRs (auth, parsers, untrusted-input surfaces) and merge/release gates for full runs, cheap deterministic scanners per-commit underneath. Same principle Google states — catch it before it lands — phased by budget.

Verdict

Not something to bolt onto CI as-is. Something to study, adapt, and earn its cost at the review points where a human would otherwise spend hours. The anti-hallucination architecture — reproduce-first, critic agents, calibrated severity — is the part worth stealing regardless of whether you ever run the harness.