Google reveals PageBreak, a Gemini-based security agent that found 500+ confirmed XSS bugs in its own web apps
On Sept 24, 2026 Google's Product Security team described PageBreak, an internal agent built on Gemini 3.1 Pro and Gemini 3.5 Flash that hunts vulnerabilities in Google's first-party web apps. It has found more than 500 cross-site scripting (XSS) bugs. It reports a bug only after a hand-written validator actually executes an exploit against a running copy of the app, which Google says gives a "near-zero false positive rate". Apps built on Google's hardened frameworks yielded only two XSS bugs.
Key facts
- Announced Sept 24, 2026 on the Google blog by information security engineer Michał Bentkowski; a companion Bug Hunters post details real findings
- Models: Gemini 3.1 Pro and Gemini 3.5 Flash; the design is described as model-flexible
- Timeline: pilot November 2025, full project launch January 2026
- Result: 'over 500 Cross-Site Scripting (XSS) vulnerabilities across Google first-party web applications'
- Deterministic validators for XSS, SQL injection, path traversal, RCE and SSRF run real payloads before anything is reported
- As of Sept 4, 2026 it found exactly two XSS bugs across hundreds of apps built on Google's high-assurance frameworks, both in internal apps or debug endpoints
- Google says it also found complex flaws such as cache poisoning and cryptographic bypasses that engineers and external bug hunters had missed; fixes are to be generated more and more by the CodeMender agent
What happened
Google's Product Security team published a description of PageBreak, an agent that has scanned Google's own web applications since January 2026 (after a pilot in November 2025). Gemini models reason about each app and propose likely vulnerabilities. Each hypothesis then goes to a specialised, hand-written validator, which runs a real payload against a live replica of the app. For XSS, for example, it injects JavaScript, loads the page in a rendering harness and checks whether the code runs. Only confirmed exploits are reported.
Google says PageBreak found more than 500 XSS vulnerabilities, some of which engineers and external bug-bounty hunters had missed. In applications built on Google's hardened, high-assurance web frameworks it found only two, both in internal apps or debug endpoints.
Why it matters
Most LLM bug hunters suffer from high false-positive rates. PageBreak shows the pattern that makes them usable at scale: the model proposes and a deterministic check confirms. The low count in hardened frameworks is also evidence that secure-by-design frameworks hold up even against automated AI attackers. This follows other lab reports in 2026 about AI agents finding vulnerabilities in large numbers.
Changelog
- 2026-09-30: created (from leads queue, lab-blog audit)
Related events
Sources (4)
- officialGoogle blog: Agentic hacks, real proofs: inside Google's PageBreak project
- officialGoogle Bug Hunters: Google's PageBreak Project, real-world findings
- pressTech City Authority: Google's PageBreak found over 500 XSS bugs. The trick was refusing to trust the AI's word
- pressCryptopolitan: AI agent PageBreak finds 500+ Google web bugs, just 2 in hardened apps
id: 2026-09-24-google-pagebreak-ai-xss-hunter · updated 2026-09-30 · open in the interactive timeline