Hacking SAML with Claude Code news update
Intermediate data structures, gadget definitions, and scope evaluation rules are structured as JSONL records to prevent token waste on out-of-scope code paths.
By Dillip Chowdary β’ Oct 10, 2026 β’ Source: oblique.security
In an empirical security research report published by oblique.security's report, security researcher Niels Provos's methodology and Anthropic's Cyber Verification Program were leveraged to evaluate SAML protocol implementations using Claude Opus. Following acceptance into Anthropic's program to remove standard safety guardrails, the researcher constructed an automated hacking harness to generate exploits across every accessible SAML implementation. Over a month of testing, the research pipeline uncovered full authentication bypasses in major projects, dozen auxiliary component signature bypasses, and unauthenticated denial-of-service vulnerabilities across multiple programming language ecosystems.
This news update examines the architectural design of the automated SAML exploitation harness built around Claude Opus, the specific vulnerabilities discovered across target frameworks, and the defensive recommendations for development teams handling SAML assertion processing. The analysis details how model-driven vulnerability research operates at scale, evaluating the performance metrics of AI-generated exploits, emergency patching responses across target maintainers, and instructions for integrating security analysis harnesses within enterprise workflows.
What shipped in Hacking SAML with Claude Code news update
The research release details an open-source automated vulnerability discovery harness engineered around Anthropic's Claude Opus model, hosted on GitHub under oblique-security/saml-research. Operating on an upgraded Claude Max 20x plan, the system breaks technical security auditing into two distinct operational phases to scale context management: a gadget discovery phase that identifies anomalous parsing primitives in XML libraries, followed by a findings confirmation phase that synthesizes identified gadgets into verified end-to-end exploits. Intermediate data structures, gadget definitions, and scope evaluation rules are structured as JSONL records to prevent token waste on out-of-scope code paths.
Rather than relying on explicit pattern-matching instructions for historical vulnerabilities, the harness feeds high-level threat models and general guidance to Claude Opus, allowing the model to independently explore attack surfaces. The automated pipeline discovered four zero-day authentication bypasses: an XML comment injection flaw in NameID handling for Authentik tracked under CVE-2026-57580, signature wrapping vulnerabilities on Response messages in PHP's litesaml/lightsaml package under CVE-2026-63182, signature wrapping in OneUptime issue 2949, and signature wrapping in Java's saml-client project under issue 149. Additionally, twelve auxiliary signature bypasses were confirmed in non-Response messages such as AuthnRequest, AttributeQuery, and LogoutRequest, including an information disclosure vulnerability in TypeScript's samlify project.
What improved in Hacking SAML with Claude Code news update
Automated security research efficacy improved through structured multi-agent coordination, intermediate data storage, and strict scoping rules that optimize LLM token usage during exploit generation. In testing against Node's xml-crypto library, the harness detected canonicalization flattening (gadget g-0005) where processing instructions were stripped to bare strings, which the confirmation phase automatically escalated into a functional email truncation authentication bypass. Memory management analysis during assertion validation led to accepted upstream patches in Go's xmldsig package to fix quadratic memory allocation during signature verification, while identifying active unauthenticated out-of-memory denial-of-service vectors in Node's xmldom library and Python packages utilizing libxmlsec1 XSLT transformations.
| Metric / Dimension | Baseline SAML Audit Method | Claude Opus Harness Method |
|---|---|---|
| Vulnerability Discovery Scope | Manual single-library code review | Multi-ecosystem automated JSONL pipeline |
| Primary Authentication Bypasses | Historical quarterly manual disclosures | 4 zero-day authentication bypasses in 30 days |
| Non-Response Signature Vulnerabilities | Rarely audited auxiliary messages | 12 confirmed project bypasses across ecosystems |
| Target Fix Cycle (OneUptime) | Slow manual issue negotiation | 3 consecutive automated AI-generated PR cycles |
Maintainer remediation responses varied across ecosystems following vulnerability submission. For the Authentik vulnerability, eight independent security researchers submitted reports concurrently using AI-assisted analysis tools. In the OneUptime repository, public disclosures triggered rapid automated AI-driven pull requests (#2949, #2981, and #2988) to resolve signature bypasses via signed error responses and processing instruction manipulations.

What you gain from Hacking SAML with Claude Code news update
Security engineering teams gain empirical evidence regarding the speed and thoroughness of model-driven protocol auditing, demonstrating that modern LLMs can discover complex logic bugs without explicit attack templates when provided appropriate harness primitives. Organizations deploying SAML authentication implementations obtain actionable vulnerability data across Go, Python, Node, PHP, and Java ecosystems, enabling immediate patching against unauthenticated memory exhaustion attacks and assertion spoofing. Developers receive concrete guidance urging the deprecation of custom SAML implementations in favor of standardized OpenID Connect protocols or hardened, vendor-maintained library APIs.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
Enterprise security operations teams gain a reusable architectural blueprint for enterprise vulnerability discovery, utilizing JSONL structured pipelines to filter false positives prior to triggering execution-heavy exploit validation steps. The research highlights critical operational realities regarding open-source security maintainability, establishing that automated slop reports and high-volume AI vulnerability disclosures require formalized, automated triage mechanisms to prevent maintainer burnout.
How to get Hacking SAML with Claude Code news update
To deploy Claude Code and Anthropic models for security analysis and protocol research, update the CLI tool directly through standard package management channels. Execute the update command to ensure access to the latest Opus model capabilities:
npm install -g @anthropic-ai/claude-codeTo switch models within an active interactive Claude Code terminal session to Claude Opus or specific research snapshots, execute the model configuration command:
/model claude-opus-3-5To launch a security analysis session with a pinned model selection from the system shell command line, use the model flag directly:
claude --model opusTo establish a persistent default model selection across all CLI executions, configure your local environment settings file or export the primary model environment variable:
export CLAUDE_DEFAULT_MODEL="claude-opus-3-5"What to watch after Hacking SAML with Claude Code news update
Security teams must track ongoing patching efforts across Python and Node SAML dependencies affected by unauthenticated XML denial-of-service vectors. Python implementations relying on libxmlsec1 require immediate monitoring for XSLT transform filter patches to block recursive template expansion attacks. Development teams using Node libraries dependent on xmldom should monitor security advisories for the resolution of private memory allocation reports.
Furthermore, enterprise open-source maintainers are expected to implement updated contribution guidelines and automated verification steps to handle the rising volume of AI-generated security disclosures. As multi-agent hacking harnesses become widely accessible, protocol maintainers must prepare for concurrent multi-reporter disclosures on complex XML signature wrapping vulnerabilities across public repositories.
Developer Action Items
- β Inventory whether Anthropic / Claude / GitHub runs in prod, CI, staging, or on laptops before you debate severity.
- β Pull the vendor advisory for CVE-2026-57580, CVE-2026-63182 and patch from that page β not from a social recap.
- β If you cannot patch today, isolate the service, rotate tokens that sat on the affected surface, and raise the logging floor.
- β Record the decision and residual risk so the next on-call does not re-litigate whether you are exposed.
Hacking SAML with Claude Code news update FAQ
What models and tools were used to hack SAML implementations in this research?
The researcher used Claude Opus through Anthropic's Cyber Verification Program alongside a custom multi-agent hacking harness written in Go, hosted at github.com/oblique-security/saml-research.
Which major SAML projects were found to have authentication bypass vulnerabilities?
Full authentication bypasses were identified in Authentik (CVE-2026-57580), PHP litesaml/lightsaml (CVE-2026-63182), OneUptime (issue 2949), and Java's saml-client (issue 149).
What denial-of-service vulnerabilities were uncovered across language ecosystems?
Unauthenticated memory exhaustion vectors were identified in Go's xmldsig due to quadratic allocations, JavaScript's xmldom library, and Python packages that fail to filter XSLT transforms passed to libxmlsec1.
Sources
Author
Dillip Chowdary
Writes Tech Bytes coverage of AI, engineering, and the tools that actually ship. Editor of Tech Pulse Daily.
Related on Tech Bytes
Unlocking hidden revenue streams with market models
Read β
Astropad Workbench 1.3 adds faster streaming, privacy curtain, and more
Read β
Camera-equipped AirPods reportedly wonβt launch in 2026, despite demo video leak
Read β
The Open-Sourcing of DeepSeek Harness Opens the Door to Modular, Unbundled AI Agentβ¦
Read β
Today's Tech Pulse briefing
Full briefing β
Advertisement