Separating Signal from Slop: Triaging CVEs in the Age of AI Security Research

Separating Signal from Slop: Triaging CVEs in the Age of AI Security Research

Share

TL;DR
  • The volume of disclosed CVEs is climbing fast, up roughly 85% over last year while the count of known exploited vulnerabilities grew far more slowly, and AI is a big reason why.
  • Tools that can read a codebase, identify a vulnerability, and generate a working proof-of-concept have lowered the skill floor for vulnerability research. They also produce a wave of bugs that look terrifying on paper (remote code execution, popular software, no authentication,) yet they often depend on a configuration rarely enabled in real deployments.
  • Across four heavily publicized 2026 Nginx CVEs, every one required an uncommon configuration. For nginx-rift, one independent study found the vulnerable rewrite rule in just 1 of 35,633 real configs. For nginx-quicburst, our own scans found just 4 hosts running an affected version with HTTP/3 enabled across a 56,000-host sample, about 0.007%. For CVE-2026-42533, the remote code execution appeared only under an intentionally vulnerable proof-of-concept config, leaving a normal deployment with a denial of service, not code execution.
  • The lesson: A critical bug class in a popular product does not equal mass exploitation when the precondition almost never appears in production.

The AI-driven CVE Surge

In our earlier post on prioritizing emerging threats, we described the firehose of roughly tens of thousands of CVEs publishing each year, and how our threat enablement team (TEA) prioritizes these to identify which display the most real-world impact. Over the past several months, we've seen that number rise steeply.

In the first eight months of 2026, the CVE Project published 57,871 records, far above any prior year at the same point. The count of known exploited vulnerabilities rose about 14% over the same period, a stark slower incline against the 85% jump in disclosures.

Figure 1: Published CVEs beside catalogued exploitation, cumulative by month. Source: zerodayclock.com.
Figure 1: Published CVEs beside catalogued exploitation, cumulative by month. Source: zerodayclock.com.

A meaningful share of the increase is AI-assisted research. Autonomous and semi-autonomous agents now scan large codebases, identify various issues, and produce proof of concepts, often with minimal human intervention.

While AI has proven to be great at finding bugs, including ones that have sat undiscovered for years, we have observed AI-assisted research producing a high volume of CVEs with minimal real-world impact. When the rate of disclosure climbs far faster than the rate of exploitation, every defender's job gets harder as they drown trying to find the signal in the flood of potential threats. 

    Identifying CVEs with Mass-Exploitation Potential

    Our triage framework looks for a cluster of attributes that, together, signal newly disclosed CVEs that are more likely to be leveraged for mass exploitation:

    • Common in enterprise: the software is widely deployed on reachable infrastructure.
    • Code execution or privileged access: the impact is worth an attacker's time.
    • Exploited in the wild: it shows up in CISA KEV or in observed attacks.
    • Public proof of concept: exploit code already available publicly.
    • Affects the default configuration: out-of-the-box configuration is vulnerable.

    The Edge-Case Configuration Problem

    This is the heart of our prioritization framework: CVSS "doesn't tell the whole story." A critical score rates the bug, not whether the configuration it needs exists in real deployments.

    The examples below were all previous Emerging Threats that initially appeared to have mass exploitation potential, but each required a non-default configuration which reduced exploitable in-the-wild instances to near zero.

    Figure 2: Each of these CVEs is common, an RCE, and ships with a public PoC, yet none affects the default configuration or is exploited in the wild, so all three land at Tier 3 (Low Threat).
    Figure 2: Each of these CVEs is common, an RCE, and ships with a public PoC, yet none affects the default configuration or is exploited in the wild, so all three land at Tier 3 (Low Threat).

    An Nginx Case Study: AI-Discovered CVEs That Are All Bark and No Bite

    Four Nginx advisories crossed our triage queue in the first half of 2026, each one followed by popular news sites amplifying the research and concerned customers reaching out about their exposure to these new vulnerabilities. Given how common Nginx is, we cautiously reviewed the details of each.

    Upon first glance, all of these looked like mass-exploitation candidates: popular products, critical bug classes, no authentication. Earlier in the year we reviewed all the publicly disclosed details, but proof of concepts were only recently released. So, we did a deep dive on the technical details of each. We tested each public proof of concept in our virtual lab and ran through the exploitation prerequisites.

    Several patterns emerged: every bug requires a non-default configuration, and every public proof-of-concept is pinned to an exact build, which prevents easy mass-exploitation. They split three ways on memory address randomization (ASLR) to achieve code execution: one needs it disabled outright, two defeat it on their own with a runtime leak, and the last defeats it only when the configuration reflects the leak back, a setup that seemed intentionally built for exploitation.

    • nginx-rift (CVE-2026-42945)
      The bug: a rewrite-module heap overflow autonomously flagged by depthfirst.
      Lab result: We reproduced full RCE (commands as the Nginx worker), but only with ASLR disabled, since the exploit hardcodes heap and libc addresses; the groom took about ten tries.
      Real-world impact: Successful exploitation requires a rewrite with ? in its replacement followed by a set on a regex capture, which is an uncommon non-default configuration. On a host with ASLR enabled (the default in modern Operating Systems) it is a worker crash, not code execution.
    • nginx-poolslip (CVE-2026-9256)
      The bug: a heap overflow in the same rewrite module, credited by F5 to Mufeed VH (Winfunc Research), Nebula Security, and Vexera AI.
      Lab result: Nebula's standalone PoC leaks the heap and Nginx base at runtime and needs no second bug; we reproduced command execution as the worker (uid=101(nginx)) on the stock nginx:1.31.0-trixie image, and an earlier poolslip and rift chain did the same on nginx:1.30.0. Both are probabilistic and build-pinned.
      Real-world impact: Exploitation requires a rewrite with overlapping captures (^/((.*))$ → $1$2), which is another uncommon non-default configuration. Unlike rift, its runtime leak defeats ASLR on its own, so on a stock host with ASLR enabled it still reaches code execution, not just a crash.
    • nginx-quicburst (CVE-2026-42530)
      The bug: a use-after-free in the HTTP/3 QPACK module, credited by F5 to six finders in coordinated disclosure.
      Lab result: Reproduced by compiling nginx 1.31.1 with HTTP/3 and AddressSanitizer pinpointed the use-after-free in ngx_http_v3_get_insert_buffer, an unauthenticated worker crash. Nebula's weaponized reverse-shell PoC defeats ASLR with a remote memory scan, so a working RCE now exists, but across three lab runs on the stock nginx:1.31.1-trixie image, the scan never successfully obtained a shell.
      Real-world impact: Exploitation requires HTTP/3 explicitly enabled (listen ... quic), which is non-default, and affects only 1.31.0 and 1.31.1. A weaponized RCE PoC exists, but it never produced a shell in our lab. So for a typical deployment, it likely results in an unauthenticated worker crash.
    • CVE-2026-42533
      The bug: a heap overflow in the stream script engine (the same length-versus-copy family), independently reported by roughly eighteen finders; depthfirst published a public PoC in July.
      Lab result: We reproduced the leak and full RCE via a crafted TLS SNI, but only when the config reflects the map result back to the client (return); a routing config (proxy_pass) still crashes the worker yet leaks nothing.
      Real-world impact: Exploitation requires Nginx built with the non-default stream and ssl_preread modules plus a map on an unnamed regex capture. The ASLR bypass needs a reflecting config, so a typical SNI-routing deployment is just a DoS.

    CVE

    Non-default config requirement

    ASLR requirement

    Our lab result

    Public PoC

    rift

    Uncommon URL rewrite config

    Must be disabled (hardcoded addresses)

    RCE confirmed (ASLR off, ~10 tries)

    depthfirst

    poolslip

    Uncommon URL rewrite config

    Defeated by runtime leak

    Standalone RCE on stock 1.31.0-trixie

    Nebula, y198nt chain

    quicburst

    HTTP/3 enabled

    Defeated by remote scan

    UAF reproduced; RCE unsuccessful in VM lab

    Nebula

    CVE-2026-42533

    Uncommon TLS SNI-routing config

    Defeated only if config reflects leak

    RCE (reflecting config); DoS (routing config)

    depthfirst

    Rift and Poolslip: How Common Are These Rewrite-Rule Configs, Really?

    Rift and poolslip both depend on the same kind of non-default rewrite rule, so the one thing that decides their real-world impact is how often that rule actually appears in the wild.

    Calif, a security research firm working on AI-assisted vulnerability discovery and led by Thai Duong, came at the prevalence question from a different angle. Their "Needle in a Haystack" study parsed 35,633 real Nginx configurations with nginx's own tokenizer (their parser, ngxray) and found one genuinely vulnerable config for rift and zero for poolslip. Its conclusion:

    Both are real and exploitable, but their real-world impact is likely low. [...] They rely on config patterns that almost never appear in production.

    Calif is candid about the limits. Since GitHub skews toward examples and small projects and the gnarliest rewrite chains live in private repos, so the real number is probably higher. But the distance between calling these configs common and measuring them at one in 35,633 is enormous.

    Figure 3: Calif (@calif_io) summarizing "Needle in a Haystack": out of 35,633 Nginx configs scanned with their open-source ngxray, one vulnerable config, in an abandoned project. Source: x.com/calif_io.
    Figure 3: Calif (@calif_io) summarizing "Needle in a Haystack": out of 35,633 Nginx configs scanned with their open-source ngxray, one vulnerable config, in an abandoned project. Source: x.com/calif_io.

    Assetnote, an attack surface management firm, got there independently. Co-founder Shubham Shah put it plainly:

    The art of publishing research is to educate, inform, and help others. When the recent wave of Nginx vulnerabilities dropped, our team spent time understanding the real impact and so did Thai Duong's team. We came to the exact same conclusions. Meanwhile, a thousand screaming articles about how important the vulnerability is, when in reality, it just isn't.

    The common thread is that two independent teams, Calif with a parsed count and Assetnote through its own review, reached the same conclusion: the configuration rift and poolslip depend on is too rare in the wild to make them mass-exploitation candidates.

    Quicburst: Determining Vulnerable Nginx HTTP/3 Prevalence

    For quicburst, we ran the numbers ourselves because its impact depends entirely on one opt-in feature: HTTP/3. So, we set out to determine of all the Nginx we can see, what share actually runs HTTP/3, and of those which fall in the vulnerable version range. We measured it three ways:

    • What code owners write (parsed GitHub configs). We collected 38,012 public Nginx config files and parsed 30,483 of them from 31,829 repositories with ngxray, Calif's scanner, which uses Nginx's own tokenizer rather than keyword matching. 133 defined a listen ... quic listener, the directive that actually turns HTTP/3 on, or about 0.44% of the configs we parsed.
    • What's exposed on the internet (Shodan). Of the 45 million Nginx hosts, Shodan indexes (product:nginx), about 147,000 advertise HTTP/3 via the Alt-Svc: h3 header. That is roughly 0.3%, and the header overcounts: it is high-recall (~99%) but only ~81% precise, since roughly a fifth are fronting stacks flying an Nginx banner. Of the genuine HTTP/3 hosts, about 66% mask their version (server_tokens off), and among those that reveal one, only about 11% run an affected build (1.31.0 or 1.31.1). That works out to roughly 5,600 Nginx-h3 hosts internet-wide reporting a vulnerable version: about 3.8% of all HTTP/3 advertisers, and roughly 0.012% of all Nginx Shodan indexes.
    • What's actually vulnerable (active attack-surface scan). Across a sample of 56,000 identified Nginx hosts, about 2,000 looked like HTTP/3 candidates, but only about 899 completed a real HTTP/3 handshake, or 1.6% of the Nginx in scope. The rest merely answered a QUIC Version-Negotiation packet, which any QUIC-ish listener does without speaking h3. Filtering for Nginx genuinely terminating that h3 (not a fronting CDN), then for an affected version, left just 4 hosts on 1.31.0 or 1.31.1, about 0.007% of the 56,000-host sample.
    Figure 4-6: Shodan narrows from all Nginx (44,715,790), to those advertising HTTP/3 (145,395), to those reporting an affected Nginx version running HTTP/3 (5,639). Source: shodan.io.
    Figure 4-6: Shodan narrows from all Nginx (44,715,790), to those advertising HTTP/3 (145,395), to those reporting an affected Nginx version running HTTP/3 (5,639). Source: shodan.io.

    CVE-2026-42533: DoS in Practice, RCE Only by Design

    Since RCE vs. DoS depends on reading memory required for the ASLR-bypass, we set out to confirm whether the leaked memory ever reaches the attacker in standard out-of-the-box configurations. Based on our review of the Nginx documentation, the normal configuration for an ssl_preread routing decision is proxy_pass to a backend, which consumes the value internally and reflects nothing to the client. All three documented Nginx examples extract the SNI "without terminating SSL/TLS", route it with proxy_pass, and fail to answer the client with return. The ASLR-bypass required for RCE needs a vastly different configuration that echoes the leaked libc and heap pointers back to the client. Based on our review, depthfirst's config appears to be an intentionally vulnerable test harness for demonstration, as it pairs ssl_preread with a return from a separate module (ngx_stream_return_module) in a way no documented example does, and that return "$1$m" answers a TLS listener with plaintext no real client could read.

    This results in two very different outcomes: a normal routing config will only ever result in a denial of service, but the ASLR bypass and remote code execution requires depthfirst's intentionally vulnerable proof-of-concept config. Our side-by-side on the same vulnerable 1.30.3 build demonstrates this difference:

    Config Setup

    Directive

    Memory leak readable?

    Outcome

    Reflecting (depthfirst PoC)

    return "$1$m"

    Yes, libc and heap addresses recoverable

    Worker crashes; leak enables a self-contained ASLR-bypass RCE

    Routing (normal)

    proxy_pass "$1$m"

    No, bytes go to the backend

    Worker crashes; DoS only

    Next, we measured how common the reflecting configuration is, parsing configs with ngxray rather than counting keywords. In the same corpus of 30,483 parsed configs, only 48 enabled ssl_preread at all, about one in 600, and not one reflected the SNI back to the client.

    To find the reflecting wiring, we had to hunt for the feature directly, which surfaced 541 configs that enabled it across GitHub. Of those, 534 routed the SNI to a backend with proxy_pass, the feature's intended use; two reflected it back to the client, the setup the exploit needs; and five did neither, routing in an included file. We read both reflecting configs, and neither was a deployment: one was depthfirst's own proof-of-concept, the other a personal debug config from 2023 that echoed the parsed SNI back only to print it while testing the feature, years before the vulnerability was known.

    Across all 30,483 parsed configs, we found zero reflecting configurations, bounding the wiring the RCE requires at fewer than one in 10,000.

    Nginx ssl_preread configs on GitHub (541 parsed)

    Directive

    Count

    Vulnerable?

    Reflecting

    return

    2

    Yes

    Routing

    proxy_pass

    534

    No

    Other

    mixed

    5

    No

    When Marketing Hype Outpaces Impact

    In May, June, and July, various news sites reported these CVEs as the next big threat. The Hacker News framed nginx-rift as an "18-year-old unauthenticated NGINX RCE" and returned in July to cast CVE-2026-42533 as a "15-year-old" critical flaw that "can crash workers and may allow remote code execution," nginx-poolslip drew "patch now" urgency from security trade press (GBHackers, Cyber Security News), and nginx-quicburst arrived with Nebula Security's tweet billing their research as "only the third NGINX vulnerability since 2014 to receive NGINX's 'major' severity rating," complete with a lab RCE video and a promised writeup "including the ASLR bypass."

    Figure 7: The amplification cycle across the four CVEs: The Hacker News billing rift and CVE-2026-42533 as "18-" and "15-year-old" pre-auth RCEs, and Nebula Security and depthfirst announcing quicburst, poolslip, and rift with branded names and lab reverse-shell demos. Sources: x.com, thehackernews.com.
    Figure 7: The amplification cycle across the four CVEs: The Hacker News billing rift and CVE-2026-42533 as "18-" and "15-year-old" pre-auth RCEs, and Nebula Security and depthfirst announcing quicburst, poolslip, and rift with branded names and lab reverse-shell demos. Sources: x.com, thehackernews.com.

    As each of these released, our customer security teams reached out about their exposure, and understandably, the headlines described critical, unauthenticated RCEs in software they run everywhere. But our Threat Enablement team returned the same answer every time. They were likely not running the non-default rewrite rules, the opt-in HTTP/3 listener, or the stream ssl_preread config these bugs require, so their real exposure was effectively zero. Even if there happened to be a host running one of these configs, code execution was unlikely. Some exploits require ASLR to be disabled to avoid crashing the worker. Those that can bypass ASLR depend on specific conditions and exact builds, making broad scanning and exploitation impractical.\

    Figure 8: A sample of customer requests that followed each disclosure, echoing the headline framing ("critical NGINX vulnerability," "15-Year-Old Pre-Auth nginx RCE") and asking us to check their exposure.
    Figure 8: A sample of customer requests that followed each disclosure, echoing the headline framing ("critical NGINX vulnerability," "15-Year-Old Pre-Auth nginx RCE") and asking us to check their exposure.

    AI-Assisted Research: Surfacing Vulnerable Code Sinks in All Configurations, Not Just Common Defaults

    When leveraging AI to identify vulnerabilities in popular software, its real advantage is thoroughness. An agent can follow potentially vulnerable code paths far more exhaustively than human review usually does, including those paths that only surface under configurations that may be less common. Historically, human researchers have had finite time and spent it where impact usually lives, the commonly deployed and reachable code paths, rationally setting aside the obscure functionality exposed in an application's edge case configuration. An agent has no such limitation: it audits the rarely enabled path with the same attention as the common one, so it surfaces real bugs in code that human review never prioritized. That thoroughness is genuinely valuable and is why these tools turn up bugs that sat unnoticed for years. Unfortunately, the same coverage that finds more bugs also disproportionately finds them in code paths rarely enabled in real-world configurations.

    The effect compounds when many teams point LLMs at the same ubiquitous target (Nginx runs roughly a third of the web) and its known-fragile functionality (the rewrite and script engine, the new QUIC and QPACK code). They surface the same latent bugs at nearly the same time. F5 credits six finders for quicburst, three independent teams for poolslip, and roughly eighteen for CVE-2026-42533. The result is a stream of near-simultaneous advisories that reads like an Nginx apocalypse but is really the same tools finding the same bugs, amplified across news outlets and social media. None of it addresses the question that matters though: whether these CVEs affect the configurations commonly deployed across defenders' attack surfaces.

    Answering that question is where the process breaks down. Across the coverage we reviewed, every writeup named the configuration the bug requires, but none measured how common it is (aside from Calif’s rewrite-config analysis). Take quickburst: HTTP/3 is advertised by roughly 0.3% of the Nginx hosts Shodan indexes, and barely a tenth of those report an affected version. “Unauthenticated RCE in Nginx” and “unauthenticated RCE in the 0.3% of Nginx running HTTP/3” are not the same story. One makes headlines; the other lands in a patch backlog. Defenders escalated on the framing that reached them, which is why conveying how common a vulnerable configuration is should be part of publishing research, as the researchers have the lab and the context to measure it. Published alongside the writeup, that number lets everyone downstream prioritize accordingly.

    Conclusion: Triaging CVEs in the Age of AI Security Research

    As AI security research scales up vulnerability discovery, the burden of separating real risk from noise shifts downstream to defenders. The research overhead of identifying impactful threats among overhyped vulnerabilities is more than most security teams can manage, but the following is a good place to start:

    • Does it affect the default configuration? If the vulnerability requires a specific non-default configuration, affected software is a subset of deployments, not all of them. Find out how large that subset realistically is before escalation.
    • Is the exploitation impact conditional? A memory corruption vulnerability whose proof-of-concept requires ASLR to be disabled or configuration trickery for an ASLR bypass, will likely result in DoS instead of RCE. Understand the technical nuances of exploitation to ensure likely impact is properly understood.
    • Is it being exploited in the wild? KEV status and observed in-the-wild activity are a good indicator of what threat actors deem to be good candidates for mass exploitation.

    This is the work our Threat Enablement and Analysis team specializes in: disclosure triage, lab reproduction, and attack surface analysis. So, by the time the next marketed RCE hits the news, our customers already know whether they are impacted.

    If you’re interested in learning more about emerging threat services delivered through our Cosmos platform, visit bishopfox.com/services/continuous-threat-exposure-management/emerging-threat-services


    Nate Robb

    By Nate Robb

    Senior Operator

    Nate Robb is a Senior Operator on the Threat Enablement Team at Bishop Fox. Prior to coming to Bishop Fox, he held roles as a security consultant and spent time as a full-time bug bounty hunter, where he worked to secure Fortune 500 companies, state and Federal Agencies, and small and medium-sized businesses.


    Banksy Fox exploder1

    By Threat Enablement & Analysis Team

    The Bishop Fox Threat Enablement & Analysis team researches emerging vulnerabilities, exploits, and attacker techniques to understand how new threats translate into real-world risk. The team combines vulnerability research, exploit development, threat intelligence, and offensive security expertise to analyze new disclosures, validate exploitability, and develop methods for identifying affected systems at scale.

    Subscribe to our blog

    Be first to learn about latest tools, advisories, and findings.