Greg Kroah-Hartman, maintainer of the Linux stable kernels, presented the talk “Security in the LLM age” at the Kernel Recipes 2026 conference, showing how much value AI tools provide when finding vulnerabilities in real system code. The example was an automated Mythos audit that claimed 79 issues in the kernel: after strict maintainer review, only about twenty were confirmed, and the final result of the whole batch was “10 real bug fixes.” For the industry, this is a rare closed-loop example where raw AI audit numbers can be compared with a human verdict determining what makes it into stable releases.

What happened
Greg Kroah-Hartman, a Linux Foundation Fellow and maintainer of the Linux stable kernels (as well as the USB and driver core subsystems), spoke on September 22, 2026, at the Kernel Recipes 2026 conference with the talk “Security in the LLM age” about how large language models are practically used to find vulnerabilities; the recording appeared on the conference’s official YouTube channel on September 29, 2026. The central example was a breakdown of a Mythos automated audit run on kernel code, and a slide fragment shared in the Hacker News discussion breaks down all 79 claimed vulnerabilities by category: 24 had no details (“something crashed”), 14 were not bugs at all, 3 contained completely fabricated data, and 15 had already been fixed in the latest release — 11 by other developers and 4 by Anthropic. Twenty items actually required fixes, mostly with caveats like “assume a malicious filesystem image” or “assume the possibility of packet injection in the middle of the network stack.” The maintainer’s final verdict on the whole batch was “10 real bug fixes.”
Context
Kroah-Hartman is responsible for accepting fixes into the stable kernel branches, so reviewing third-party security reports is his daily work, and his verdict is the human check that public claims about “AI-found vulnerabilities” usually lack. The value of the case is that it is closed-loop: the full denominator of claimed findings is known, and the result of strict triage is known, whereas vendor AI audit numbers are usually published before any review. Part of the flow has a mundane explanation: reports of findings already fixed by the time of review are a typical symptom of an outdated code snapshot on which the system ran, and a trivial check against git history before submission would have removed almost a fifth of the flow almost for free. Another part of the findings relies on extended threat-model assumptions: formally such reports are valid, but practically inapplicable, and it is exactly these that strict triage filters out first.
Why this matters for the industry
For the industry, the talk turns a long-standing maintainer complaint into a measured quantity: generating a list of findings has become cheap, the bottleneck has shifted to human triage, and the honest metric for a pipeline is not the number of findings but the share of accepted reports per expert hour spent. A product of the form “ran code through an LLM — got a list of vulnerabilities” does not work in this view: only a small part of the list reached the status of completed fixes, which is a direct indication that what should be sold to the buyer is verified contribution, not a raw counter. At the same time, AI companies are indeed moving upstream fixes: among the already-fixed findings on the slide is Anthropic, just the scale of such contribution is more modest than public claims about found vulnerabilities. A reasonable industry response looks like pre-validation before humans — mandatory reproduction steps, deduplication against git history and releases, specifying the model and code snapshot version — as well as an expected shift in metrics from “found N vulnerabilities” to “accepted upstream after review.”
Why this matters for users
For readers, the case provides a measured reason for skepticism: headlines like “AI found N vulnerabilities in the Linux kernel” should reasonably be read with the question “how many of these passed expert triage,” because in the public example the raw counter shrank by a factor of several even before the question of applied fixes. Those who use LLMs themselves to find vulnerabilities should add input filters to their pipeline before passing reports to humans: mandatory reproduction steps, checking against already-fixed versions, and an explicitly defined threat model, otherwise findings will end up in the same category of formal assumptions. Those who send reports to open-source projects should note the lesson about reputation: a flow of under-detailed and outdated reports consumes maintainer time and devalues subsequent reports from the same sender. The primary source for verifying any retellings remains the full talk recording on YouTube.
What is still unknown / limitations
All numbers refer to a single observation: one run of one system, Mythos, on one codebase, and the sample of 79 findings is too small to generalize the “claimed — confirmed” ratio to the whole class of AI audits. The run methodology — which model was used, with what prompts, whether there were duplicates, and on which code snapshot version the audit ran — is not known from available materials, so strict accuracy comparisons are premature. The slide with the findings breakdown is known from a fragment in the Hacker News discussion, and percentages should be rechecked against the full talk recording rather than retellings. The final estimate of the number of applied fixes is Kroah-Hartman’s own formulation after triage, not an independent verification. Expectations that vendors will switch to the “accepted upstream” metric and that common input filters for AI reports will appear are an interpretation of the talk, not a completed fact.
Sources
- Kernel Recipes 2026 - Security in the LLM age (talk recording, official Kernel Recipes YouTube channel)
- Hacker News: discussion of the talk recording with a slide fragment about 79 Mythos vulnerabilities
Author
Look at AI, editorial team
