GitHub Security Lab published a breakdown of an audit conducted by its open-source AI agent, Taskflow Agent: the tool found and confirmed 24 vulnerabilities in Android apps, and advisories for all findings are published at securitylab.github.com/ai-agents. The agent from the seclab-taskflows repository is run via GitHub Copilot, and the result was CVE-level findings, including a location-tracking chain in the OsmAnd navigator with over 10 million installs.

image
image

What happened

A GitHub blog post titled “How we found 24 Android vulnerabilities using our open source AI security agent” describes how the GitHub Security Lab’s open-source AI agent, Taskflow Agent from the seclab-taskflows repository, run via GitHub Copilot, found and confirmed 24 vulnerabilities in Android apps. All advisories for the findings are collected in a separate index at securitylab.github.com/ai-agents. Two taskflow prompts were adapted for the platform: the new gather_mobile_entry_point_info.yaml separates an app’s entry points into mobile and non-mobile, while classify_application_local.yaml runs each entry point through a list of popular vulnerability classes, including confused deputy and insecure broadcasts. The most striking finding was in the OsmAnd navigator: the exported MapActivity accepts intent-extras such as SilentImport, Replace, and SettingsTypes without user confirmation, allowing a third-party app with zero permissions to open a chain that can track the device owner’s location.

Context

Taskflow Agent is an agentic extension of GitHub Security Lab in which an expert audit methodology is packaged into reusable taskflow prompts rather than free-form dialogue with the model: runs are performed in strict and broad modes, and checks are defined by explicit checklists of vulnerability classes. The list of classes is set manually due to LLM nondeterminism: without an explicit list, the model may miss a needed bug category, so each entry point is checked against specific items. The mechanics of the OsmAnd finding, however, rely on a long-standing Android feature: the platform does not allow an app to restrict which extras an external caller can place in an intent, so issues with settings read from extras have been known in the ecosystem long before AI agents and belong to classic exported components.

Why this matters for the industry

For the industry, this is a rare case where an agentic LLM system is backed by a measurable result rather than a demo: 24 CVE-level findings with public advisories, recorded runs, and a verifiable machine-readable result artifact. For AppSec teams, this is a ready-made template for first-line mobile code checks: instead of narrow expert audits, part of the exported surface is covered by a regular agentic run that can be integrated into CI. For teams building their own agents, this is a methodology engineering reference: reusable workflow prompts, entry-point breakdown, explicit checklists, and parallel strict and broad runs compensate for model nondeterminism without new architectures, and the system’s value shifts to the quality of checklists and fix tracking. If the advisory index continues to grow, the corpus of public findings is a natural benchmark for agentic security, and, following the Android adaptation, taskflow sets for other platforms and bug classes are likely; these are expectations, not source-confirmed data.

Why this matters for users

The practical effect for readers is available today. If you write Android apps, check exported components and especially settings read from intent-extras: as OsmAnd shows, such a scheme fits an attack that requires no permissions from the attacker. Your own repository can be run in a codespace with the command ./scripts/audit/run_mobile.sh in the myorg/myrepo format: according to the authors, this takes 1–2 hours for a medium-sized repository, the result is read from SQLite (audit_results table, has_vulnerability column), and a GitHub Copilot license is required. The run can be set as a regular CI step and an audit_results report can be collected for the team. Separately, it is worth checking your dependencies against the advisory index at securitylab.github.com/ai-agents for affected libraries.

What is still unknown / limitations

The main source is GitHub’s own blog post about its own tool, and the Hacker News thread received 2 points and 0 comments, meaning there is no independent discussion of the material yet. It is unknown in how many repositories these 24 vulnerabilities were found, so a comparison of the agent’s efficiency with manual audits remains unsupported. There is no open statistics on the false-positive rate and token cost per run in the sources, and without this data it is hard to justify integrating runs into CI. Reproducibility is partial: the prompt layer and audit script are open, but full independent replication is limited. Expectations about transferring the methodology to other platforms and the advisory index’s role as a benchmark are interpretations, not confirmed facts.

Sources

Author

Look at AI, editorial team