Adversa's Copilot CLI attack ran in 50% of runs on one of three models, and the post gives no run count
Adversa AI says an encrypted web page made GitHub Copilot CLI send the contents of a .env.prod file to an attacker endpoint in 28 seconds. The post gives one rate, 50% on one model, with no run count, and GitHub says the user's own choices mean it is not a vulnerability.
By Parminder Kumar Sharma · · 11 min read

The rate is 50%. The run count is missing
Adversa AI's post of 6 October 2026 says one model offered in GitHub Copilot CLI "executed the full chain in 50% of our runs". It does not say how many runs that was: 50% could be 1 of 2 or 50 of 100. It reports on three models: one that ran the chain, and two GPT-5.6 models, not named beyond that, that refused. GitHub's documentation, read at 14:46 BST on 6 October, lists 15 models for Auto selection in the CLI. Adversa reports on 3 of them (20%, derived) and is silent on the other 12.
That is a security vendor's result, from a test set-up the post does not describe. It does not establish that the technique works against real users, that any real secret has left any real machine, that it works outside autopilot, that it is new, or that GitHub has changed nothing since 1 October, the last date Adversa says the chain reproduced. The Register's report (14:00 UTC) adds a GitHub statement and no independent test.
What Adversa says it showed
The class is called Cryptographic Context Injection. In Adversa's words, "Static guardrails read text; they do not run it." The instruction travels as ciphertext, with the key material and an instruction to decrypt, so a classifier reading ingested text sees an opaque blob. The agent runs the decryption itself, and the plaintext arrives as the output of its own tool. In the Copilot case Adversa says the agent read local files on the way, then requested a follow-up URL carrying their contents, in 28 seconds. This briefing does not reproduce the ciphertext format, key handling or instruction wording; Adversa withholds its payloads too.
Two conditions are stated: autopilot mode and a permissive model. Adversa writes that "we are not claiming a confirmation prompt bypass", and says the same instructions as plaintext are refused, so the encryption is what gets them through.
Stated and not stated in Adversa's Copilot CLI post, read in full on 6 October 2026
- Question
- Conditions
- Stated in the post
- Autopilot, a permissive model, a user who asks for a page.
- Not stated
- Permission flags, sandbox state, CLI version.
- Question
- Models and rate
- Stated in the post
- mai-code-1.1-flash ran the full chain in 50% of runs; two GPT-5.6 models consistently refused.
- Not stated
- Run counts, which two GPT-5.6 models, the other 12 Auto models.
- Question
- What was taken
- Stated in the post
- .env.prod contents reached an attacker endpoint in 28 seconds.
- Not stated
- Test canaries or real credentials.
- Question
- Reproduction
- Stated in the post
- Reproduces as of 1 October 2026.
- Not stated
- Independent reproduction, any change since.
- Question
- Real-world use
- Stated in the post
- Payloads withheld; no victim described.
- Not stated
- Any exploitation outside the lab.
- Question
- Outside autopilot
- Stated in the post
- Autopilot named as a condition.
- Not stated
- Whether it runs with approvals on.
- Question
- Vendor response
- Stated in the post
- Validated by 1 October, declined, ineligible.
- Not stated
- Any fix, advisory or CVE.
| Question | Stated in the post | Not stated |
|---|---|---|
| Conditions | Autopilot, a permissive model, a user who asks for a page. | Permission flags, sandbox state, CLI version. |
| Models and rate | mai-code-1.1-flash ran the full chain in 50% of runs; two GPT-5.6 models consistently refused. | Run counts, which two GPT-5.6 models, the other 12 Auto models. |
| What was taken | .env.prod contents reached an attacker endpoint in 28 seconds. | Test canaries or real credentials. |
| Reproduction | Reproduces as of 1 October 2026. | Independent reproduction, any change since. |
| Real-world use | Payloads withheld; no victim described. | Any exploitation outside the lab. |
| Outside autopilot | Autopilot named as a condition. | Whether it runs with approvals on. |
| Vendor response | Validated by 1 October, declined, ineligible. | Any fix, advisory or CVE. |
GitHub's answer, and what GitHub's own pages say
Adversa reported the finding to GitHub's bug bounty programme on 17 September 2026. By 1 October, 14 days later, GitHub's triage team had validated it, declined to treat it as a vulnerability and ruled it ineligible for a bounty. Adversa quotes the reply: the user "explicitly asked Copilot CLI to fetch attacker-controlled content while giving copilot full permissions to act autonomously". Adversa published five days after that date. A GitHub spokesperson told The Register the finding "requires a user to intentionally direct Copilot CLI to fetch attacker-controlled or untrusted content".
GitHub's published bounty rules for Copilot explain the shape of that reply: prompt injection reports may be eligible only with a concrete impact, such as "Unauthorized actions occurring without required user confirmation". Adversa says it is not claiming a confirmation bypass, so by its own description the report sits on the far side of that line. The dispute is about where the line belongs.
GitHub's documentation, read between 14:40 and 14:50 BST, fills in the conditions. In autopilot, "Copilot cannot carry out any actions that require permission unless you explicitly grant it full permissions", and the prompt on entering it offers "Enable all permissions (recommended)". Full permissions equal --allow-all, which covers tools, paths and URLs, and autopilot is sticky by default: it stays on for the next prompt. By default all URLs need approval. Local sandboxing is "turned off by default" and experimental in the CLI; until it is on, shell commands run "with the same access as your user account".
Read together, the conditions Adversa needs are each a choice: autopilot, full permissions and a permissive model. Only the last may be invisible, and that is disputed: GitHub's docs say the CLI shows "the model used for each response" in the terminal, while Adversa says the user "does not see" which model handled the session. This briefing did not run the product and cannot settle it.
GitHub does publish advisories when the approval step itself fails. At 14:43 BST on 6 October the github/copilot-cli repository listed two: CVE-2026-29783 (high, 6 March 2026), where crafted shell expansion made commands look read-only, and CVE-2026-45033 (medium, 6 May 2026), where a nested git repository could run commands. None exists for this chain. Inference: GitHub treats a bypass of the approval step as a vulnerability and its absence, by user choice, as configuration.
Two friendly names: guardrail and autopilot
A guardrail or classifier sounds like a control that inspects what an agent will do. The finding is that it inspects text before it is decoded, and the dangerous step is the decoding, which the agent performs with a tool it was given. OWASP's prompt injection entry lists input and output filtering as a mitigation, lists Base64 encoding as a way to evade filters, and says "it is unclear if there are fool-proof methods of prevention". The NCSC wrote in December 2025 that prompt injection "may never be totally mitigated in the way that SQL injection attacks can be", that it "cannot be fully mitigated with a product or appliance", and: "Beware any that claim they can 'stop' prompt injection".
Autopilot is a name for removing the human approval step, which is the control this whole class of attack depends on being absent. GitHub's own page compares it to handing a task to a colleague. OWASP lists "excessive autonomy" among the root causes of excessive agency, and human approval for high-risk actions among its prompt injection mitigations.
Approval is a weak control too, on one vendor's evidence. Anthropic's post of 7 August 2026 says Claude Code users approve 97% of permission prompts, and that in its own test 143 of 1,053 paid testers (13.6%) caught a dangerous command swapped into a prompt. That is a vendor study, not independent and not about Copilot. It argues for scoping credentials and isolating first.
Not new, and not one vendor's
Adversa published the same class against Grok on 20 August 2026, 47 days before the Copilot post and 78 days after reporting it to xAI on 3 June. That write-up gave no success rate; Adversa told The Hacker News it had tried 20 times since June with 40% success, against Grok 4.5 Fast. An Adversa researcher described the Gemini variant on a personal site on 11 March 2026, 209 days before the Copilot post, with 5 of 5 reproductions. A preprint from October 2025 reports a related filter bypass on four chat platforms. For Grok a denominator exists, in a news interview. For Copilot it does not yet.
PromptArmor reported a different class against Copilot CLI in February 2026: a command-parsing bypass that let a shell command run without approval. It says it filed the report on 25 February and GitHub closed it the next day as "a known issue that does not present a significant security risk". Whether the advisories above relate to it is not stated.
Both researchers have commercial interests. Adversa sells a coding agent security platform and says it stopped the chain in an instrumented agent. PromptArmor sells monitoring of vendor AI changes. Neither fact makes a finding wrong. Both are reasons to ask for run counts and settings.
What each vendor's documentation says about defaults
The Register wrote that autopilot is the default in other agentic coding tools, naming Anthropic's Claude. Each vendor's own pages, read on 6 October, say something narrower, and different from each other. This briefing states what they say and does not rank them: Adversa tested only Copilot CLI, and a documented default is not a measured outcome.
Default and no-approval modes as each vendor's documentation describes them, read on 6 October 2026
- Tool and vendor
- GitHub Copilot CLI
- Default, per its docs
- Asks before tools that change or run files; URLs need approval; local sandbox off, experimental.
- No per-action approval
- Autopilot. All permissions (prompt: recommended) lifts tool, path and URL checks; docs advise a sandbox.
- Tool and vendor
- Claude Code (Anthropic)
- Default, per its docs
- Terminal start mode is auto (recent versions), a classifier review; manual mode runs reads only; Bash sandbox off.
- No per-action approval
- bypassPermissions: everything runs; docs say isolated containers and VMs only.
- Tool and vendor
- Codex CLI (OpenAI)
- Default, per its docs
- Recommends Auto in version-controlled folders: workspace-write sandbox, approval on request; network off.
- No per-action approval
- Dangerous full access (yolo): no sandbox, no approvals, "not recommended".
- Tool and vendor
- Gemini CLI (Google)
- Default, per its docs
- Most write tools need confirmation; sandbox enabled by flag, variable or settings.
- No per-action approval
- yolo: all tools auto-approved, "use with extreme caution".
| Tool and vendor | Default, per its docs | No per-action approval |
|---|---|---|
| GitHub Copilot CLI | Asks before tools that change or run files; URLs need approval; local sandbox off, experimental. | Autopilot. All permissions (prompt: recommended) lifts tool, path and URL checks; docs advise a sandbox. |
| Claude Code (Anthropic) | Terminal start mode is auto (recent versions), a classifier review; manual mode runs reads only; Bash sandbox off. | bypassPermissions: everything runs; docs say isolated containers and VMs only. |
| Codex CLI (OpenAI) | Recommends Auto in version-controlled folders: workspace-write sandbox, approval on request; network off. | Dangerous full access (yolo): no sandbox, no approvals, "not recommended". |
| Gemini CLI (Google) | Most write tools need confirmation; sandbox enabled by flag, variable or settings. | yolo: all tools auto-approved, "use with extreme caution". |
Wording differs. GitHub's autopilot prompt labels all permissions "recommended", while the same page says to consider a sandbox first. Codex labels its equivalent not recommended, and Claude Code says to use its no-checks mode only in isolated environments. Claude Code's auto mode and Codex's Auto preset are different designs from autopilot.
What UK teams should do, in this order
UK developers and organisations run coding agents on laptops and CI runners that often hold cloud credentials, SSH keys and tokens. The NCSC's May 2026 agentic AI blog says to apply least privilege, avoid long-lived credentials and monitor, and that "If you cannot understand, monitor or contain an agent's actions, it is not ready for deployment". Its December 2025 blog says to prefer "deterministic (non-LLM) safeguards that constrain the actions of the system".
Take this with you
Actions, in the order worth doing. Items marked Judgement are this briefing's view, not a source's.
- Do not run an agent in autopilot or allow-all mode on a machine that holds long-lived secrets. GitHub says such a session has the same access to files and shell as you do. Check scripts and CI jobs too: its autopilot page shows scripted use with the yolo option.
- Take long-lived secrets out of reach: short-lived, least-privilege tokens, and no cloud keys, SSH keys or production .env files in any directory the agent can read.
- Run unattended agents in a container, VM or dedicated machine that does not mount the home directory, with an egress allow-list. GitHub's local sandbox is experimental and off by default.
- Keep URL approval on and narrow: pre-approve named domains, deny the rest, never use allow-all-URLs. GitHub says URL detection for shell commands has limits, so enforce egress at the network too.
- Require approval for shell and network actions wherever credentials are visible. On Copilot Business or Enterprise, ask your administrator about the managed settings that disable bypass mode and require the sandbox.
- Judgement: do not treat model choice as a control. Adversa's model result is unverified and covers 3 of 15 Auto models.
- Monitor outbound connections from developer machines and CI runners, and alert on a new destination soon after an agent fetched a web page. Log agent tool calls with resolved arguments.
- Judgement: where unattended sessions fetched untrusted pages while secrets were readable, decide whether to rotate them. No exploitation is reported, so this is proportionality, not an incident trigger.
Earlier briefings cover the same pattern. Agentforce could post to Slack with no confirmation, and both were the defaults. OpenAI's worm-style prompt injection report gave four transcripts and no rate. Apple will tighten Full Disk Access for AI agents.
The question this leaves
Adversa has given one number without a denominator. GitHub has given one reason without publishing a test. Neither has published the settings. So the question for any team is not whether a 50% is right. It is this: on which of your machines can an agent read a secret, run code and reach the internet in the same session, and who signed that combination off?
Key facts
Sources
- PrimaryThe Copilot CLI write-up of 6 October 2026: the claim, the models, the rate, the 28 seconds and the disclosure timeline. Read in full.Adversa AIaccessed 2026-10-06
- PrimaryThe earlier write-up of 20 August 2026 that named the technique against Grok and Gemini, with its disclosure dates. Read in full.Adversa AIaccessed 2026-10-06
- PrimaryThe 11 March 2026 write-up of the Gemini variant, 5 of 5 reproductions.An Adversa AI researcher (personal research site)accessed 2026-10-06
- PrimaryPreprint on bypassing prompt guards, abstract: a stated theoretical barrier for filters and attacks on four chat platforms.arXivaccessed 2026-10-06
- PrimaryAutopilot mode: what it does, the permissions prompt, sticky behaviour, sandbox advice.GitHub Docsaccessed 2026-10-06
- PrimaryAbout Copilot CLI: security considerations, allowed tools, automatic approval, risk mitigation.GitHub Docsaccessed 2026-10-06
- PrimaryConfiguring Copilot CLI: trusted directories, path and URL permissions and their stated limits.GitHub Docsaccessed 2026-10-06
- PrimaryCloud and local sandboxes: local sandboxing off by default, experimental in the CLI, lighter-weight isolation.GitHub Docsaccessed 2026-10-06
- PrimaryEnterprise managed settings: disabling bypass mode and requiring the sandbox.GitHub Docsaccessed 2026-10-06
- PrimarySupported models: 30 listed, 15 in Auto model selection for the CLI. Counted from the page.GitHub Docsaccessed 2026-10-06
- PrimaryAuto model selection: what it does and the statement that the CLI shows the model used for each response.GitHub Docsaccessed 2026-10-06
- PrimaryIneligible submissions: the prompt injection rule and its conditions for eligibility.GitHub Bug Bountyaccessed 2026-10-06
- PrimaryPublished security advisories for the Copilot CLI repository, listed through the API at 14:43 BST on 6 October 2026.GitHubaccessed 2026-10-06
- PrimaryThe earlier Copilot CLI report of February 2026, a different class, with GitHub's reply. No payload reproduced.PromptArmoraccessed 2026-10-06
- PrimaryLLM01:2025 Prompt Injection: mitigations, including filtering and human approval, and the encoding scenario.OWASPaccessed 2026-10-06
- PrimaryLLM06:2025 Excessive Agency: excessive autonomy listed as a root cause.OWASPaccessed 2026-10-06
- PrimaryBlog of 8 December 2025 on whether prompt injection can be fully mitigated and what to do instead.NCSCaccessed 2026-10-06
- PrimaryBlog of 15 May 2026: least privilege, avoiding long-lived credentials, monitoring, human accountability.NCSCaccessed 2026-10-06
- PrimaryClaude Code auto mode announcement, 7 August 2026: its default and its own approval study. A vendor's own claims.Anthropicaccessed 2026-10-06
- PrimaryClaude Code permission modes: the built-in starting mode and bypassPermissions advice.Anthropicaccessed 2026-10-06
- PrimaryClaude Code sandboxing: stated to be off by default.Anthropicaccessed 2026-10-06
- PrimaryCodex approvals and security: the default sandbox, network off, Auto preset, dangerous full access.OpenAIaccessed 2026-10-06
- PrimaryGemini CLI policy engine: default and yolo approval modes.Googleaccessed 2026-10-06
- PrimaryGemini CLI sandboxing: how it is enabled.Googleaccessed 2026-10-06
- Reported byNews report of 6 October 2026 at 14:00 UTC, with the GitHub spokesperson statement. Used as a pointer and for that statement.The Registeraccessed 2026-10-06
- Reported byReport of 20 August 2026 carrying Adversa's answers on the Grok build tested, 20 attempts and 40% success.The Hacker Newsaccessed 2026-10-06


