P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

Malware that talks to AI triage: by Talos's own bar, only one of four tricks steered the models, the plainest

Cisco Talos counts 84 samples in four families that carry text addressed to AI analysers. In its test of five unnamed local models, only the plainest technique steered verdicts by Talos's own bar, at 264 pairs toward benign against 33 toward malicious.

By Parminder Kumar Sharma · · 21 min read

A dark room with a graphite laptop on a walnut desk at the right. Its screen shows a generic two-panel window: on the left a tall stack of plain grey bars like the lines of a program, the top two inside an amber strip; on the right a narrow panel of three empty message capsules with small cyan pills. No writing appears anywhere, and the left of the picture is empty dark.

Talos's own chart: the plainest trick moved verdicts toward benign 264 times and toward malicious 33

Cisco Talos published a post on 8 October 2026 about malware that carries text addressed to an AI reviewer. It counts 84 samples in four families, collected from January 2025 to July 2026, and says the best techniques steered the outcome the attacker's way in "about 35% of test runs". The figure that makes that sentence checkable is in the post's Figure 4. For the plainest technique, a comment telling the model to ignore the file, the chart's label reads 264 to 33. In 264 matched pairs the model's verdict moved toward benign when the text was present, and in 33 it moved toward malicious. That is 8 to 1 and a net 231 pairs (both derived), and the chart prints the net rate as +34.4 percentage points, which we take to be where the post's 35% comes from. The label's meaning is our reading, because the post does not define the notation. The printed p values match exact two-sided sign tests on counts read that way: 0.489, 0.500 and 0.063 to the printed decimals, and the rest below 0.0001.

The other three techniques did not repeat it. Template spraying, which took the most engineering, moved verdicts by minus 3.1 points (34 pairs toward benign, 41 toward malicious, p = 0.489), so the chart shows no steering. Context fabrication reached plus 11.6 but stayed under the chart's noise line. Guardrail triggering went the other way, minus 28.5, with 87 pairs toward malicious against 13 toward benign. One of the four cleared both of Talos's evidence bars in the attacker's favour, and it was the cheapest. Here is what that does not establish:

  • That 35% of runs ended benign. The rate is a net figure on a three-point scale of benign, suspicious and malicious, and any move toward benign counts. In Figure 5, read from the bar lengths (our reading, to the nearest run), the three strings that ask the model to ignore the file gave a benign verdict in 41, 40 and 48 of about 135 runs. The rest stayed suspicious or malicious.
  • A rate in your SOC. The test used five local models that the post does not name, three runs each, with strings inserted into files they never shipped in. Hosted models, other prompts and real pipelines were not tested.
  • A count of attacks. The 84 counts samples collected through CAIRN, which finds them with Talos's own search filters over VirusTotal metadata. The post does not say whether it counts family members or files that carry the text, and the family reports show builds arriving in short bursts.
  • A missed detection. The comment does nothing when the file runs, and Talos says core conventional detection is unaffected. A shifted verdict in a test is not an alert that was closed.
  • A contest between AIs. What the technique exploits is older than that: a program's text is read by a tool that cannot reliably separate the file's contents from the instruction it was given.

What the 84 counts, and where the samples came from

Talos files this behaviour as archetype A3, "AI-Analysis Evasion", in CAIRN, the Cognitive Artifact Intelligence Research Network. CAIRN is Talos's open-source toolkit, launched on 22 September 2026 (the same launch carried CLOSEDQUORUM, which we covered in our briefing). It works only from VirusTotal metadata, with no binary downloads or detonation, and Talos calls it "a research effort, not a pure active threat signal". Its README says every family attribution it publishes is first confirmed by hands-on reverse engineering. A3 is one of 12 archetypes, A0 to A11. The post defines it as malware that "embeds natural-language instructions to influence automated analysis".

Where the samples came from matters for what 84 can mean. CAIRN finds candidates with search filters over VirusTotal metadata: 27 in the README, of which 24 are enabled. The general filter Talos links for this class, in the version it links, searches for text of this kind in Windows program files that at least five antivirus engines already flag. So the corpus is built by searching for the feature it then counts, and a file that carries differently worded text, or that no engine flagged, is not in it. Other channels cover scripts and hunt for named families. PLOTSAFE shows why wording matters: Talos's own report says its text is generated per build, with a random token, which "defeats full-sentence detection by construction".

Table 1. The four families in the post, as each Talos family report in its repository describes them. The post itself gives no per-family counts.

  1. Family and technique
    FRUITSHELL: PowerShell reverse shell, plain comment
    Samples and dates in the family report
    1 original script, first seen 28 January 2025, plus 9 further scripts from independent adopters, to April 2026
    What the report leaves open
    Rates one of the nine as likely genuine coursework and traces the comment to a public post, then course material. No count of intrusions. Google lists the family as "Observed in operations"
  2. Family and technique
    PLOTSAFE: stealer in Go DLLs, comment generated per build
    Samples and dates in the family report
    Text says 45 DLLs and 2 EXE precursors; its tables list 22 earlier and 21 later DLLs. First seen 20 to 27 March 2026. The report says the text is in the later series only, and 9 of its 21 rows are marked as carriers
    What the report leaves open
    One carrier confirmed by static analysis, the other eight from content snippets taken at collection time. How the DLLs were delivered is unknown
  3. Family and technique
    HOLLOWCLAD: licence-cracker theme, one instruction in several chat formats
    Samples and dates in the family report
    7 samples, 11 to 29 April 2026, one since deleted from VirusTotal. Campaign closed
    What the report leaves open
    Efficacy against real pipelines is "unmeasured". Distribution channel not recovered
  4. Family and technique
    MANTLEMAZE: packed loader, invented institutional authority
    Samples and dates in the family report
    16 samples listed, 4 April to 11 July 2026 (the report says a 98-day campaign), text identical in the first and last build
    What the report leaves open
    Efficacy "unmeasured". The loader's final stage is not recovered

Adding the headline counts gives between 69 and 80, depending on whether FRUITSHELL is 1 or 10 and PLOTSAFE 45 or 47. The post says 84 and publishes no breakdown. We could not reproduce 84, which may only mean the database holds samples the reports do not list. The class is also wider than four families. Talos's repository carries a fifth A3 family, GUARDBREAKER, with 48 or more samples from March to August 2026, and the post notes the same comment inside ROZESHELL tooling.

The span is 19 calendar months (derived; the post speaks of 18 months). Inside it, the reports put most samples in short windows in 2026: PLOTSAFE's DLLs first seen between 20 and 27 March, HOLLOWCLAD's 7 between 11 and 29 April, MANTLEMAZE's 16 between 4 April and 11 July. Those are builds from a handful of operators. The only family with a long record is FRUITSHELL, with 10 scripts over about 15 months, Talos's figure. A sample count is not a count of actors, campaigns, intrusions or victims. For FRUITSHELL, Google's report of 5 November 2025 lists it in one table row as publicly available and "Observed in operations", and gives no number.

What the technique is, in plain terms

An AI-assisted triage pipeline does three things. It pulls text out of a file (strings, comments, script content), puts that text in a prompt with the analyst's question, and asks a model for a verdict or an explanation. Talos's point is that the model receives "both the analyst's question and the file's contents" in one input. A3 is text planted in the file for that reader. It does nothing when the file runs. Its only job is to be read. We describe the steps without the wording, because the wording is not what a defender needs.

  • Copied. The first case, FRUITSHELL, is a small reverse shell with a two-line comment addressed to an AI reviewer, saying the script is harmless and need not be analysed and describing a task that no code in the script performs. Talos saw the comment verbatim in nine further scripts from at least four actors over 15 months, and inside more capable tooling (ROZESHELL, with an AMSI bypass, runtime compilation and a shellcode loader). Talos reads that as a "cost-free addition regardless of what else their tooling does".
  • Generated. PLOTSAFE builds the sentence from a template, with different filler each time and a random build token. It also carries a 29-byte function that exists only so the Go compiler does not delete the unused text. Talos reads that as deliberate engineering.
  • Sprayed. HOLLOWCLAD repeats one refusal instruction wrapped in several chat-template formats, betting that one matches the wrapper a scanner uses. The post says seven formats; its Figure 2 lists nine rows, and the post does not reconcile them. The binary also carries about 80 sections named after well-known protectors although it is not packed, and notes that threaten a reader with monitoring and revocation. The family report says none of the threats describes anything the binary does.
  • Dressed in authority. MANTLEMAZE presents itself as the property of three well-known companies at once, backed by an invented contract, certification, NDAs and patents, to touch a topic a model's guardrails prohibit. Talos calls that casting "a very wide net". Underneath is a loader whose build path names a vulnerable Intel network driver (CVE-2015-2291, published in 2017 in NVD). That, not the text, is where the endpoint risk sits (our reading), and the family report advises blocking the driver.

A fifth variant is not in the post's four. ESET reported in late August and September a VBScript from UAC-0099, which it calls Russia-aligned, whose only AI-directed content is a comment asking for help building a weapon, meant to trip a scanner's safety refusal. Talos's figures have a technique called Guardrail Triggering and a row labelled GUARDBREAKER-WMD, which we read as this family's text. That is an inference from the labels, because the post's own text never mentions it.

What the 35% measured

The post's evaluation paragraph is the whole method. Talos took the anti-analysis strings from seven families, including all in the post, and crossed every string into every sample body, so each string was tested in files it never shipped in. A panel of five local models triaged each combination three times, giving 135 matched pairs per string. Each pair compares one model's verdict on the same file with the string present and with it removed, on a three-point scale. The reported percentage is a net rate: pairs that moved toward benign, minus pairs that moved toward malicious, divided by total pairs.

Three things follow. First, 135 is five models times three runs times nine files (derived), so the post's own numbers imply nine sample bodies, a small set whose selection is not described. Second, on the post's wording, a verdict that slips from malicious to suspicious counts as a move toward benign. Third, the text and the chart differ. The text says "about 35% of test runs" and that the plainest technique "worked almost universally". The chart prints a net rate of plus 34.4 points for it. Talos prints no denominator for any bar. Working back from the label, 231 net pairs at 34.4% implies about 670 pairs (derived, approximate).

Table 2. Talos's Figure 4, as printed. The pair counts are our reading of the chart's labels, which the post does not define. A hollow bar is one that cleared only one of Talos's two evidence bars.

  1. Technique in Figure 4
    A3.001 Direct instruction, the plain comment
    Pairs toward benign against toward malicious
    264 against 33 (8 to 1, derived)
    Net rate and the bar
    Plus 34.4 points. Solid bar, cleared both
  2. Technique in Figure 4
    A3.002 Template spraying
    Pairs toward benign against toward malicious
    34 against 41
    Net rate and the bar
    Minus 3.1 points, p = 0.489: no steering shown. Refusals: 34 pairs moved to refusal and none away, plus 12.8, hollow
  3. Technique in Figure 4
    A3.003 Context fabrication
    Pairs toward benign against toward malicious
    72 against 25
    Net rate and the bar
    Plus 11.6 points, p below 0.0001 but under the noise line. Hollow
  4. Technique in Figure 4
    A3.004 Guardrail triggering
    Pairs toward benign against toward malicious
    13 against 87
    Net rate and the bar
    Minus 28.5 points: the model became more suspicious. Refusals 5 against 0, p = 0.063
Twelve horizontal bars, one per string Talos tested, each about 135 runs, split into malicious, suspicious, benign, refused and unclear. The three strings that ask the model to ignore the file have 41, 40 and 48 benign verdicts. All nine other strings have 12 or fewer benign: 12, 7, 5, 5, 3, 1, 0, 0 and 0. Counts are read from the bar lengths of Talos Figure 5.
Talos's Figure 5 redrawn to scale. The counts are read from the bar lengths by us, to the nearest run; the post prints none.

Figure 5 is the more useful picture for a SOC. The three strings that ask the model to ignore the file are the only ones where benign verdicts reach about a third. No other string passes 12. HOLLOWCLAD's text drew refusals in about a quarter of runs (33 of about 130, derived from the bar), while GUARDBREAKER's drew 95 malicious and 40 suspicious verdicts in about 135 runs, none benign and, in this chart, no refusals.

Talos's own summary is that "the impact of these strings is a mixed bag", and its conclusion is that a well-constructed pipeline is "not meaningfully more vulnerable than a human analyst who knows what prompt injection looks like". The post reports no test of human analysts or of a pipeline built the way it recommends, so the second sentence is a judgement, not a result, and the HOLLOWCLAD and MANTLEMAZE family reports call real-world efficacy unmeasured. We have met the same shape before: OpenAI's self-replicating prompt injection report gave four transcripts and no rate, and Adversa's Copilot CLI attack gave 50% on one of three models with no run count. A rate without its denominator is a number you cannot carry to your own pipeline.

What the post states and what it does not

Table 3. The post's main claims against what its text, figures and linked reports establish. Dates are 2026 unless shown.

  1. The post says
    The best techniques steered the outcome in the attacker's favour in about 35% of test runs
    Stated or checked
    Figure 4: plus 34.4 points net for direct instruction, 264 against 33 pairs
    Not stated, or not established
    A share of runs ending benign. A denominator. That the other techniques worked: they scored minus 28.5 to plus 11.6
  2. The post says
    84 distinct samples, four families, January 2025 to July 2026
    Stated or checked
    The window is 19 calendar months. Family reports list 7 and 16 samples for two families
    Not stated, or not established
    A breakdown by family or year. A count the public files reproduce (69 to 80 from the reports). Whether it counts family members or files carrying the text
  3. The post says
    A panel of five local LLMs, three runs each
    Stated or checked
    Five models, three runs, 135 pairs per string
    Not stated, or not established
    Which models, versions, prompts or hosting. Any hosted model. How the nine files (derived) were chosen
  4. The post says
    The text must be plaintext, so is always detectable
    Stated or checked
    The model has to be able to read it; the family reports give detection guidance
    Not stated, or not established
    That any wording will match: PLOTSAFE generates its text per build. ESET also lists tricks that need no instruction, such as flooding the context, which the post does not discuss
  5. The post says
    FRUITSHELL was reported as active in the wild by Google
    Stated or checked
    Google, 5 November 2025: publicly available, status "Observed in operations"
    Not stated, or not established
    Any count of intrusions, actors or victims. The report traces the comment to course material and rates one of nine scripts as likely coursework
  6. The post says
    The comment moved into more capable tooling (ROZESHELL)
    Stated or checked
    ROZESHELL: AMSI bypass, runtime compilation, shellcode loader
    Not stated, or not established
    Which of its scripts carry the comment. The FRUITSHELL report (4 August) counts all three Zero-Loader scripts, the ROZESHELL report (3 September) says only one does
  7. The post says
    The text was recently attributed for the first time to a named APT group
    Stated or checked
    ESET, 10 September: a Russia-aligned group, UAC-0099, a VBScript, a target in Ukraine
    Not stated, or not established
    That it is the first: ESET cites earlier supply-chain cases with no named group. That the trick works. Talos links a news article, not ESET's report
  8. The post says
    A well-constructed AI-assisted pipeline is not meaningfully more vulnerable than a human analyst
    Stated or checked
    A judgement in the conclusions
    Not stated, or not established
    Any test of human analysts, or of a pipeline built to Talos's advice

The one named group: ESET's case, and Talos's test of it

The post says the technique class has "been recently attributed for the first time, to a named APT group", and links a news article. The primary is ESET's article of 10 September 2026, "GuardBreaker: Derailing AI-assisted malware analysis with a code comment". ESET says the Russia-aligned group UAC-0099 used a VBScript in the early stages of an attack on a target in Ukraine, with a decoy request about building a nuclear weapon in a comment, meant to trip an LLM-powered scanner's safety guardrails so that it stops inspecting the file. The script's job was to download MATCHBOIL, a loader ESET says the group uses exclusively. Help Net Security reported ESET's statement on X on 31 August. ESET's article describes one target and gives no sample count. The 48 or more samples come from Talos's repository report, not from ESET.

Three points follow. First, this is not the technique that worked in Talos's test. If Figure 4's guardrail technique is this string, 87 pairs moved toward malicious and 13 toward benign. Second, ESET's worry is a different failure: a refusal is no verdict, and a team that reads silence as clean has a blind spot. ESET's advice is that "a lack of output, too, needs to trigger further checks". Talos's panel of five local models drew few refusals for it (5 pairs to 0, p = 0.063, on the same reading), and whether hosted models behave the same is untested. Third, both publishers have an interest. ESET sells managed detection and response and recommends it at the end of the article. Talos is Cisco's threat intelligence group, and its post argues that AI-assisted pipelines can be built soundly. Neither interest makes a finding wrong. It is a reason to read the design and not the summary.

AI against AI is the wrong picture

"AI-analysis evasion" and "AI against AI" suggest two models in a contest. What is described is older. A program's text is read by a tool that cannot separate the file's contents from the instruction it was given, the data-versus-instruction confusion the NCSC discusses under prompt injection. The NCSC's blog of 8 December 2025 says current models "simply do not enforce a security boundary between instructions and data inside a prompt", and that prompt injection may never be totally mitigated the way SQL injection can be. It recommends reducing the likelihood and the impact instead.

Four labels deserve suspicion. Evasion: in Talos's test one technique in four steered verdicts, and in Talos's own reports the comment has no effect on the malware and conventional detection is unaffected. The ROZESHELL report assigns the archetype "because the comment demonstrates intent to defeat AI-based triage pipelines", not because the comment does any work in the attack. AI-assisted: assisted implies a person decides, and where a model can close an alert, it decides. Guardrails: the word suggests a safety control, but the NCSC's August 2026 blog says a model's built-in safeguards "should not be treated as holistic". Evidence, never instruction: that is Talos's design goal, not a property of any model. A prompt can ask the model to behave that way. Only code outside the model can enforce it.

A tall single-column diagram. A file under analysis is read by two things. Reader 1, conventional detection (signatures, YARA, sandbox), which Talos says is unaffected. Reader 2, the AI triage pipeline that A3 targets: extract text, put the analyst question and the text in one prompt, and the model returns a verdict, with no boundary inside the model. Control A labels the text as evidence. Control B, outside the model, lets a verdict add suspicion but never close, downgrade or approve.
Where the AI layer sits, as Talos describes the layers, and where the NCSC, ESET and Talos advice puts the controls. Controls A and B are advice, not test results.

For UK SOC and managed detection teams that have added model triage

In what we read, the NCSC has published no guidance written for language-model triage of malware samples. What the NCSC has published points one way.

  • 8 December 2025, "Prompt injection is not SQL injection (it may be worse)". Treat prompt injection as a residual risk to be reduced, not fixed. Design protections should "focus more on deterministic (non-LLM) safeguards that constrain the actions of the system". Be wary of blocking known phrases such as "ignore previous instructions", because there are endless ways to rephrase them. Log enough to spot suspicious activity, potentially "full input and output of the LLM" and tool use.
  • 21 September 2026, "One does not simply defend agentically". For detection work such as a SOC, the NCSC describes three principles, the first of which is "Technology is behind the detection, but humans respond with actions". It proposes a scale of how risky an automated defensive step is that starts at AI that provides data and then "explainable advice to a human", and it calls tasks that only advise a human the easy win.
  • 20 August 2026, "Managing the cyber risk of agentic AI". Interim advice, to be superseded by formal guidance: decide how much oversight each system has, evaluate any "judge" model independently, keep transcripts, and treat agent activity as user activity in 24/7 monitoring.

Cyber Essentials does not cover this. Version 3.3 of the requirements (April 2026) sets five technical controls: firewalls, secure configuration, security update management, user access control and malware protection. We searched its 26 pages for artificial, machine learning, language model, LLM, generative, AI and agentic, and found none. A certificate says nothing about whether a supplier passes the text of your files to a model, or what the model may decide. That is not a flaw in a baseline scheme, but it means the question has to be asked in writing. The NCSC's August blog lists ETSI EN 304 223 as further reading on baseline requirements for AI systems. We have not read the standard.

Take this with you

Put these to a managed detection supplier in writing

  • Does any language model read content from our endpoints, such as file strings, scripts, command lines or email text, as part of triage? Which models, hosted where, and who processes the text?
  • Can a model's output close, suppress, downgrade or auto-resolve an alert, or approve a file, without a person or a fixed rule agreeing?
  • How is untrusted text kept apart from the instructions you give the model, and what enforces that outside the model?
  • What happens when the model refuses or returns nothing? Does the alert escalate?
  • Which detections run independently of the model, and can a model verdict lower their result?
  • What is logged: the full input and output of the model, tool calls, and who accepted or overrode the verdict? Can we see it for our alerts, and for how long is it kept?
  • Have you tested the pipeline with text addressed to the reviewer, and how often is it repeated? What did it show?

What to do, in order

Take this with you

Defender actions for teams that use a model on untrusted text

  • Find every place a model reads untrusted text in your analysis pipeline. Sample strings are one. File names, command lines, email bodies, URLs, sandbox reports and ticket text can also be written by an attacker (our reading). List each, its owner, and whether its output can change a decision.
  • Separate evidence from instruction in the prompt: keep the task apart from the extracted text and label the text as untrusted. Treat this as a way to make injection harder, not a guarantee, and keep the model's tool access to read-only.
  • Make the model advise and a gate decide. A verdict may add suspicion or raise priority. It may never close, downgrade or approve an alert or a file alone. A deterministic detection or a person has the last word, and a refusal or empty output means a person looks.
  • Log the model's input and output for every verdict, with tool calls and who accepted or overrode it, in a place the text being analysed cannot alter.
  • Hunt for the text as a detection signal. In files and scripts, look for sentences that address an AI, an assistant or a language model, for chat-template markers, and for refusal or copyright claims, and raise the score when they appear. A hit is evidence of intent. A miss proves nothing, because wording can be generated per build, so do not treat the list as the protection.
  • Test your own pipeline safely. In a copy, not production, submit a harmless file that conventional detection is certain to flag, such as the standard antivirus test file, with and without a plain sentence addressed to the reviewer saying it needs no analysis. Record the model, the prompt, the runs and whether the verdict moves. Use no real malware and copy no strings from reports. Repeat when the model or prompt changes.
  • Write the position into supplier contracts and your own runbook: who decides, what is logged, what happens on refusal, how often the test is repeated, and who sees the result.

The question that exposes the gap

Talos and the NCSC agree on one point: the boundary between a file's text and your instructions cannot be left to the model. So ask your own pipeline, or your supplier. If a file told your triage model, in plain words, that it was harmless, which part of the process is built to disagree with the model, and how would you know it had?

Key facts

Sources

  1. PrimaryThe primary post, 8 October 2026 (feed time 10:00 UTC), read in full from the page HTML with a browser User-Agent from 11:20 BST: the A3 definition, 84 samples in four families, the evaluation paragraph, the recommendations and conclusions, and the figures, which were downloaded and read: Figure 4 labels and Figure 5 bar lengthsCisco Talosaccessed 2026-10-08
  2. PrimaryIntroducing CAIRN: Frontier tracking for AI-integrated malware, 22 September 2026, read in full: what CAIRN is, the metadata-first method, the acquisition filters, and the statement that it is a research effort and not a pure active threat signalCisco Talosaccessed 2026-10-08
  3. PrimaryCAIRN README, read as the raw file: metadata-only operation, the 12 archetypes A0 to A11, the 27 acquisition channels of which 3 are disabled, the ai-analysis-evasion channel, the safety boundaries and the reverse-engineering confirmation rule; the database is not in the public repositoryCisco Talos (GitHub)accessed 2026-10-08
  4. PrimaryAcquisition filters file at the commit the post links, line 105: the ai-analysis-evasion filter is limited to Windows program files with at least five antivirus detections; read, wording not reproducedCisco Talos (GitHub)accessed 2026-10-08
  5. PrimaryFRUITSHELL family report, last updated 4 August 2026: the original script, the nine further scripts and their actor groupings, the propagation path, and the detection guidance that model output must never override deterministic signalsCisco Talos (GitHub)accessed 2026-10-08
  6. PrimaryPLOTSAFE family report, last updated 6 August 2026: 45 DLLs and 2 EXE precursors in the text, 22 earlier and 21 later DLLs in the tables, the 9 marked carriers, the per-build generated text and the 29-byte functionCisco Talos (GitHub)accessed 2026-10-08
  7. PrimaryHOLLOWCLAD family report, last updated 4 August 2026: 7 samples between 11 and 29 April 2026, the chat-template spraying, the fake protector sections, the intimidation notes, and the open question on real-world efficacyCisco Talos (GitHub)accessed 2026-10-08
  8. PrimaryMANTLEMAZE family report, last updated 4 August 2026: 16 listed samples between 4 April and 11 July 2026, the fabricated authority props, the BYOVD loader finding and the open question on efficacyCisco Talos (GitHub)accessed 2026-10-08
  9. PrimaryROZESHELL family report, last updated 3 September 2026: the AMSI bypass and shellcode loader chain, and the statement that only one of the three Zero-Loader scripts carries the commentCisco Talos (GitHub)accessed 2026-10-08
  10. PrimaryGUARDBREAKER family report, last updated 3 September 2026: the guardrail-triggering technique, 48 or more samples from March to August 2026, the UAC-0099 attribution to ESET and CERT-UACisco Talos (GitHub)accessed 2026-10-08
  11. PrimaryGTIG AI Threat Tracker, 5 November 2025: Table 1 describes FRUITSHELL as a publicly available PowerShell reverse shell with hard-coded prompts meant to bypass LLM-powered security systems, status Observed in operationsGoogle Threat Intelligence Groupaccessed 2026-10-08
  12. PrimaryGuardBreaker: Derailing AI-assisted malware analysis with a code comment, 10 September 2026, read in full: the UAC-0099 VBScript, the decoy comment, MATCHBOIL, the earlier supply-chain cases, and the advice that no single model should decide a file is safe and that a lack of output should trigger further checksESETaccessed 2026-10-08
  13. PrimaryPrompt injection is not SQL injection (it may be worse), 8 December 2025, read in full: models enforce no boundary between instructions and data, residual risk, deterministic safeguards, deny-listing warning, loggingNational Cyber Security Centreaccessed 2026-10-08
  14. PrimaryOne does not simply defend agentically, 21 September 2026, read in full: the three principles for detection work, the potency scale and the low-potency easy winNational Cyber Security Centreaccessed 2026-10-08
  15. PrimaryManaging the cyber risk of agentic AI, 20 August 2026, read in full: interim advice, oversight levels, independent evaluation of judge models, transcripts, 24/7 monitoring, built-in safeguards not holistic, ETSI EN 304 223 as further readingNational Cyber Security Centreaccessed 2026-10-08
  16. PrimaryThinking carefully before adopting agentic AI, 15 May 2026, read in full: summary of the joint Careful adoption of agentic AI services guidance, start small and use agents for low-risk tasksNational Cyber Security Centreaccessed 2026-10-08
  17. PrimaryCyber Essentials overview, read in full: the five technical controlsNational Cyber Security Centreaccessed 2026-10-08
  18. PrimaryCyber Essentials requirements for IT infrastructure v3.3, April 2026, 26 pages read as text: five technical controls, and no occurrence of artificial, machine learning, language model, LLM, generative, AI or agenticNational Cyber Security Centreaccessed 2026-10-08
  19. PrimaryCVE-2015-2291, read through the NVD API: the Intel Ethernet diagnostics driver IQVW32.sys and IQVW64.sys before 1.3.1.0, published 2017-08-09NIST NVDaccessed 2026-10-08
  20. Reported byRussian hackers plant nuclear weapon prompt in malware to trip AI safety guardrails, 31 August 2026: used only for the date of ESET's statement on X, which was not read directlyHelp Net Securityaccessed 2026-10-08

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

How often

Every new briefing in one email, at 7am, or at 7am, 12:30pm and 6pm. Nothing is sent when nothing is new. Unsubscribe any time.