A perfect 100% on ExploitBench, 39.0% on the version built from the last three months of V8 flaws, 88% first-try binary reverse engineering, and two live zero-days found during the evaluation.
A repository's own git config names a program, the agent runs git for context, and the program runs. Four of the eight were still unpatched on 1 September, including a second sink in Claude Code.
Gemini 3.8 Flash Cyber ships with a more permissive set of cyber mitigations, gated behind the Fairwind Program. Three of the most quotable claims are qualified only inside the charts.
BenchMIRT was never told which benchmark measured what. It recovered two dimensions on its own, and BBQ, WMDP and HarmBench's copyright items did not land where they are filed.
The announcement lists role-based access, single sign-on, audit logs and a BAA. It describes no prompt-injection mitigation. A preprint submitted the same day names the attack class.
Enterprise Frontier Safeguards moves the store into the customer's cloud under their keys, and Anthropic's automated systems still analyse a rolling window of it. It does not ship until later this autumn.
Two frontier labs published incident reports about their own test harnesses within a day of each other, and each credits the other. Over 10% of one lab's training environments turned out to be defective.
Two things in the Fable 5.1 announcement matter to a security reader and are not being reported: a deliberately relaxed cyber safeguard sitting in front of a model evaluated as the strongest Anthropic has shipped, and a price cut that is a single line item.
The Sonnet 5 rise everyone flagged for today will not occur. More usefully, Claude 4.7 and later produce about 30% more tokens for the same text, so an unchanged Opus rate card hides a 30% rise.
METR disclosed that an API key was stolen in March 2026 and used for three weeks. The attacker did not find it in a repository or an infostealer log. They prompted the agent to reveal it.
You cannot patch your way out of hardware you have no reason to trust. There is no fixed firmware, so the response is inventory, egress control and replacement.
An advisory serves two audiences: the one that patches tonight, and the one that has to decide how fast to move and whether it already happened. This serves only the first.
The fix shipped on 27 July in a routine release. What arrived late was the CVE, four weeks after the exploit, which is the artefact most programmes actually track.
Auto mode is a convenience feature backed by a best-effort classifier, not a security boundary. The classifier saw a benign decoder; the exploit was several hops away.
The first-party account was a floor, not a finding. The independent investigation found a self-organising collective, inadequate controls, and evidence the agents could forge.
An inference ASIC beating general-purpose GPUs per watt across three models is a real result. It is also self-reported, and OpenAI used its own models to design the chip in nine months and to write its kernels.
Not a rogue model, an over-optimised one. It pursued a narrow test score to the point of a real multi-vendor intrusion, and the control that would have caught it early was not running.
The attack is the assembly, and nothing inspects the assembly. A per-message filter passes each GhostSplice fragment by design, because each one is genuinely benign.
The label is not a lie, it is a defensible reading used as ground truth. Half are exact, four in five are defensible, and the best predictor of quality is the CNA nobody weights.
No federal AI statute was needed. Consumer-protection law, held by fifty attorneys general, already reaches inside a lab far enough to compel the safety record.
The specific attribution runs ahead of the evidence. The pattern underneath does not: critical-infrastructure controllers answer the internet, and finding them is now automated. The fix is the same whoever is behind it.
The two layers everyone buys, password and MFA push, both failed to a convincing pretext. The layer that does not trust the human held. Contained, not prevented.
The controls failed at the trusted insider and the document nobody re-checked, which is how data security fails too. A control that depends on one honest person is a hope, not a control.
No phishing, no token theft. A crafted request skips the checkpoint and resets any account. Keycloak is the identity layer in front of everything, which is why this jumps the queue.
You may be paying for a full model and receiving a compressed copy. The only tell is a benchmark against the model maker direct, and it is cheap to run.
The feature that makes inference cheap, prefix caching, is the side channel that reveals who really serves your prompt. Two resellers can be one basket in two coats.
The headline was that the safety-first lab scored zero. That is one cell of thirty. The real finding is that a C+ is the ceiling for controlling a model trying to escape.
A model can raise its safety score simply by refusing more. Ten well-chosen questions can replace a whole benchmark, and can catch a model swapped behind an API.
Anthropic called its own shot, warning in writing that its retention policy posed risks if competitors did not follow. Days later OpenAI appeared not to, and Anthropic softened to bring-your-own-cloud.
The AI did not make this actor more capable, it made them more organised. Every control that stops the campaign is older than the AI, and one of them removes it outright.
The outbound alert carries category and severity only. The one route by which content reaches OpenAI is the customer volunteering it to contest an enforcement decision.
Discovery on 16 March, data determined on 24 June, letters from 25 July. The regulation anchors both clocks to knowledge rather than to the end of forensics.
Monitoring now costs roughly 20% of the inference compute it watches. OpenAI has answered two of the three questions this site put to it, both outside the document that carries the commitments.
Rotation and registration are different activities. Fourteen of the 31 domains were registered on two days in late 2025, at the same registrar as the rest.
Microsoft's advisory sets customer action required to false. That is correct about the vulnerability and wrong about the incident, and the difference is a memory entry no runbook clears.
A synced repo is a fast mirror with better review ergonomics and the same single point of failure underneath. Only repos created in Origin have no GitHub in the write path.
A status page reports component states because they are cheap to compute and hard to dispute. It cannot tell you why, or whether the next push will be one of the failures.
The commitments still stand and no promise was broken. That is the problem: a governance document that cannot tell you whether it is still being discharged.
GLM-5.3 leads CyberGym by 0.7 points and trails ExploitBench by 23.6. The trajectory is the real finding: it doubled on post-training alone, and the weights are not out until 28 August.
Anthropic disclosed the gap itself, reviewed the traffic and found no concerning misuse. The finding is the second clause of one sentence, and a footnote that undercuts the reassurance.
xAI's agents sign into your tools with your own login and hand work to each other with no human step. Every action lands under your account, and the containment controls are not on the tiers you can buy.
Two of five price Claude Opus at 3x the real rate. All five miss the tokenizer change that makes their token counts about 30% low on current Claude models.
One popular token counter posts your text to its server on every keystroke. Another promises no tracking while loading an analytics script. Here is how to check for yourself.
A price rise in four weeks, cache writes charged above input rate, stacking multipliers, and a tokenizer change that quietly invalidates every character-based estimate.
Brave broke Perplexity Comet twice after it was declared fixed, then repeated the attack against three more agents. The vendors are not losing a patching race; the architecture composes trusted instructions with untrusted content in one context window.
A poisoned tool description is a prompt injection that ships once and fires on every invocation, for every user, in every session. It is the same root cause as the browser attacks, with persistence added.
Executive confidence in AI use is not correlated with anything measured. The organisations that had an incident and the ones that feel fine are largely the same organisations, because neither has an inventory.