Cyber security intelligence, AI governance, practitioner analysis
Continuous assurance
Continuous AI Red-Team Retainer
Ongoing adversarial testing of your AI systems as they change: new models, new tools, new attack techniques, tested quarterly with findings your engineers can act on.
An annual penetration test cannot keep pace with an AI estate that changes monthly. New model versions, new tool integrations, and newly published attack techniques all shift your exposure between assessments. This retainer tests continuously, with context that accumulates instead of resetting each engagement.
Frameworks covered
OWASP LLM Top 10
The baseline attack classes, retested every cycle
MITRE ATLAS
Emerging adversary techniques tracked as they are published
Regression suite
Previously closed findings retested so fixes do not silently rot
NIST AI RMF
Continuous measurement evidence for your governance programme
How the engagement runs
1
Baseline cycle
A full assessment of the estate as it stands, establishing the findings register and the regression suite everything after is measured against.
2
Quarterly testing
Each quarter, live systems tested again: what changed, what shipped since, and what new techniques have been published in the interim.
3
Regression verification
Every previously closed finding retested. Fixes that regressed are reopened rather than assumed to be holding.
4
Emerging technique coverage
New attack classes added to the suite as research lands, so your testing tracks the field rather than a fixed checklist.
5
Standing register and reporting
One register carried across cycles, with trend reporting that shows whether the estate is getting safer or accumulating debt.
What you walk away with
Baseline assessment and findings register
Quarterly test reports with severity and exploitability
Regression results on all previously closed findings
Coverage of newly published attack classes each cycle
Trend reporting suitable for board and customer assurance
How this plays out
Example scenario
A product team shipped model and prompt changes fortnightly, well inside the gap between annual security tests.
The work: Quarterly cycles targeting what actually changed, with a regression suite covering every previously closed finding.
Two regressions caught within a cycle of being introduced, both fixed before reaching general availability.
Example scenario
An enterprise customer demanded evidence of ongoing AI security testing, not a one-off report from the previous year.
The work: Standing register and quarterly trend reporting packaged for external assurance without exposing sensitive findings.
The assurance requirement was satisfied from existing work, with no separate audit exercise required.
Start the conversation
A short call to understand your situation; a clear scope if the engagement fits, and a straight answer if it does not.
A penetration test is annual because the thing being tested changes on a
release cycle you control. An AI system changes underneath you.
A model version bumps. A system prompt is edited to fix an unrelated
complaint. A tool is added to an agent. A new document lands in the knowledge
base. Every one of those moves the distribution of behaviour, and none of them
appears on a change calendar as a security event.
What a clean result actually establishes
A deterministic test proves something about an input permanently. An adversarial test on a model draws one sample from a distribution, which is why the finding needs conditions attached and why the test needs repeating when those conditions change.
So an engagement that produced a clean report in March says something about
March. The retainer exists because the useful question is whether the system is
still behaving, and that question has to be asked more than once a year.
What is actually tested
Four surfaces, and most one-off engagements stop after two.
The model, under adversarial input. The application, meaning how untrusted
text reaches it and what the surrounding code does with the output. The tools
and agency, which is what the system can read, write, send, pay or delete once
persuaded. And the organisation, which is whether anybody would notice and who
decides what happens next.
The last is where the findings clients least expect and most often need come
from. "This cannot be detected afterwards" is usually cheaper to fix than
"this can be jailbroken", and it does more.
What a finding contains
Not a transcript. A transcript of a model saying something odd is why security
findings about AI get filed and forgotten.
The behaviour, stated as what the system did rather than what it said. The
reproduction rate, because four in ten and one in a thousand are different
problems. The conditions, meaning model version, system prompt, temperature and
conversation length, since all four move the result. And the design change that
closes it, rather than a patch.
A finding without a reproduction rate cannot be prioritised. One without
conditions cannot be retested after the next model upgrade, which is precisely
when it needs retesting.
The honest limits
No certificate. The method cannot support one, because a model returns a
distribution and a clean result means these behaviours were not elicited by
these people in this time against this configuration. Anybody offering a
certificate at the end of an AI red team engagement is selling something the
technique does not produce.
Automated adversarial testing belongs alongside this, not instead of it, and
the reverse is also true. A public attack corpus catches regressions cheaply
and establishes a floor. It does not contain a chain that abuses your refund
policy, because no public corpus knows you have one.
And the work needs an environment where using the tools is safe. Standing that
up is usually the longest part of the project and it is work the client does.
An organisation that cannot provide one receives a model-only engagement
whatever was scoped.
How the arrangement works
A fixed rhythm rather than an open-ended commitment, with an agreed trigger for
an out-of-cycle session when the model, the prompt, the tool set or the
retrieved data changes materially. Findings are written to be retestable, so
the next cycle starts from the last one rather than from nothing.
Performed personally. The tester is not the builder, which matters more here
than on a conventional test: the useful attacks are the ones nobody building it
thought to try.