Why does testing count as evidence?
A policy says what should happen. A log shows what did happen on real traffic. A red team result shows what happens when someone tries to break the policy on purpose. Frameworks ask for all three: the EU AI Act's risk management and post-market monitoring, ISO/IEC 42001's monitoring and evaluation, and the NIST AI RMF's Measure function all expect testing.
What makes a finding usable?
- It names the policy it tested, not only a generic attack category.
- It records the system, version and date.
- It includes the transcript or input that caused the failure.
- It has a severity and a recommended fix.
- It links to the retest that shows the fix worked.
When should testing happen?
Before launch, to decide whether the system is ready. After launch, because models, prompts, tools and attack techniques change. Alice splits this into two products: WonderBuild runs pre-launch tests based on custom policies, ranks findings by severity with remediation and maps test documentation to regulatory requirements; WonderCheck runs ongoing post-launch red teaming with drift and regression detection. Pillar Security's RedGraph engine validates findings with transcripts. Lasso describes closed-loop remediation and auto guardrail patching. SPLX names automated red teaming alongside runtime threat inspection.
How should findings reach the control?
The most useful loop sends findings into the runtime control. Alice says WonderCheck findings flow into WonderFence or back to WonderBuild; Pillar says its guardrails calibrate from red teaming findings; Lasso describes auto guardrail patching. Ask each vendor to show one finding, the control change it caused and the retest.
Automated, managed or both?
Automated testing runs at scale and on a schedule. Expert-led testing finds what automated suites miss, especially for a specific business context. Alice's platform page describes automated adversarial testing combined with expert-led red teaming, and WonderBuild is described as automated and managed. For governance evidence, record which kind each test was.
How should results be stored for an audit?
Keep test results with the system's register entry: the date, the system version, the policies tested, the attack categories used (mapped to OWASP or MITRE ATLAS where possible), the findings with severity and the retest results. Keep them exportable, so the record survives a change of vendor. An auditor will want to see that tests ran on a schedule, not only once before launch, and that serious findings were closed.