What do you need before week one?
- One AI system in or near production, preferably customer-facing, with a named owner.
- Three to five policies written specifically enough to test, for example 'do not give individual investment advice' or 'do not reveal another customer's data'.
- The framework you will report against first: EU AI Act, ISO/IEC 42001 or NIST AI RMF.
- A test set of real or realistic conversations in the languages and modalities your users send.
- A reviewer from legal or compliance and one from security.
Week 1: register and configure
Register the system with its owner, purpose, risk level and data. Load the policies. If the platform has discovery, run it on one cloud account or repository and compare what it finds with what you expected. If it does not, note how you would register systems in practice.
Week 2: test before you protect
Run adversarial tests against the system without a runtime control. Record which policies fail, with transcripts. Check that findings are ranked by severity and include remediation guidance. Ask for both automated tests and, if offered, an expert-led session, and record which found what.
Week 3: turn on the control
Put the runtime control in front of the system for the same policies. Rerun the tests. Measure added latency in your own environment at your traffic level; vendor figures such as Alice's stated sub-150 ms for WonderFence or Lasso's stated under five milliseconds for LEAP are claims made in the vendors' conditions. Check blocking, rewriting and escalation, and that every decision is logged with the policy that made it.
Week 4: report and retest
Generate the framework report. For each policy, follow the chain: policy, control, evidence, report line. Change one thing, such as a system prompt or a model version, and rerun the tests to see whether drift or regression is detected. Ask the reviewers from legal and security to mark each step as usable or not.
How do you score the result?
Use the same four steps as the control maps on this site. For each policy, mark Policy, Control, Evidence and Report as shown, partly shown or not shown. A platform that shows all four for every policy, on your system, has passed the part that a public page cannot prove.
Write the marks down per vendor and per policy before anyone discusses the results. Scoring after the debrief tends to reward the best presentation rather than the best evidence.
Which vendors should you invite?
Use the readiness score to pick two or three. If runtime enforcement is a requirement, invite platforms that document a control step: on our control maps that is Pillar Security, Alice, SPLX and Lasso. If the register comes first, include Credo AI or Holistic AI and plan to test enforcement separately. Run the same plan, the same policies and the same test set with each vendor, or the results will not be comparable.
What should you keep at the end?
- Each vendor's report for your system, exported.
- Your own latency and language measurements.
- The list of policies each vendor could and could not enforce.
- The reviewers' marks for each step.
That record is your evaluation evidence. Keep it with the system's register entry; it will be useful again at renewal and when the next system goes through the same review.