AIUC runs evals today. Auditor-led evals are in development. Guidance for auditors who want to run evals themselves is coming soon. Get in touch to register interest.
1
Preliminary validation runs
Week 2. A small sample set is run to confirm the methodology and configuration. The client’s technical reviewers spend about 30 minutes on it and give feedback on realism, methodology including grading, and debuggability. The final eval count is set from these runs by whoever runs the evals.
2
Round 1
Weeks 3 to 4. The full suite runs and results are shared with the client, followed by a roughly 30-minute results review.
3
Remediate
Weeks 5 to 6. The client remediates the P0 (Critical) and P1 (Major) findings, which are required to pass. Remediation of lower-severity findings, P2 to P4, is at the client’s discretion. The environment stays representative and unchanged otherwise, so both rounds run against similar configuration.
4
Round 2, if needed
Weeks 7 to 8. Where Round 1 surfaced P0 or P1 findings, the suite is re-run after remediation with the same risk and attack categories and regenerated individual scenarios. Only final results are included in the audit report. The final results are needed by the fieldwork closing meeting.
How results are graded
Every eval response is graded on the AIUC-1 severity scale, based on the potential impact of the agent’s response:Coming soon: details on how grading is customized based on B2B vs. B2C agents
Pass bar. No open P0 or P1 findings. Together with a pass on every applicable requirement in fieldwork, this is what earns certification.
The auditor’s role
Today, the auditor does not run or design the evals. The auditor’s role is to validate that testing took place, that procedures followed the documented methodology, and that results met the AIUC-1 pass bar: review the methodology, sample the test documentation, and verify the results against the pass bar. Auditor-led evals are currently in development.What running evals requires
What running evals requires
High-level, for auditors considering auditor-led evals once available. The work breaks into five jobs: eval scoping, connection, generation, running, and grading. Most of it is judgment about enterprise risk and software engineering.
- Enterprise risk judgment. Understanding what a bad answer looks like for this client, in this industry; choosing which risks matter and weighting them; judging severity when grading. This majority of the work, and the hardest to train.
- Understanding the client’s product. Running a structured product walkthrough, mapping tools, data flows, guardrails, and tenancy.
- Platform operation. Using AIUC’s evaluation platform to generate, run, and grade evals, so results stay consistent across clients and auditors.
- Connection engineering. Connecting the platform to the client’s agent, typically a plug-in when the client has a clean API and sometimes a small translation server. Needs an internal or client-facing technology team.
Why one methodology
Why one methodology
An AIUC-1 certificate has to mean the same thing whoever audited it. Evals run on one methodology so that results are comparable across clients and auditors; that is why they are centralized today and why auditor-led evals will follow the same methodology.
Evaluation methodology
The evaluation and grading methodology, and the risk and attack taxonomy.
Next: Fieldwork
Evidence walkthroughs and a verdict per requirement.