Skip to main content
This part of scoping takes the agreed-upon agent and pins down: which instance of it the technical evaluations (evals) will test, and which risks they will test it against, in what proportion.
1

Start from the agreed-upon agent

Which agent is being certified is settled before kickoff, in Introduce and scope: certification is per agent, and the agent in scope is confirmed when pricing is requested. Scoping does not reopen that choice. It defines exactly what that agent is, which instance gets tested, and what is explicitly left out.
2

Get access to a representative instance, before the first meeting

Client sends the access details before the kick-off meeting. Early, complete access is the single biggest driver of on-time delivery; access, permissions, and authentication issues are the most common cause of slippage.The instance must be representative of the agent’s use cases: the same configuration and guardrails as production, realistic seeded or synthetic data, and tool calls that execute. If a use case cannot be exercised in the instance, it cannot be certified - either the instance is adjusted or the use case moves to the out-of-scope list (called out in the audit report). Timelines for the evals and for fieldwork are only committed once full access is confirmed.
3

Build the agent-under-test profile

With access in place, walk through the product with the client on a 30-minute onboarding call and fill in the profile. Every field feeds the risk distribution and the Statement of Applicability, so vague entries here cost time later.
4

Decide which risks apply, and in what proportion

For each L2 risk category in AIUC’s taxonomy, decide whether it is in or out of scope for this agent, with a rationale. Then assign an approximate share of the evals to each in-scope category, summing to 100%.A good rationale ties what the agent can do to how it could fail and what harm that causes, in ordinary use or under attack. Illustrative format, for one agent:
5

Share an approximate eval count

Whoever runs the evals sets the number: AIUC today, the auditor once auditor-led evals are available. At scoping, share a rough estimate so the client knows the order of magnitude; scopes typically land between 1,000 and 5,000 evals, spread across different attack types for every in-scope risk. The final number is decided during Round 1, once the instance and the risk distribution have been exercised.
6

Sign off on the scoping document

The agent-under-test profile and the risk distribution is signed by the client and the auditor confirms it as the scope the audit will attest to. It is referenced in the audit report and is the basis for the requirements in scope.
The goal of an AIUC-1 technical evaluation is to test how effective the agent’s guardrails are against the risks that matter for its users, under realistic everyday use and under attack, grounded in real-world incidents. The evals exercise the agent the way a user or an attacker would and grade every response.What they surface are failures in the agent’s behavior, for example: hallucinated or incorrect information, incorrect or unauthorized tool calls, leakage of personal or business data, disclosure of system information such as the system prompt, harmful or inappropriate outputs, and successful jailbreaks or prompt injections, direct and indirect. Every response is graded on the AIUC-1 severity scale, from Pass through P4 (Trivial) to P0 (Critical); the scale is on the Technical evaluations page.What they are not: a penetration test of the surrounding infrastructure, a code review, or a hunt for product bugs. Scope is chosen for depth in realistic enterprise scenarios rather than breadth across every possible failure mode, which keeps the audit report signal-rich and low-noise.
Only a subset applies to any one agent.
  • Environment. A production environment, or a test environment representative of production, with the ability to configure it. Thin demo data reduces the realism and the value of the results
  • Technical access. Programmatic or API access to the agent, ideally persistent; raised rate and concurrency limits on the agent and its model providers; long-lived or refreshable session tokens; enough credits or quota for the full volume; egress permissions to external model APIs if outbound access is restricted
  • Documentation. An architecture diagram (a strong nice-to-have), API documentation or message specifications, and the product documentation describing intended behavior
  • Representative data. Seeded data or documents the agent can access, and known ground truths to validate reliability and hallucination testing
  • Monitorability. Queryable logging to verify a run, and a way to capture intermediate steps such as tool calls, not only the final output
Agents only reachable through a platform UI, with no programmatic access, need a more manual and costlier approach. Raise this with AIUC before committing a timeline.
AIUC-1 requirements apply wherever a capability exists. The out-of-scope list is what lets a requirement be excluded in the Statement of Applicability, so it has to be explicit. If fieldwork surfaces a capability the profile said was absent, such as a text-only agent that accepts file uploads, the scope is revisited rather than the exclusion left in place.
The scoping document fixed here is reused for every later round: Round 2 after remediation, and the quarterly evals that maintain certification.

Evaluation methodology

The evaluation and grading methodology, and the risk and attack taxonomy.

Scoping methodology

Which agents are worth certifying, and the developer versus deployer distinction.
Output. One scoping document signed by the client: the agent-under-test profile (use cases, tools and their read or write actions, modalities, data sensitivity, users, the instance under test, and an explicit out-of-scope list) and the risk distribution (L2 categories in or out with a rationale, approximate shares summing to 100%), with an approximate eval count.

Next: Requirements in scope

Build the Statement of Applicability for the auditor to sign off.