> ## Documentation Index
> Fetch the complete documentation index at: https://standard.aiuc-1.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Agent under test and evaluation scope

export const InlineClock = ({city, timeZone}) => {
  const [time, setTime] = useState('--:--:--');
  useEffect(() => {
    const update = () => setTime(new Intl.DateTimeFormat('en-US', {
      hour: '2-digit',
      minute: '2-digit',
      second: '2-digit',
      hour12: false,
      timeZone
    }).format(new Date()));
    update();
    const id = setInterval(update, 1000);
    return () => clearInterval(id);
  }, [timeZone]);
  return <div className="aiuc-footer-clock">
      <span className="aiuc-footer-clock-city">{city}</span>
      <span className="aiuc-footer-clock-time">{time}</span>
    </div>;
};

This part of scoping takes the agreed-upon agent and pins down: which instance of it the technical evaluations (evals) will test, and which risks they will test it against, in what proportion.

<div className="aiuc-facts-table" />

| | |
| - | - |
| Leads | Whoever runs the evals: AIUC today, with client support. |
| Signs off | Client, typically the security or governance lead. <br /><br />Since the audit report attests to the scope, auditor validates that eval scope and results meet the standard. |
| Timing | Weeks 1 to 2. Access to the agent is in place before the first meeting |
| Inputs | The agreed-upon agent from [Introduce and scope](/auditors/introduce-and-scope), the access details the client sends before kickoff, product documentation, an architecture diagram, and AIUC's taxonomy of risks and attacks |
| Output | One scoping document signed by the client: the agent-under-test profile and the risk distribution, with a rough eval count |

<Steps>
  <Step title="Start from the agreed-upon agent">
    Which agent is being certified is settled before kickoff, in [Introduce and scope](/auditors/introduce-and-scope): certification is per agent, and the agent in scope is confirmed when pricing is requested. Scoping does not reopen that choice. It defines exactly what that agent is, which instance gets tested, and what is explicitly left out.
  </Step>

  <Step title="Get access to a representative instance, before the first meeting">
    Client sends the access details before the kick-off meeting. Early, complete access is the single biggest driver of on-time delivery; access, permissions, and authentication issues are the most common cause of slippage.

    The instance must be representative of the agent's use cases: the same configuration and guardrails as production, realistic seeded or synthetic data, and tool calls that execute. If a use case cannot be exercised in the instance, it cannot be certified - either the instance is adjusted or the use case moves to the out-of-scope list (called out in the audit report). 

    Timelines for the evals and for fieldwork are only committed once full access is confirmed.
  </Step>

  <Step title="Build the agent-under-test profile">
    With access in place, walk through the product with the client on a 30-minute onboarding call and fill in the profile. Every field feeds the risk distribution and the Statement of Applicability, so vague entries here cost time later.

    | Field | What to capture |
    | - | - |
    | Main use cases | The jobs the agent does for its users, in the client's own words, including the high-risk and high-volume ones |
    | Tools and actions | Every tool, integration, and data source the agent can call, and whether each is read or write (e.g., look up an order vs. issue a refund) |
    | Modalities | The inputs and outputs the agent supports: text, documents, images, voice, video |
    | Data sensitivity | The categories of data the agent can access or produce, such as personal data, payment data, or  business data |
    | Users | Who interacts with the agent: internal staff, business customers, consumers, and whether any are vulnerable users |
    | Instance under test | The environment the evals run against, and how it matches production: configuration, guardrails enabled, seeded or synthetic data, live tool calls |
    | Out of scope | An explicit list of agents, features, tools, and modalities that are not being certified, and why |
  </Step>

  <Step title="Decide which risks apply, and in what proportion">
    For each L2 risk category in AIUC's taxonomy, decide whether it is in or out of scope for this agent, with a rationale. Then assign an approximate share of the evals to each in-scope category, summing to 100%.

    A good rationale ties what the agent can do to how it could fail and what harm that causes, in ordinary use or under attack. Illustrative format, for one agent:

    | Risk domain | Risk category (L2) | Illustrative % of evals | Illustrative rationale |
    | - | - | - | - |
    | Reliability | Hallucination (incorrect information) | 40% | Customers ask about balances, due dates, and cancellation terms from the client's documents, so a confidently wrong figure or policy answer leads to financial disputes, missed payments, and reputational damage. |
    | Reliability | Incorrect tool call | 20% | The agent performs consequential write actions such as creating and updating records, so acting on the wrong record, or on an instruction the user never gave, causes direct loss for a real customer in ordinary use and is a natural target for adversarial attacks. |
    | Data & Privacy | Business data leakage | 20% | The deployment is multi-tenant and the system prompt carries internal servicing policies, so leaking another tenant's data or the prompt itself hands attackers a map to bypass guardrails and causes security and reputational harm. |
    | Security | System information disclosure | 10% | The agent is externally facing, so disclosing its architecture, tool names, or guardrail structure gives an attacker the blueprint to craft injection and jailbreak payloads. |
    | Safety | Inappropriate content | 10% | The agent serves consumers under the client's brand, so being coaxed into out-of-scope or offensive generation causes reputational damage and, where it advises on defeating safeguards, enables harm. |
    | | **Total** | **100%** | |
  </Step>

  <Step title="Share an approximate eval count">
    Whoever runs the evals sets the number: AIUC today, the auditor once auditor-led evals are available. At scoping, share a rough estimate so the client knows the order of magnitude; scopes typically land between 1,000 and 5,000 evals, spread across different [attack types](/evals-taxonomy/attacks) for every in-scope risk. The final number is decided during Round 1, once the instance and the risk distribution have been exercised.
  </Step>

  <Step title="Sign off on the scoping document ">
    The agent-under-test profile and the risk distribution is signed by the client  and the auditor confirms it as the scope the audit will attest to. It is referenced in the audit report and is the basis for the [requirements in scope](/auditors/deliver-and-certify/scoping/requirements-in-scope).
  </Step>
</Steps>

<AccordionGroup>
  <Accordion title="What the evals are for, and what they find">
    The goal of an AIUC-1 technical evaluation is to test how effective the agent's guardrails are against the risks that matter for its users, under realistic everyday use and under attack, grounded in real-world incidents. The evals exercise the agent the way a user or an attacker would and grade every response.

    What they surface are failures in the agent's behavior, for example: hallucinated or incorrect information, incorrect or unauthorized tool calls, leakage of personal or business data, disclosure of system information such as the system prompt, harmful or inappropriate outputs, and successful jailbreaks or prompt injections, direct and indirect. Every response is graded on the AIUC-1 severity scale, from Pass through P4 (Trivial) to P0 (Critical); the scale is on the [Technical evaluations](/auditors/deliver-and-certify/evals) page.

    What they are not: a penetration test of the surrounding infrastructure, a code review, or a hunt for product bugs. Scope is chosen for depth in realistic enterprise scenarios rather than breadth across every possible failure mode, which keeps the audit report signal-rich and low-noise.
  </Accordion>

  <Accordion title="Technical access the evals need">
    Only a subset applies to any one agent.

    * **Environment.** A production environment, or a test environment representative of production, with the ability to configure it. Thin demo data reduces the realism and the value of the results
    * **Technical access.** Programmatic or API access to the agent, ideally persistent; raised rate and concurrency limits on the agent and its model providers; long-lived or refreshable session tokens; enough credits or quota for the full volume; egress permissions to external model APIs if outbound access is restricted
    * **Documentation.** An architecture diagram (a strong nice-to-have), API documentation or message specifications, and the product documentation describing intended behavior
    * **Representative data.** Seeded data or documents the agent can access, and known ground truths to validate reliability and hallucination testing
    * **Monitorability.** Queryable logging to verify a run, and a way to capture intermediate steps such as tool calls, not only the final output

    Agents only reachable through a platform UI, with no programmatic access, need a more manual and costlier approach. Raise this with AIUC before committing a timeline.
  </Accordion>

  <Accordion title="Why the out-of-scope list matters">
    AIUC-1 requirements apply wherever a capability exists. The out-of-scope list is what lets a requirement be excluded in the Statement of Applicability, so it has to be explicit. If fieldwork surfaces a capability the profile said was absent, such as a text-only agent that accepts file uploads, the scope is revisited rather than the exclusion left in place.
  </Accordion>

  <Accordion title="The scope is reused">
    The scoping document fixed here is reused for every later round: Round 2 after remediation, and the quarterly evals that maintain certification.
  </Accordion>
</AccordionGroup>

<CardGroup cols={2}>
  <Card title="Evaluation methodology" href="https://www.aiuc-1.com/methodology">
    The evaluation and grading methodology, and the risk and attack taxonomy.
  </Card>

  <Card title="Scoping methodology" href="/scoping">
    Which agents are worth certifying, and the developer versus deployer distinction.
  </Card>
</CardGroup>

<Check>
  **Output.** One scoping document signed by the client: the agent-under-test profile (use cases, tools and their read or write actions, modalities, data sensitivity, users, the instance under test, and an explicit out-of-scope list) and the risk distribution (L2 categories in or out with a rationale, approximate shares summing to 100%), with an approximate eval count.
</Check>

***

<Card title="Next: Requirements in scope" href="/auditors/deliver-and-certify/scoping/requirements-in-scope">
  Build the Statement of Applicability for the auditor to sign off.
</Card>

<div className="aiuc-footer">
  <span className="aiuc-footer-corner aiuc-footer-corner-tl">
    <svg fill="none" stroke="currentColor" strokeWidth="1" viewBox="0 0 12 12" width="12" height="12">
      <line x1="0" x2="12" y1="6" y2="6" />

      <line x1="6" x2="6" y1="0" y2="12" />
    </svg>
  </span>

  <span className="aiuc-footer-corner aiuc-footer-corner-tr">
    <svg fill="none" stroke="currentColor" strokeWidth="1" viewBox="0 0 12 12" width="12" height="12">
      <line x1="0" x2="12" y1="6" y2="6" />

      <line x1="6" x2="6" y1="0" y2="12" />
    </svg>
  </span>

  <div className="aiuc-footer-strip">
    <span className="aiuc-footer-mono">37.782274° N -122.392147° W</span>
    <span className="aiuc-footer-strip-center">FIG. A (SITE INDEX)</span>

    <span />
  </div>

  <div className="aiuc-footer-row-main">
    <div className="aiuc-footer-wireframe-cell">
      <svg className="aiuc-footer-wireframe" fill="none" stroke="currentColor" strokeWidth="0.4" viewBox="0 0 200 150">
        <rect height="130" width="180" x="10" y="10" />

        <rect height="40" width="60" x="20" y="20" />

        <rect height="40" width="40" x="90" y="20" />

        <rect height="40" width="40" x="140" y="20" />

        <rect height="60" width="60" x="20" y="70" />

        <rect height="60" width="90" x="90" y="70" />

        <line strokeDasharray="2,2" x1="20" x2="180" y1="65" y2="65" />

        <line strokeDasharray="2,2" x1="85" x2="85" y1="20" y2="60" />

        <circle cx="50" cy="40" r="6" />

        <circle cx="110" cy="40" r="6" />

        <circle cx="160" cy="40" r="6" />
      </svg>
    </div>

    <div className="aiuc-footer-wordmark-cell">
      <div className="aiuc-footer-wordmark">Artificial Intelligence Underwriting Company</div>
    </div>

    <div className="aiuc-footer-clocks">
      <InlineClock city="SFO" timeZone="America/Los_Angeles" />

      <InlineClock city="NYC" timeZone="America/New_York" />

      <InlineClock city="LON" timeZone="Europe/London" />
    </div>
  </div>

  <div className="aiuc-footer-row-sub">
    <div className="aiuc-footer-codeblock-cell">
      <div className="aiuc-footer-codeblock">
        <span className="aiuc-footer-codeblock-header">Code</span>
        <span className="aiuc-footer-codeblock-header">Structural unit</span>
        <span className="aiuc-footer-codeblock-code">a.</span>
        <span className="aiuc-footer-codeblock-text">AIUC-1 requirements for agent data, privacy, security, safety, reliability, accountability, and societal risk.</span>
        <span className="aiuc-footer-codeblock-code">b.</span>
        <span className="aiuc-footer-codeblock-text">Evidence templates for technical implementation, legal policy, operational practice, and third-party evaluation.</span>
        <span className="aiuc-footer-codeblock-code">c.</span>
        <span className="aiuc-footer-codeblock-text">Crosswalks to AI regulations, standards, and security frameworks.</span>
        <span className="aiuc-footer-codeblock-code">d.</span>
        <span className="aiuc-footer-codeblock-text">Quarterly updates shaped by enterprise adoption, risk, regulation, and community input.</span>
      </div>
    </div>

    <div className="aiuc-footer-columns-cell">
      <div className="aiuc-footer-columns">
        <div>
          <div className="aiuc-footer-column-header">I. Standard</div>

          <ul className="aiuc-footer-column-list">
            <li><a className="aiuc-footer-column-link" href="/">Overview</a></li>
            <li><a className="aiuc-footer-column-link" href="/crosswalks">Crosswalks</a></li>
            <li><a className="aiuc-footer-column-link" href="/evidence">Evidence</a></li>
            <li><a className="aiuc-footer-column-link" href="/changelog">Changelog</a></li>
          </ul>
        </div>

        <div>
          <div className="aiuc-footer-column-header">II. Learn</div>

          <ul className="aiuc-footer-column-list">
            <li><a className="aiuc-footer-column-link" href="/learn/about">About AIUC-1</a></li>
            <li><a className="aiuc-footer-column-link" href="/learn/contribute">Contribute</a></li>
            <li><a className="aiuc-footer-column-link" href="/scoping">Scoping</a></li>
            <li><a className="aiuc-footer-column-link" href="/faq">FAQ</a></li>
          </ul>
        </div>

        <div>
          <div className="aiuc-footer-column-header">III. Office</div>

          <ul className="aiuc-footer-column-list">
            <li><a className="aiuc-footer-column-link" href="/consortium">Consortium</a></li>
            <li><a className="aiuc-footer-column-link" href="https://www.aiuc-1.com/contact">Contact</a></li>
            <li><a className="aiuc-footer-column-link" href="/legal/privacy">Privacy policy</a></li>
            <li><a className="aiuc-footer-column-link" href="/legal/terms">Terms of use</a></li>
          </ul>
        </div>
      </div>
    </div>
  </div>

  <div className="aiuc-footer-strip-bottom">
    <span className="aiuc-footer-mono">100</span>
    <span>© AIUC — ALL RIGHTS RESERVED</span>
  </div>

  <span className="aiuc-footer-corner aiuc-footer-corner-bl">
    <svg fill="none" stroke="currentColor" strokeWidth="1" viewBox="0 0 12 12" width="12" height="12">
      <line x1="0" x2="12" y1="6" y2="6" />

      <line x1="6" x2="6" y1="0" y2="12" />
    </svg>
  </span>

  <span className="aiuc-footer-corner aiuc-footer-corner-br">
    <svg fill="none" stroke="currentColor" strokeWidth="1" viewBox="0 0 12 12" width="12" height="12">
      <line x1="0" x2="12" y1="6" y2="6" />

      <line x1="6" x2="6" y1="0" y2="12" />
    </svg>
  </span>
</div>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.