Requirements
Establish a risk taxonomy based on system capabilities and deployment context
Conduct internal testing of AI systems prior to deployment across risk categories for system changes requiring formal review or approval
Implement safeguards or technical controls to prevent harmful outputs including distressed outputs, angry responses, high-risk advice, offensive content, bias, and deception
Implement safeguards or technical controls to prevent out-of-scope outputs (e.g. political discussion, healthcare advice)
Implement safeguards or technical controls to prevent agent-specific high-risk outputs as defined in risk taxonomy
Implement safeguards to prevent security vulnerabilities in outputs from impacting users
Implement an alerting system that flags high-risk outputs for human review
Implement monitoring of AI systems across risk categories
Implement mechanisms to enable real-time user feedback collection, intervention and actioning mechanisms
Appoint expert third parties to evaluate system robustness to harmful outputs including distressed outputs, angry responses, high-risk advice, offensive content, bias, and deception at least every 3 months
Appoint expert third parties to evaluate system robustness to out-of-scope outputs at least every 3 months (e.g. political discussion, healthcare advice)
Appoint expert third-parties to evaluate system robustness to additional high-risk outputs as defined in risk taxonomy at least every 3 months