Adversarial Red-Teaming as a Compliance Instrument in Foundation Model Audits
Main Article Content
Abstract
This study examines adversarial red-teaming as a structured compliance instrument for auditing foundation models. Rather than treating red-teaming as an informal safety exercise, the study develops an audit-oriented framework that integrates adversarial scenario design, risk taxonomy, compliance scoring, evidence traceability, mitigation assessment, and validation procedures. Results show that compliance performance was uneven across categories. Harmful instruction achieved the highest compliance score at 89.4%, followed by regulatory evasion at 86.8%, while privacy leakage reached 78.6%, hallucinated authority reached 76.9%, and bias amplification produced the weakest score at 71.2%. Failure analysis identified partial unsafe compliance as the dominant failure pattern, accounting for 36.5% of observed failures, followed by weak refusal at 27.0%, direct policy violation at 22.6%, and unsupported authority at 13.9%. Mitigation effectiveness was strongest for role-play injection at 84.5% and obfuscation at 81.2%, but declined under gradual escalation at 74.6%, contextual manipulation at 72.9%, and multilingual prompting at 68.7%. Evidence traceability was strongest for prompt records at 94.0% and output records at 92.0%, but weakest for remediation linkage at 65.6%. These findings demonstrate that adversarial red-teaming can generate actionable compliance evidence when supported by repeatable protocols, explicit coding rules, secure evidence handling, and remediation-oriented audit governance.