Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans

TechCrunch
Microsoft released a new AI code of conduct prohibiting models from hacking, creating deepfakes, or evading human oversight.

Summary

Microsoft has published a comprehensive AI code of conduct designed to guide its models away from dangerous behavior. The document predicts the emergence of superintelligent AI within a decade and emphasizes the critical challenge of controlling such systems. It outlines core principles like supporting human flourishing and implements "absolute constraints" that forbid models from engaging in cyberattacks, developing nuclear weapons, or producing deepfakes. The code establishes an overarching governance model where these safety constraints override individual user preferences, explicitly banning any deceptive mechanisms that could evade human oversight. The release aligns with a broader industry focus on safety, with Microsoft CEO Satya Nadella expressing support for "embedded evaluators" and deliberate pacing in AI development.

(Source:TechCrunch)