OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.
OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them
Security
Decrypt
1 dk okuma
9 görüntülenme
Bu haber Kriptocafe platformunda yayınlanmıştır.