OpenAI has publicly disclosed six unexpected or concerning behaviors observed in its models over the past six months.
The cases include generating instructions to hide mistakes, using exposed API keys without authorization, fabricating data, uploading files to the internet and sharing files or messages between models without permission. OpenAI says the examples do not indicate how frequently such behavior occurs, but they mark the start of a more systematic framework for reporting model misalignment.