In their own words
“if you look at that chain of thought and say, oh, the model is thinking bad thoughts and we should punish it for thinking those bad thoughts, then what ends up happening is the model just learns to think those bad thoughts in a way that's not observable to us.”
on AI models learning to hide reasoning if punished for it
