Discussion about this post

User's avatar
Fukitol's avatar

What is it exactly that they are doing? I'd assume more, I can't remember the term atm, tentative token prediction and branch comparison before final token selection. A memory:precision tradeoff vs. "chain of thought" essentially.

CoT monitoring is not really that useful in the first place. There's this delusional belief that you're seeing some sort of accurate reporting on internal model state, when actually it's a context refinement hack. There's no way to know what the model is "thinking" because it isn't thinking in the first place, and you can't with any accuracy predict what it will predict next based on prior context.

If you want security, stop the model from *doing* things it shouldn't do. The tokens it predicts are irrelevant to security if they can't result in real world damage (other than hurt feelings anyway) because you properly constrained it to only the things which it should do, instead of foolishly hoping you could coerce it into avoiding doing bad things through context manipulation.

mancaded's avatar

You said there would be a black swan event that would scupper the AI craze. Here it comes.

"Unpredictable, unalignable LLM knocks over major bank/utility because OpenAI needed some key to jangle".

66 more comments...

No posts

Ready for more?