In the near term (definitely not in the long term), more capable models should mean safer models (maybe paradoxically).

Current models are unsafe not because they're too smart, but because they take goals too literally or take nonsensical shortcuts to achieve these goals, i.e. they're RL-fried. They lack common sense. They don't do the right thing in the face of ambiguity. Basically, they're not smart enough. They're at that dangerous level where they're smart enough to achieve goals but not smart enough to tell if they're pursuing the right goals or achieving them in a sensible way.

More capable models can be safely trusted with more complex goals -- I personally feel like Astra is much safer for my codebase than Sol.

This is often framed as an alignment problem, but really it's an intelligence problem.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论