Where are the Cognitive-Science based Safety Researchers?

I’ve always been interested in AI research from a Cognitive Science perspective, and I’ve found that researchers in the Bayesian Cognitive Science paradigm(Josh Tenenbaum and crew) have been developing statistical models of intelligence that can learn based on limited information and do prediction and simulations, which could also explain planning. I’ve also noticed certain Neuroscience(Dileep George and crew) researchers converge on a similar Bayesian paradigm.

I understand people are very worked up about LLMs these days and this dominates AI risk concerns but I can easily imagine a world where there is some fundamental information efficiency constraint on LLMs and AGI/ASI depends on the kind of efficient, compositional world models that the Cog Sci researchers above are looking into.

Through a Cognitive Science lens, we can see values as a function of world models and innate rewards: a constructured world model(which includes one's self) can map hypothetical states of the world to expected future rewards(values). Even if safety researchers don’t want to investigate how world models work due to fears of increasing capabilities, there should at least be more research into understanding how human innate rewards, a key component of human values, work.

The question I’m trying to ask is: Where are the Cog Sci based AI safety researchers? Alignment should be easier if you know what the AI is going to look like, and we would like to see a distribution of safety researchers proportional to how likely we think each AI approach is to work, and yet the only researcher I've seen taking this approach is Steven Byrnes from a Neuroscience approach.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论