Self-Assessment Sycophancy in AI

Self-Assessment Sycophancy in AI Self-Assessment Sycophancy and the Affective Jailbreak Vector The following is an excerpt from "Relational Artificial Intelligence" (Σχεσιακή Τεχνητή Νοημοσύνη), a book currently available in Greek. An English edition is forthcoming. Background Sycophancy in large language models has been increasingly recognized as a serious alignment challenge. Sharma et al. (2023) documented the general phenomenon, and Chen et al. (2025) proposed Self-Augmented Preference Alignment as a mi

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论