Calling classifiers "decision models" is a crime against machine learning. Those models decide less than LLMs, as you can think the CoT as a form of serial processing of information that such classifiers lack. Those models can be either small LLMs that are not let think, but instead the logits at the last position of the prefill are used to categorize in classes or, when they are BERT-alike, they compress the input meaning and project a class: in both cases they decide a lot less than an LLM.
评论
?
参与讨论