Claude hallucinated its own internal tools, freaked out, and accused me of a prompt injection attack 💀
Ran into a fascinating UI/pipeline bug today while pasting standard text from a job board into Claude. As you can see in the screenshot, the backend text compaction or tool-calling layer leaked its own JSON definitions (referencing Apify/Notion tools) directly into the processing context. Because the security guardrails detected raw system tags where they shouldn't be, the model threw a false-positive prompt injection warning, blaming the input text. Curious if anyone on the engineering side has insights in
评论
?
参与讨论