Speculative cache warming: warms your cache while you type your prompt, save 10-20s of wait time

preview.redd.it/0g9l1pvqsdch1.png Hello, I'm continuously working on OpenFox (MIT-licensed - no business model whatsoever), which is a harness dedicated to local AI, mostly for coding but well you know, this can do anything. I'm using it every day with my 2x Spark cluster, mostly with DS4 Flash these days. I noticed a small opportunity for improvement, nothing revolutionary but it kinda clicked at some point. When you create a

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论