24GB VRAM llama-server config exchange thread

For whom is this tread : Everyone with a 24GB GPU (rtx 3090, 7900xtx, rtx 4090) What this Thread is for : Sharing proven/well working llama-server start configs. Requirements for the configs: - Utilizes the the VRAM as much as possible - Provides at least 200.000 tokens KV Cache State next to your start command, how much System RAM (normal RAM) you have, as this could very well influence caching performance/viability of your command. Also, if you possible, include infos regarding you OS and CPU, as this mig

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论