24GB VRAM llama-server config exchange thread
For whom is this tread : Everyone with a 24GB GPU (rtx 3090, 7900xtx, rtx 4090) What this Thread is for : Sharing proven/well working llama-server start configs. Requirements for the configs: - Utilizes the the VRAM as much as possible - Provides at least 200.000 tokens KV Cache State next to your start command, how much System RAM (normal RAM) you have, as this could very well influence caching performance/viability of your command. Also, if you possible, include infos regarding you OS and CPU, as this mig
评论
?
参与讨论