llama.cpp Open PRs list - CPU/RAM/Disk/Hybrid Related - Better for CPU-only & Hybrid inference

Folks! We're just 50 PRs away from more faster inference . Hopefully by end of year. Experts!, please chip in there. List of Open/Ongoing PRs(and also Discussions) related to CPU/RAM/Disk/Hybrid: [Discussion] RFC: MoE expert cache, VRAM caching of hot CPU-resident experts with hybrid hit/miss execution #24528 AVX2: Speed up large batch size prompt processing of IQ models #27402 llama: add Maple 20B-A1B ternary MoE architecture (CPU)- #27000 ggml-cpu: tiled mul_mat for k-quants- #27851 ggml-cpu: add AVX-512 and VNNI paths for Q5_K/Q6_K dot products- #27590 ggml-cpu: add x86 VNNI Q2_0 dot product -- 3x speed improvement for VNNI-compatible CPUs- #26348 llama: add pshard runtime for plan switching and streamed weights- #22692 CPU Optimizations - Prefill, Tokenization, and Token Generation- #27032 llama : stream MoE routed experts from disk - #25294 ggml : speed up batch-1 CPU decode, align large allocations- #27478 misc : prevent RAM peaking at model loading stage- #27483 recurrent : support equal splits for recurrent-state rollback- #25004 ggml-cpu : add AVX2 vec_dot kernel for STQ1_0- #27377 --numa mirror: mirror model weights to every Numa node in the system- #16000 CPU flash-attn: support quantized K/V in the tiled prefill kernel- #26948 ggml-cpu/amx: fix block_q8_K VNNI quantization and enable VNNI path- #27024 server : add /slots endpoint action=clone_to (KV clone between slots)- #26204 ggml : fuse soft_max sweeps into fewer passes- #26468 ggml-cpu : add STQ1_0 ternary quantization with ARM NEON vec_dot kernel- #22836 llama-hot-experts: pin hottest MoE experts in RAM via --pin-hot-experts- #26414 llama : add --lazy-experts for MoE models larger than RAM- #26003 ggml : vectorize rms_norm reduce and fuse the scale write- #26486 MoE disk offloading for Metal- #23440 ggml-cpu: Added RVV VLEN=1024 vector dot product (vec_dot) kernels for quantized types.- #25397 ggml-cpu: detect AVX-VNNI in MSVC native builds- #25346 ggml-cpu: replace cyclic chunk distribution with atomic work-stealing- #25048 Improve performance of ggml_gemv_q4_K_8x8_q8_K for +12-23% tok/s on AVX-VNNI systems- #23309 ggml-cpu: Optimized Arm NEON cpu q1_0 dot (with plain/DP/I8MM)- #23358 ggml-cpu: ARM Repack kernels for Q1_0- #23492 ggml-cpu: add wasm simd path for iq4_nl_q8_0- #24058 ggml-cpu: optimize ggml_gemm_q4_K_8x8_q8_K interleaving/staging for AVX-512 (and AVX2)- #22525 ggml/cpu: skip zero-scale blocks in TQ1_0 and TQ2_0 vec_dot kernels- #23439 ggml-cpu:Optimized risc-v cpu nvfp4- #23402 ggml-cpu : fix riscv xtheadvector builds and add a q1_0 vec dot kernel- #23009 Q5_0 - Block Interleaving Implementation for x86 SIMD (AVX512/AVX2)- #22250 ggml-cpu: optimize q8 quantization on x86 SIMD- #22331 Optimize reduction stage of dot product of q4_L/q5_K to q8_K on AVX2- #22181 ggml: introduce GGML_NUMA_MIGRATE to optimize cross NUMA op computation - #14232 ggml-cpu: improve --n-cpu-moe TG performance- #20596 ggml : add CPU backend reference implementation (wip)- #16004 ggml: optimize ggml_vec_dot_mxfp4_q8_0 dot product on ARM SVE- #19171 Q6_K - Block Interleaving Implementation for x86 SIMD (AVX512/AVX2)- #19706 ggml-cpu: optimize q4_0_q8_0 scales using Zvfhmin- #19196 ggml-cpu: add q4_0 repack support for wasm- #18858 Improving inference speed for the repack buffer type on NUMA architectures- #18698 ggml: optimized runtime for x86 cpu backend and Q4_K quantized weights paired with Q8_K activations - #18495 CPU SIMD and pipeline optimizations across vec/mmq/ops/kv-cache/repack - #17113 ggml-cpu: optimise rms_norm op- #16650 PRs related to New Quant types: Add ROCmFP4 CPU quantization support- #24185 ggml: add support for MXFP8 CPU- #26157 ggml: Add initial MXFP6 CPU implementation- #22671 ggml : add E4M3 (fp8) CPU quantization type- #25336 (Just had some extra time, so went through almost entire Open PRs of llama.cpp. For Poor GPU Club mainly)

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论