We swapped AdamW's optimizer states for a Fast Fourier Transform (FFT) to cut VRAM in half. Anyone else trying non-quantization methods?

Hey everyone, Like most of you, we have been fighting constant OOM errors while trying to fine-tune 8B and 70B models on consumer GPUs. The AdamW optimizer states are always the biggest bottleneck. We didn't want to rely on aggressive 8-bit quantization because we were seeing degradation in convergence, so we tried an experiment: tackling the optimizer states in the frequency domain. The methodology: Instead of storing the full gradients, we transform them using an FFT. This isolates the high-energy signal from the noise. We dynamically drop the low-impact frequencies and compress the state. When we inverse-transform back, it maintains the directional integrity but uses roughly 50% less VRAM. The catch: Running FFT operations adds compute overhead. It takes slightly longer per step, but the trade-off is completely avoiding OOM crashes and pushing batch sizes way up on standard hardware. We are currently giving out access to our internal Colab environment and baseline weights to anyone who wants to poke holes in our math or try to break it. We are really curious if anyone else here is exploring frequency-domain stuff or other non-quantization methods for VRAM reduction?

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论