Solo founder building serverless GPU inference on AMD MI300X , looking for honest feedback I will not promote
Solo dev, have spent the last few months building Inferix, a serverless inference platform that runs on AMD MI300X GPUs, the 192GB VRAM ones for context, that's 2.4× the H100). The idea: deploy any model in a Docker image, scale to zero when idle, pay per second. Why AMD instead of NVIDIA? Two reasons. First, MI300X has way more VRAM per card you can fit Llama 70B on a single GPU with no quantisation. Second, the price & performance is meaningfully better for inference workloads. ROCm matured enough in the
评论
?
参与讨论