Solo founder building serverless GPU inference on AMD MI300X , looking for honest feedback I will not promote

Solo dev, have spent the last few months building Inferix, a serverless inference platform that runs on AMD MI300X GPUs, the 192GB VRAM ones for context, that's 2.4× the H100). The idea: deploy any model in a Docker image, scale to zero when idle, pay per second. Why AMD instead of NVIDIA? Two reasons. First, MI300X has way more VRAM per card you can fit Llama 70B on a single GPU with no quantisation. Second, the price & performance is meaningfully better for inference workloads. ROCm matured enough in the

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论