My V4 Flash Vision Exp test was good, but local serving still looks expensive

Until this release, I used DeepSeek for text and sent image work to another model. V4 Flash Vision Exp is good enough that I am reconsidering that split. I tried it through ZenMux because that was already wired into my test script, and it got most of the image content I checked right. This was a small personal test, not a benchmark. The local hardware is the difficult part. The checkpoint is 305B, and DeepSeek's published vLLM example uses a single node with four GB300 GPUs. That is well outside a normal desktop budget, especially with accelerator and memory prices where they are now. A team processing images all day might still make the numbers work. The hosted API bill disappears after buying the machine, but power, cooling, maintenance, and idle time still count. High utilization and a need for predictable latency would make local serving easier to justify. My workload is bursty, so most of that capacity would sit unused. The hosted test cannot tell me how much visual quality or speed changes after quantization and local serving. Actual results that include GPU count, quantization, sustained TPS, and any loss in recognition quality would make the hardware decision much easier.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论