25 t/s with the tensor parallel execution 50% on each MacBook instead, via RDMA. Probably could be optimized a lot more in this use case, since I get 15 t/s with SSD streaming on a single machine.
评论
?
参与讨论
25 t/s with the tensor parallel execution 50% on each MacBook instead, via RDMA. Probably could be optimized a lot more in this use case, since I get 15 t/s with SSD streaming on a single machine.