How would you go about using multiple models together from a singe router (?) or a end point ?

I have 3 machines, My main one can run Qwen3.8 Flash Next at 13-15 tps, i also have an MacMini 16GB whcih can run Orninth 9B or Gemma4 12B easily and i have a Pi5 8B that can run a 3B model well. I want to run an EndPoint/Router that is connected to the harness, that breaks down the task and distributes it among these models. Some Background to this, I recently started using Claude Code, I have been noticing how it distributes work among, that is what makes it so fast. Ithis was not the case with Codex and Sol/Astra. I am wondering if there is any preexisting way to do this ?

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论