The Model Server
Over the last two articles we went through everything needed to run a model: an SDK that runs inference, and a management layer that downloads models, sizes them for your hardware and loads them. You can use the SDK directly from your Go code, but one of the most popular ways of consuming LLMs is through an HTTP API, and that’s what’s left: the server that lets your existing OpenAI client talk to all of it.
评论
?
参与讨论