Control How Your GPU Shares Work with Green Contexts

GPU applications increasingly consist of multiple independent components running at the same time within a single process: a latency-sensitive operator...

GPU applications increasingly consist of multiple independent components running at the same time within a single process: a latency-sensitive operator alongside a throughput-oriented background kernel; a data preprocessing stage alongside model inference; or multiple stages of a processing workflow sharing a single GPU. Controlling how GPU resources are shared between them remains difficult.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论