Literature Review: LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load | Bnechmarking LLMs on Phones [R]

Just finished reading the paper: LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load I am starting to benchmark LLMs on edge devices, particularly phones thus been reading a lot on the what has been done and what is currently being done and wanted to share you my journey of reading such papers and my takes on them. This is one of the only papers I have read that have benchmarked RPi5-Hailo (Hailo's 10H) iPhone 16 Pro (A19 Pro) S24 Ultra (Snapdragon 8 Gen 3

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论