Literature Review: LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load | Bnechmarking LLMs on Phones [R]
Just finished reading the paper: LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load I am starting to benchmark LLMs on edge devices, particularly phones thus been reading a lot on the what has been done and what is currently being done and wanted to share you my journey of reading such papers and my takes on them. This is one of the only papers I have read that have benchmarked RPi5-Hailo (Hailo's 10H) iPhone 16 Pro (A19 Pro) S24 Ultra (Snapdragon 8 Gen 3
评论
?
参与讨论