I built a transcription app that keeps your recordings on your iPhone and still handles speaker separation
I’m a solo iOS developer, and when I started building LoroNote , I had one thing I didn’t want to compromise on: transcription quality. The easy options were obvious. Send the recording to a cloud transcription API. Or use a smaller Whisper model so it’s easier to run on an iPhone. I didn’t want either. I wanted Whisper Large V3 Turbo running directly on the device. That meant dealing with memory limits, heat, battery usage, performance differences between iPhone models, and long recordings. Eventually I ended up building my own Core ML inference engine and optimizing the pipeline around Apple silicon. Now the transcription itself runs fully on-device. And I pushed the same idea further with speaker diarization . So LoroNote can not only transcribe the recording locally, but also separate speakers locally without sending the audio to a server. The current flow is basically: Record → Large V3 Turbo transcription → Speaker diarization → Review All on the iPhone. A few other things I’ve added around that: Background transcription Audio file import Apple Watch recording iPad support Tags and organization Apple Intelligence summaries The Apple Intelligence summary is separate from the transcription pipeline. The actual transcription and speaker separation are handled on-device by LoroNote. A lot of the recent work has been less about adding flashy features and more about making the hard parts reliable: long recordings, different devices, real-world background noise, speaker changes, memory usage, and edge cases users actually send me. That’s probably the part I’m most proud of. It would have been much easier to call an API and move on. But having the whole transcription pipeline under my control means I can keep optimizing it for the hardware instead of paying for every minute of audio or depending on a server. LoroNote works on iPhone, iPad, and Apple Watch. I’m still improving the engine and diarization with every update. If you use transcription apps, I’m curious: Would you rather have the best possible model running locally, or a cloud model if it gave you slightly better accuracy? LoroNote on the App Store