Verbatim v2 update. "self-hosted transcription that turns a whole channel into transcripts + AI analysis (Whisper/Gemini, Docker)"

preview.redd.it/edgmcyv5yojh1.png This is a tool that allows for mass-transcription of audio/video into transcripts, get a fully interesting analyses to somebody you like (Youtube-channel and other platforms). Still using and developing. the v2 differs: 1. increase model-selection section in the "Setting" bar 2. compression sets for the audio after downloading ensures a less stroage taken 3. Using Metal(from MacOS) to speed up whisper process(MLX / Apple Silicon GPU backend for Whisper) 4. Adding classification tag in the "Library" for users to find the sources that are uploaded by themselves or through the "Pipelining" process. 5. Finishing "Rednotes"(Xiaohongshu in Chinese) content pipelining, trigger scraping from the UI, stream progress live, stop a running job, then use multimodal analysis to turn scraped notes into a report. 6. increase reliability of the transcribe. (implement backward strategy when several source meet content block. 7. Adding language selector 8. modification toward several processes that are important to the cost, key exposure protections, XSS fixes... It is a personal project, glad to hear some comments and feedbacks!

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论