Gepard : 0.6B streaming TTS built for real-time dialogue - 20× realtime factor, ~50ms time-to-first-audio, vLLM-native, Apache 2.0
We just open-sourced Gepard 1.0 , a TTS model built for real-time conversation. It’s streaming-first: audio starts the moment text arrives, generated frame by frame instead of waiting for a full sentence. - ~555M params : Qwen3.5 0.8B backbone (14 layers) + Nemo NanoCodec (FSQ, 22.05kHz) - ~20 x RTF , ~50ms TTFA on one RTX 5090 via vLLM - Up to 256 parallel sequences on a single RTX Pro 6000 Balckwell with 96GB VRAM - Zero-shot voice cloning from a few seconds of reference - Languages: English (US/UK), Span
评论
?
参与讨论