FireRedAudio & FireRedTTS3 by FireRedTeam - Huggingface

FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation HuggingFace : huggingface.co/FireRedTeam/FireRedAudio GitHub : github.com/FireRedTeam/FireRedAudio Demo : fireredteam.github.io/demos/fireredaudio Overview FireRedAudio is a general-purpose audio language model built on a shared 9B-parameter LLM with decoupled continuous representations : an Audio Encoder handles understanding, while a RedAE pathway handles generation. A single model supports ASR, audio understanding, zero-shot TTS, instruct TTS, semantic/acoustic speech editing, and accurate temporal grounding over recordings up to one hour long . Highlights ✨ 🧩 Purpose-built representations, one shared backbone — The Audio Encoder pathway serves understanding, while the RedAE-Patch pathway serves speech generation. Their representations remain decoupled but share the same language and reasoning backbone. To the best of our knowledge, this is the first publicly disclosed design of its kind in a unified audio-language model. 📊 One model, a full audio stack — FireRedAudio spans ASR, broad and fine-grained audio understanding, zero-shot TTS, Instruct TTS, and free-form speech editing, achieving competitive or leading results across MMAU, MMSU, Seed-TTS-Eval, InstructTTSEval, and Ming-Freeform-Audio-Edit. 🎙️ Create and edit speech with natural language — Clone a voice from a reference clip, design a voice from a description, or edit what was said and how it sounds through one continuous-latent generation pathway. ⏱️ Go from minutes to hour-long recordings — Understand recordings up to one hour with precise time-to-content alignment. Organize audio into timestamped structures, produce grounded summaries, retrieve content by time (or time by content), and reason over evidence distributed across the recording. FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations HuggingFace : huggingface.co/FireRedTeam/FireRedTTS3 GitHub : github.com/FireRedTeam/FireRedTTS3 Demo : fireredteam.github.io/demos/firered_tts_3 arXiv : arxiv.org/abs/2608.17492 Overview FireRedTTS3 is a unified speech generation and editing system built on semantically enriched continuous speech representations . It comes in two variants: FireRedTTS3-Base — zero-shot voice cloning across 24 languages and 21 Chinese dialects FireRedTTS3-Instruct — natural-language voice design and speech editing (semantic + acoustic) in one unified model Highlights ✨ 🌍 Multilingual — 24 Languages — Best average WER/CER (avg 3.754%) and best average speaker similarity on MiniMax-MLS-Test (avg 84.8%), plus best-in-class cloning WER/CER (avg 3.04%) and similarity on Seed-TTS-eval (avg 78.8%). Supported languages: Arabic · Cantonese · Chinese · Czech · Dutch · English · Finnish · French · German · Greek · Hindi · Indonesian · Italian · Japanese · Korean · Polish · Portuguese · Romanian · Russian · Spanish · Thai · Turkish · Ukrainian · Vietnamese 🗣️ Multi-Dialect — 21 Chinese Dialects — Zero-shot voice cloning across major Chinese dialect groups. Supported dialects: Anhui · Fujian · Gansu · Guizhou · Hebei · Henan · Hubei · Hunan · Jiangxi · Liaoning · Minnan · Ningxia · Shaanxi · Shandong · Shanghai · Shanxi · Sichuan · Tianjin · Wenzhou · Wu · Yunnan 🎨 Instruction-Controlled Voice Design — Generate a brand-new voice from a natural-language description (gender, age, timbre, emotion, pace, accent…) with no reference audio, guided by an explicit textual plainning step before synthesis. ✂️ Free-Form Speech Editing — Semantic editing (insertion / deletion / substitution) and acoustic editing (speed / pitch / volume) driven by free-form instructions. Project : fireredteam.github.io Their Opensource Projects : Really worth to check the project page. Nice Ecosystem. They also published some papers. OpenStoryline : An Agentic Framework for Autonomous, Human-Aligned Video Creation FireRedChat : A Fully Self-Hosted Solution for Full-Duplex Voice Interaction IVC-Prune : Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning FireRedTTS-2 : Towards Long Conversational Speech Generation for Podcast and Chatbot InstanceAssemble : Layout-Aware Image Generation via Instance Assembling Attention InstantID : Zero-shot Identity-Preserving Generation in Seconds DynamicPose : A Robust Image-to-Video Framework for Portrait Animation Driven by Pose Sequences PhotoPoster : A High-Fidelity Two-Stage Pose-Driven Image Generation Framework CQ-DINO : Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection FireRedASR : Open-Source Industrial-Grade Automatic Speech Recognition Models FireRedTTS-1S : An Upgraded Streamable Foundation Text-to-Speech System The Xiaohongshu Speech Synthesis System for Blizzard Challenge 2023 StoryMaker : Towards Consistent Characters in Text-to-Image Generation LayerDiffuse-Flux InstantStyle : Free Lunch towards Style-Preserving in Text-to-Image Generation

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论