Github

代码库

FireRedTTS3: Multilingual and Multi-Dialect Voice Cloning with Instruction-Guided Voice Design and Speech Editing
Python
A Fully Self-Hosted Solution for Full-Duplex Voice Interaction
Python
Long-form streaming TTS system for multi-speaker dialogue generation
Python
FireRed-OpenStoryline is an AI video editing agent that transforms manual editing into intention-driven directing through natural language interaction, LLM-powered planning, and precise tool orchestration. It facilitates transparent, human-in-the-loop creation with reusable Style Skills for consistent, professional storytelling.
Python
agentchatbotlangchainmcpskillsvideovideo-cutvideo-editingvideo-editing-tools
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.
Python
asrautomatic-speech-recognitionconformerindustrial-gradellmmultimodal-llmopen-sourcespeech-recognitionspeechllmtransformer
FireRed-Image-Edit is a powerful image editing foundation model achieving open-source state-of-the-art performance with precise instruction following, high-fidelity generation, superior identity consistency, and seamless multi-element fusion.
Python
aigcdeep-learningdiffusion-modelsimage-generationimage2imagepytorch
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singing ASR. FireRedVAD supports speech/singing/music in 100+ langs. FireRedLID supports 100+ langs and 20+ zh dialects. FireRedPunc supports zh and en.
Python
asrasr-pipelineaudio-event-classificationaudio-event-detectionautomatic-speech-recognitionindustrial-gradelanguage-identificationlidllmmultimodal-llmopen-sourcepunctuation-predictionpunctuation-restorationsotaspeech-recognitionspeechllmvadvoice-activity-detection