VoxSherpa TTS
Studio-quality offline neural text-to-speech for Android.
Hindi ยท English ยท British ยท Japanese ยท Chinese ยท and more โ No cloud. No limits. No compromise.
๐ก๏ธ Privacy-focused user? Please check our Documentation before downloading.
๐ Featured In
VoxSherpa TTS is listed in the official README of k2-fsa/sherpa-onnx โ the core inference library powering this app.
Why VoxSherpa?
Most TTS apps make you choose between quality and privacy. Cloud-based tools like ElevenLabs sound incredible โ but they require internet, send your text to remote servers, and charge per character.
VoxSherpa breaks that tradeoff.
It runs two professional-grade neural engines entirely on your device:
| Engine | Quality | Speed | Best For |
|---|---|---|---|
| ๐ง Kokoro-82M | Studio-grade ยท rivals ElevenLabs | Slower on budget hardware | Audiobooks, voiceovers, professional content |
| โก Piper / VITS | Natural ยท clear ยท multi-speaker | Fast on any device | Daily use, dialogue synthesis, quick synthesis |
Screenshots
| Generate | Models | Library | Settings |
|---|---|---|---|
![]() | ![]() | ![]() | ![]() |
Features
๐๏ธ Dual Neural Engine
- Kokoro-82M โ 82 million parameter neural model. Multilingual support including Hindi, English, British English, French, Spanish, Chinese, Japanese and 50+ languages. Same architecture used by top-tier commercial TTS services.
- Piper / VITS โ Fast, lightweight, natural. Generates speech in seconds on any Android device.
๐ฃ๏ธ Multi-Speaker Dialogue Synthesis (New in v4.0)
- Piper now supports multiple speakers within a single script โ generate full conversations in one pass instead of stitching clips together
- New
[speaker]tag system to assign each line to a distinct voice directly inside your script:[speaker:1] Hello, how are you? [speaker:2] I'm good, thanks! How about you? - Voice Style & Tone controls โ fine-tune how each speaker sounds to match the mood of the script
- Adjustable sentence gap / silence timing โ control the pause duration between lines for natural, realistic pacing
- Ideal for audiobooks with multiple characters, podcasts, skits, and narrated dialogue
๐ 100% Offline & Private
- All processing happens on your device
- No internet required after model download
- No account, no telemetry, no data collection
- Your text never leaves your phone
๐ Document to Audio
- PDF to Audio โ listen to any document hands-free
- TXT to Audio โ convert plain text files instantly
- Share any text directly to VoxSherpa from any app
๐ฆ Model Management
- Download models directly from the app
- Filter voice models by language or type
- Sample voice preview before selecting a model
- Import your own
.onnxmodels from local storage - Multiple models installed simultaneously
- Smart storage tracking
- Optional MMS model support โ enable via a toggle in Settings if needed
๐ System-Wide TTS
- Set VoxSherpa as your default Android TTS engine
- All downloaded models exposed to System TTS โ use any voice in Chrome, WhatsApp, TalkBack, and more
- Pitch & speed control in System TTS mode
- Sample voice preview for all models
๐ง Audio Controls
- Real-time waveform visualization
- Adjustable speed and pitch
- Interactive audio seeking with mini player controls
- MediaStyle notification with full playback controls
- Export as WAV with correct sample rate per model
๐ Speech Library
- Save all generated audio locally
- Favorites system for quick access
- View generation history with timestamps
- Voice model attribution per recording
- Regenerate audio on voice change
โ๏ธ Smart Settings
- Smart Punctuation โ natural pauses after sentence breaks
- Emotion Tags โ
[whisper],[angry],[happy]support - Per-model voice selection (Kokoro supports 50+ speakers)
- Theme-aware UI
Technical Architecture
User Text
โ
โโโโ Kokoro Engine (KokoroEngine.java)
โ โโโ Sherpa-ONNX JNI โ ONNX Runtime โ CPU/NNAPI
โ โโโ kokoro-multi-lang-v1_0
โ
โโโโ Piper / VITS Engine (VoiceEngine.java)
โโโ Sherpa-ONNX JNI โ ONNX Runtime โ CPU
โโโ VITS model (language-specific)
Built with:
- Sherpa-ONNX โ on-device neural inference
- Kokoro-82M โ multilingual neural TTS model
- Piper โ fast local TTS, now with multi-speaker dialogue support
- Android AudioTrack API โ low-latency PCM playback
Performance
Generation speed depends entirely on your device's processor:
| Device Tier | Kokoro | Piper |
|---|---|---|
| ๐ข Flagship (Snapdragon 8 Gen 3) | ~20โ40 sec/min audio | ~5 sec/min audio |
| ๐ก Mid-range (8-core) | ~60โ90 sec/min audio | ~10 sec/min audio |
| ๐ด Budget (6-core) | ~2โ3 min/min audio | ~20 sec/min audio |
Kokoro prioritizes quality over speed by design. It uses the same 82M parameter architecture that powers premium commercial TTS โ running it entirely offline on a mobile CPU is genuinely pushing the hardware limits.
Installation
Requirements: Android 11+ ยท ARM64 ยท ~500 MB free storage recommended (for models)
Model Import (Technical Users)
VoxSherpa supports importing custom .onnx models without any server:
- Place your
.onnxmodel +tokens.txton device storage - Open Models tab โ tap + โ Import Local Model
- Select your files
Compatible with any Sherpa-ONNX compatible TTS model.
Contributing
VoxSherpa is open source. Contributions welcome:
- ๐ Bug reports via Issues
- ๐ก Feature requests via Discussions
- ๐ง Pull requests for fixes and improvements
License
Copyright (C) 2025 CodeBySonu95
This program is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version.
This program is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details.
https://www.gnu.org/licenses/gpl-3.0.html
Acknowledgements
- k2-fsa/sherpa-onnx โ the inference engine that makes this possible
- hexgrad/Kokoro-82M โ the neural model behind studio-quality synthesis
- rhasspy/piper โ fast local TTS engine, now with multi-speaker dialogue support
Built with obsession. Runs without internet.
VoxSherpa โ Because your voice deserves to stay yours.



