42,256 stars · AGPL-3.0, read from /blob/main/LICENSE on 2026-09-28; the standard unmodified text, with section 13 the material obligation. SEPARATE AND STRICTER: the project's own README says the default OmniVoice pretrained weights are labelled CC-BY-NC, so the model forbids commercial use even though the code does not · v0.5.6 (2026-09-23), read from /releases/latest and confirmed by ungh to the second · Track this in Scout
Voice cloning, voice design, video dubbing, dictation, transcription and audiobook production, all running on the reader's own machine with no upload.
▶Repo detailsthe review · specs · pros & cons · install
What it is
A desktop program and a local server that bundle several speech models behind one screen. It covers voice cloning, voice design, video dubbing, dictation, transcription and long-form audiobook production, and it claims 646 languages.
What it is good for. Anyone who records narration and is paying a per-minute fee to a cloud voice service. It removes the fee and the upload. It also suits people who cannot send recordings to a third party at all, because the audio never leaves the machine. Anyone recording in a language a cloud service handles badly may get a better result here, because the model can be swapped.
- No account, no key and no per-minute bill. The cost is electricity and the machine.
- One screen covers cloning, dubbing, dictation and audiobook output, so there is no chain of four separate tools to wire together.
- Analytics is off by default, and when it is switched on the project states it sends no text, no audio, no filenames and no voices.
- The licence on the code and the licence on the voice are different, and the second one is stricter. The code is AGPL-3.0, read from the LICENSE file. But the project's own README says the default pretrained weights are labelled CC-BY-NC, which means no commercial use. Selling anything made with the default voice means finding other weights first.
- It wants real hardware. The project asks for 8 GB of memory as a minimum and recommends 16 GB or more, plus 10 GB of free disk as a minimum and 20 GB on an SSD as the recommendation. A graphics card is optional, with 4 GB of video memory as the floor and 8 GB recommended.
- Narrow platform support. On macOS the local part runs on Apple Silicon only, and Linux needs glibc 2.39 or newer, which rules out older systems. The version is 0.5.6, so it is well before 1.0.
- coqui-ai/TTS
45,998 stars, the best-known toolkit for the same job, and its last code landed on 16 August 2024, over two years ago, with no archived banner to warn anyone.
Track this in Scout - fishaudio/fish-speech
32,725 stars, code on 16 September 2026; a strong speech model with voice cloning, but no dubbing, dictation or audiobook workflow around it.
Track this in Scout - myshell-ai/OpenVoice
36,911 stars, last code 19 April 2025; a cloning model on its own, with no screen and no pipeline, and quiet for about seventeen months.
Track this in Scout
# the project's own installer, macOS or Linux curl -fsSL https://voicestudio.sh/install | sh # or Docker (a way to run a program inside its own sealed box, # so it cannot break anything else on the machine) docker run -d -p 127.0.0.1:3900:3900 \ -v omnivoice-data:/app/omnivoice_data \ --name voicestudio palashdeb/omnivoice-studio:stable




