VoiceStudio logo

VoiceStudio

Open-source, fully-local ElevenLabs alternative for voice cloning, dubbing, and transcription

Alternative to: ElevenLabs

VoiceStudio screenshotVoiceStudio screenshot

About Versions (38)

v0.5.2

2026-09-10

Introduces Studio's Convert method, a synchronized audiobook player, batch queue watch folders, and significant UI refinements for dubbing and engine management across Windows, macOS, and Linux.

v0.5.1

2026-08-28

Version 0.5.1 introduces VoiceStudio as a local speech platform with new API transports, improved dubbing tools, expanded hardware support for Apple Silicon and AMD GPUs, and a simplified one-command installation process.

v0.5.0

2026-08-14

OmniVoice-Studio is renamed to VoiceStudio, introducing a new Model Catalogue, remote GPU sharing via join codes, enhanced server security, and a redesigned Dub workspace.

v0.4.2

2026-07-27

Improved update notifications, prevented data loss during updates, fixed model download repairs, and corrected localization errors across multiple languages.

v0.4.1

2026-07-26

This update fixes AMD GPU acceleration, improves error reporting, adds MCP host allowlist support, and resolves various stability and compatibility issues across Windows and macOS.

v0.4.0

2026-07-21

Introduces comprehensive audiobook tools with multi-voice casting, a designed-voice Gallery, Paste Translation for dubbing, and significant performance and security optimizations.

v0.3.22

2026-07-13

This release significantly improves dubbing quality with stereo audio, better timing, and consistent voice matching, while introducing opt-in analytics and critical memory and performance fixes for 16GB machines.

v0.3.21

2026-07-12

This update fixes major memory leaks and OOM crashes by optimizing model unloading and introduces a comprehensive new Storage settings menu for granular data management.

v0.3.20

2026-07-12

Added an in-app uninstall feature, fixed a race condition causing incorrect backend connection errors, and improved cleanup of log folders on Linux and Windows.

v0.3.19

2026-07-12

This update improves error reporting accuracy, introduces streaming audio playback, adds dedicated uninstallers for all platforms, and optimizes the first-run wizard experience.

v0.3.18

2026-07-11

Introduces one-click installation for IndexTTS-2, automated Hugging Face endpoint detection for restricted networks, and improved settings concurrency and test coverage.

v0.3.17

2026-07-11

This release bundles FFmpeg and yt-dlp, redesigns the Engines and Models pages, fixes a critical dubbing session crash, and resolves numerous settings and scaling issues.

v0.3.16

2026-07-10

Introduces generation takes with restore, per-sentence caching for audiobooks, enhanced dubbing consistency and fit prediction, global text normalization, and a persistent audio mini-player.

v0.3.15

2026-07-09

Fixes cold-start timeouts, prevents updates from deleting manual engine installs, resolves a voice cloning performance regression, and adds new Agent Skills.

v0.3.14

2026-07-09

Adds visible engine pickers for TTS, ASR, and LLM in Settings, fixes Linux AppImage white-screen detection, and adds Windows installation documentation.

v0.3.13

2026-07-08

This release adds OpenAI-compatible transcription, fixes macOS microphone permissions, resolves Linux AppImage white-screening, and improves voice cloning and dubbing stability.

v0.3.12

2026-07-08

Improves engine selection consistency, fixes first-run network and SSL proxy blockers, and resolves multiple crashes and community-reported bugs across Windows, Linux, and macOS.

v0.3.11

2026-07-05

Improves multi-language dubbing workflows with per-language translation and caching, adds self-documenting backend crash reports, and fixes various stability issues across different platforms and network configurations.

v0.3.9

2026-07-04

Rebuilds dictation with live waveforms and faster commits, adds detailed LLM connection diagnostics, introduces per-feature LLM routing, and implements extensive reliability fixes for VRAM and system crashes.

v0.3.10

2026-07-04

This release fixes dialogue hallucinations in Cinematic/Autofit dubbing, resolves audiobook rendering crashes, ensures correct speaker counts, and prevents stale backend versions from persisting after updates.

v0.3.8

2026-07-01

Introduces live local dictation, a user pronunciation dictionary, a redesigned Settings hub, and critical stability fixes for backend connectivity and GPU memory crashes.

v0.3.7

2026-06-20

This stable release improves non-English language correctness, adds cross-platform audio and UI fixes for Linux and Android, introduces two new opt-in TTS engines, and resolves various stability and installation bugs.

v0.3.6

2026-06-16

Introduces a Longform suite for audiobooks and stories, engine-routing to prevent silent CPU fallback, enhanced dubbing tools, and a transition to AGPL-3.0 licensing.

v0.3.5

2026-06-03

Fixed a speaker diarization failure caused by PyTorch 2.6's secure unpickler by adding necessary safe-globals to the allowlist.

v0.3.4

2026-06-03

Fixed transcription failures on Windows with NVIDIA GPUs by introducing PyTorch Whisper as a fallback backend to eliminate cuDNN 8 dependencies.

v0.3.3

2026-06-03

Corrected CPU architecture reporting in Docker/web builds and improved CI checksum portability for macOS runners.

v0.3.2

2026-06-03

Fixed an issue where Docker UI users encountered 'Loopback origin required' errors and blank version information due to restricted admin routes.

v0.3.1

2026-06-02

Version 0.3.1 fixes voice-clone export crashes in Docker/browser builds, corrects version display and updater UI, and improves transcription error reporting.

v0.3.0

2026-06-02

Introduced a frameless OS-level dictation widget, refactored the capture component, updated Docker GPU detection, and overhauled the README.

preview

2026-06-01

Version 0.3.0 introduces Pro Studio for audiobook creation, extensive i18n support for 21 languages, network sharing via LAN/Tailscale, and significant stability and security improvements across Windows, macOS, and Linux.

v0.2.7

2026-05-03

Introduced a frameless global dictation widget, refactored the capture component into a standalone widget, and overhauled the README documentation.

feat/frameless-dictation-widget

2026-05-03

Introduced a new frameless dictation widget.

v0.2.6

2026-04-30

This is an auto-generated release based on the commit log.

v0.2.5

2026-04-30

Introduces system-wide dictation with global hotkeys, an autonomous batch video dubbing pipeline, and dual-mode ASR improvements.

v0.2.4

2026-04-28

Introduces a WebSocket event bus for real-time UI updates, updates the default UI scale, and fixes several cross-platform stability and boot issues.

v0.2.2

2026-04-23

This is an auto-generated release with changes detailed in the commit log.

v0.2.1

2026-04-23

This is an auto-generated release with details available in the commit log.

v0.2.0

2026-04-23

This is an auto-generated release with changes detailed in the commit log.