Native macOS voice studio for Apple Silicon. Design a voice, direct its emotion, and voice your Gloam.fm AI DJ — every sample rendered locally with MLX.
Drop a reference clip and clone a voice in seconds. Save, edit and export as portable .gvoice packs.
Five emotion variants — flat, neutral, warm, excited, hype — plus 0.5×–2.0× playback speed control.
Generate two takes side-by-side and compare waveforms and playback before you commit.
Every take is saved with backend, voice, emotion and RTF. Play, delete, or one-click reuse into the editor.
Reference clips auto-transcribe. Dictate anywhere, or convert any audio file with ⇧⌘T — Apple or Whisper, always local.
Download weights with live progress and an automatic storage preflight. No terminal required.
Swap engines per take. From real-time lightweight cloning to premium research-grade quality and multilingual voice design — all downloaded and run on your Mac.
Fast, lightweight real-time cloning. The everyday workhorse for quick iterations and live-feel synthesis.
The faster sibling of Chatterbox — RTF ~1–2×. Trades Chatterbox's exaggeration knob for a quicker, more balanced take.
Premium research-grade quality. Per-session temperature control for expressive, nuanced delivery.
Multilingual and commercial-friendly. Clone at two sizes or direct a preset speaker — or invent a voice from a text prompt in Voice Design.
Same speaker, same line — steered from flat to hype. Tap an emotion to hear how the delivery changes.
No clip to clone from — just write what you want and Qwen3-design conjures a speaker to match, entirely on-device.
A built-in text-model catalog runs Gemma-4 and Qwen3 entirely on your Mac — the engine that will script DJ banter, station IDs and show intros. Pick a tier to match your RAM; heavier weights, sharper writing.
A multi-line editor where each line carries its own voice and emotion direction. Reorder freely, generate the whole session in one click, and export a single stitched WAV.
Design and clone the hosts, MCs and DJs that drive your Gloam.fm station. Export them as .gvoice packs and drive them live through the local API — the same interchange format the Python engine speaks.
Flip on the built-in HTTP server and drive synthesis from any script. Loopback-only, port 8790, minimal sandbox entitlements — nothing exposed to the network.
No sign-in, no sync, no telemetry. Just open the app and go.
Minimal entitlements; data lives in Application Support and Caches only.
The optional server binds to localhost only — never the open network.
Voices and takes are plain files you own — .gvoice, WAV, JSON.
Grab the app or build from source. macOS 14+, Apple Silicon. Pick a model and it downloads in-app.
Record or drop a short clip. It auto-transcribes on-device; an optional transcript hint sharpens quality.
Type your line, choose emotion and speed, and synthesize. Compare A/B takes and export.
Free and open source. Nothing to sign up for — download and generate your first voice in minutes.
Gloam Voice Studio is made by Tiny Trash Labs on a stack of remarkable open projects. Thanks to everyone who builds and maintains them.