Everything runs
on your Mac.
No subscriptions for the AI. No data leaving your network. No cloud inference. Every model downloads once; after that, it's yours.
YouTube import
Paste any YouTube link. OkeDoke downloads the video and audio using yt-dlp in the best available quality. Age-restricted content works when you supply a cookies file. Audio extracts automatically via ffmpeg.
yt-dlp · cookies-aware · mp4 + mp3 · ffmpeg extraction
[YT] Connecting to YouTube...
[YT] Format: bestvideo[avc1]+bestaudio[m4a]
[YT] ⏳ Downloading... 67%
[YT] Extracting audio...
[YT] ✅ Downloaded: ยังคงคอย
Vocal separation
BS-Roformer separates the lead vocal from the instrumental using a state-of-the-art roformer architecture — SDR 12.97. Mel-Band Roformer then isolates the lead vocal from backing harmonies. Both models run on Apple Silicon GPU via CoreML. No cloud, no upload.
BS-Roformer SDR 12.97 · Mel-Band Roformer SDR 10.19 · CoreML GPU
[UVR] BS-Roformer loading...
[UVR] CoreML GPU ✅
[UVR] 47% ████████░░░░
[UVR] 100% ████████████
[UVR] ✅ vocals.mp3 + instrumental.mp3
[BVS] Mel-Band separating harmonies...
[BVS] ✅ lead.mp3 + backing.mp3
Multi-stem separation
htdemucs 6-stem separates the instrumental into individual tracks — drums, bass, guitar, piano, and other. Useful for practice (play along to just the rhythm section), remixing, or turning contest mode into a band karaoke.
htdemucs_6s · drums / bass / guitar / piano / other · CPU background
[Demucs] htdemucs_6s loading...
[Demucs] CPU background processing
[Demucs] drums ✅
[Demucs] bass ✅
[Demucs] guitar ✅
[Demucs] piano ✅
[Demucs] ✅ 6 stems complete
Synchronized lyrics
Whisper large-v3 runs via Apple Neural Engine (MLX), transcribing the vocal track in chunks — never sending silent audio to the model, which prevents hallucination during bridges and instrumental breaks. WhisperX force-aligns every word to ±30ms. Thai songs use Thonburian Whisper. Paste your own reference lyrics and OkeDoke aligns them instead.
Whisper large-v3 MLX · WhisperX · wav2vec2-th · Thonburian Whisper
[Transcribe] is_thai=True
[VAD] 8 vocal chunks extracted
[Chunk 1/8] 0–24s...
[Chunk 8/8] 192–216s...
[Realign] WhisperX ±30ms
[Realign] 64 updated, 3 kept
✅ 71 lines aligned
AI karaoke coaching
Enable Contest Mode and OkeDoke listens through your mic, scoring Pitch, Timing, and Expression in real time. A live arc fills as you sing. Three AI coach personas react mid-song to what you just did. After the song, Gemma 4 generates a personalized 2-sentence critique per coach, then OmniVoice speaks it in their own voice.
Gemma 4 2B Q4 GGUF · llama.cpp Metal GPU · OmniVoice MPS · ScoringEngine
[Scoring] mic input 44.1kHz
[Score] pitch=84 timing=71 expr=68
[Coach] warm: reaction → good_pitch
→ เสียงสวยมากค่ะ 🎵
[Gemma] Generating coach comments...
[OmniVoice] Speaking: โบกี้ ไทเกอร์
LAN server
OkeDoke starts a FastAPI server the moment the app opens. It registers itself on your local network via Bonjour so every device finds it without an IP address. The REST API serves the full song library, audio/video streaming, and real-time playback control. Server-Sent Events push state to all connected screens simultaneously.
FastAPI + uvicorn · Bonjour/mDNS :58423 · SSE events · REST API
✓ OkeDoke Pro started
[API] Bonjour: 192.168.1.42:58423
→ Apple TV connected
→ iPhone connected
→ iPad connected
[Queue] Now playing:
ยังคงคอย — Hers
Legacy library
Import your existing .dat, .mpg, .avi, and .wmv karaoke disc files. OkeDoke scans the folder, reads the vocal channel configuration, and splits left/right channels on demand via ffmpeg. No re-ripping, no conversion upfront — files are processed as needed.
.dat · .mpg · .avi · .wmv · ffmpeg channel split · on-demand
[Legacy] Scanning /Volumes/KARAOKE/
[Legacy] 847 files found
[Legacy] vocal_channel=RIGHT
[Legacy] prepare → ffmpeg split
[Legacy] ✅ ready
Models downloaded on first launch
Gemma 4 2B Q4
Coach comments · song intelligence
OmniVoice
TTS · coach voice synthesis
BS-Roformer
Lead vocal separation
Mel-Band Roformer
Backing vocal isolation
htdemucs 6-stem
Multi-stem instrument separation
Thonburian Whisper (ct2)
Thai lyrics transcription
Whisper large-v3 MLX
Fast transcription via Neural Engine
Total: ~10 GB. Downloaded once, stored in ~/Music/OkeDokePro/models/. No internet required after that.