Everything runs
on your Mac.

No subscriptions for the AI. No data leaving your network. No cloud inference. Every model downloads once; after that, it's yours.

YouTube import

Paste any YouTube link. OkeDoke downloads the video and audio using yt-dlp in the best available quality. Age-restricted content works when you supply a cookies file. Audio extracts automatically via ffmpeg.

yt-dlp · cookies-aware · mp4 + mp3 · ffmpeg extraction

[YT] Connecting to YouTube...

[YT] Format: bestvideo[avc1]+bestaudio[m4a]

[YT] ⏳ Downloading... 67%

[YT] Extracting audio...

[YT] ✅ Downloaded: ยังคงคอย

Vocal separation

BS-Roformer separates the lead vocal from the instrumental using a state-of-the-art roformer architecture — SDR 12.97. Mel-Band Roformer then isolates the lead vocal from backing harmonies. Both models run on Apple Silicon GPU via CoreML. No cloud, no upload.

BS-Roformer SDR 12.97 · Mel-Band Roformer SDR 10.19 · CoreML GPU

[UVR] BS-Roformer loading...

[UVR] CoreML GPU ✅

[UVR] 47% ████████░░░░

[UVR] 100% ████████████

[UVR] ✅ vocals.mp3 + instrumental.mp3


[BVS] Mel-Band separating harmonies...

[BVS] ✅ lead.mp3 + backing.mp3

Multi-stem separation

htdemucs 6-stem separates the instrumental into individual tracks — drums, bass, guitar, piano, and other. Useful for practice (play along to just the rhythm section), remixing, or turning contest mode into a band karaoke.

htdemucs_6s · drums / bass / guitar / piano / other · CPU background

[Demucs] htdemucs_6s loading...

[Demucs] CPU background processing

[Demucs] drums ✅

[Demucs] bass ✅

[Demucs] guitar ✅

[Demucs] piano ✅

[Demucs] ✅ 6 stems complete

Synchronized lyrics

Whisper large-v3 runs via Apple Neural Engine (MLX), transcribing the vocal track in chunks — never sending silent audio to the model, which prevents hallucination during bridges and instrumental breaks. WhisperX force-aligns every word to ±30ms. Thai songs use Thonburian Whisper. Paste your own reference lyrics and OkeDoke aligns them instead.

Whisper large-v3 MLX · WhisperX · wav2vec2-th · Thonburian Whisper

[Transcribe] is_thai=True

[VAD] 8 vocal chunks extracted

[Chunk 1/8] 0–24s...

[Chunk 8/8] 192–216s...

[Realign] WhisperX ±30ms

[Realign] 64 updated, 3 kept

✅ 71 lines aligned

AI karaoke coaching

Enable Contest Mode and OkeDoke listens through your mic, scoring Pitch, Timing, and Expression in real time. A live arc fills as you sing. Three AI coach personas react mid-song to what you just did. After the song, Gemma 4 generates a personalized 2-sentence critique per coach, then OmniVoice speaks it in their own voice.

Gemma 4 2B Q4 GGUF · llama.cpp Metal GPU · OmniVoice MPS · ScoringEngine

[Scoring] mic input 44.1kHz

[Score] pitch=84 timing=71 expr=68

[Coach] warm: reaction → good_pitch

→ เสียงสวยมากค่ะ 🎵


[Gemma] Generating coach comments...

[OmniVoice] Speaking: โบกี้ ไทเกอร์

LAN server

OkeDoke starts a FastAPI server the moment the app opens. It registers itself on your local network via Bonjour so every device finds it without an IP address. The REST API serves the full song library, audio/video streaming, and real-time playback control. Server-Sent Events push state to all connected screens simultaneously.

FastAPI + uvicorn · Bonjour/mDNS :58423 · SSE events · REST API

✓ OkeDoke Pro started

[API] Bonjour: 192.168.1.42:58423

→ Apple TV connected

→ iPhone connected

→ iPad connected


[Queue] Now playing:

ยังคงคอย — Hers

Legacy library

Import your existing .dat, .mpg, .avi, and .wmv karaoke disc files. OkeDoke scans the folder, reads the vocal channel configuration, and splits left/right channels on demand via ffmpeg. No re-ripping, no conversion upfront — files are processed as needed.

.dat · .mpg · .avi · .wmv · ffmpeg channel split · on-demand

[Legacy] Scanning /Volumes/KARAOKE/

[Legacy] 847 files found

[Legacy] vocal_channel=RIGHT

[Legacy] prepare → ffmpeg split

[Legacy] ✅ ready

Models downloaded on first launch

Gemma 4 2B Q4

Coach comments · song intelligence

2.9 GB

OmniVoice

TTS · coach voice synthesis

~2 GB

BS-Roformer

Lead vocal separation

610 MB

Mel-Band Roformer

Backing vocal isolation

871 MB

htdemucs 6-stem

Multi-stem instrument separation

52 MB

Thonburian Whisper (ct2)

Thai lyrics transcription

~3 GB

Whisper large-v3 MLX

Fast transcription via Neural Engine

auto

Total: ~10 GB. Downloaded once, stored in ~/Music/OkeDokePro/models/. No internet required after that.