Desktop dictation, live transcription, and speaker capture for Linux and Windows.
gdictate turns speech into text where you already work. Hold a hotkey, speak, and the app sends recognized text to the focused window. It uses a local desktop shell, a Python daemon, and a browser Web Speech bridge, so there is no Google Cloud setup and no paid API key.
Download latest · Project page · Build workflow
- Two hold-to-talk channels: one bind for your microphone, one bind for speaker output.
- Live popup: compact lower-center transcript overlay with timer and voice level strip.
- Click-through overlay: the popup can stay above apps without stealing clicks or focus.
- Sequential paste: confirmed chunks can be pasted while you dictate, with final fallback on release.
- Compact GUI: one main app screen with a focused settings tab.
- Shared core: GUI, tray, daemon, and CLI use same settings and runtime modules.
- Two live STT engines: Chrome Web Speech gives interim text; ChatGPT bridge or any OpenAI-compatible
/v1/audio/transcriptionsendpoint gives normalized final text. - Speaker capture: Linux auto-detects the active speaker monitor; Windows supports loopback recording endpoints.
- File transcription: optional local audio/video transcription with subtitle and JSON exports.
- Release packages: AppImage, deb, rpm, Arch pacman package, Windows setup exe, and Windows MSI.
Use the latest GitHub Release.
| Platform | Asset to pick |
|---|---|
| Linux portable | AppImage |
| Debian / Ubuntu | deb package |
| Fedora / RPM distros | rpm package |
| Arch / CachyOS / pacman distros | pacman package |
| Windows | setup exe or MSI |
For Arch-style packages:
sudo pacman -U ./gdictate-*-x86_64.pkg.tar.zstAfter install, run gdictate-app from the app launcher or terminal.
Tauri app / tray / native controls
|
v
Python daemon and shared core
|
+--> hotkeys
+--> audio routing
+--> paste backend
+--> live popup events
+--> file transcription jobs
|
v
Hidden Chrome / Chromium / Edge Web Speech page ──> streaming results
or OpenAI-compatible STT endpoint ──> final result
Chrome engine runs local HTTPS/WebSocket service and hidden speech-proxy.html for interim/final events. chatgpt captures current PipeWire/Pulse default source on release, posts WAV to local ChatGPT bridge, then emits same normalized final event. openai uses any compatible endpoint. All engines share hotkeys, audio routing, popup, daemon IPC, and paste path.
ChatGPT Web transcription endpoint is unofficial and may change. Keep its bridge loopback-only. Never expose session credentials or bridge port.
| System | Status |
|---|---|
| Linux + Wayland + GNOME/KDE | Primary target: GUI popup, tray, PipeWire/Pulse routing, clipboard paste, evdev hold hotkeys. |
| Arch / CachyOS | Main package path through the pacman release asset. |
| Debian / Ubuntu | deb package plus source bootstrap path. |
| Fedora / RPM distros | rpm package plus source bootstrap path. |
| Windows | Desktop package, microphone dictation, paste backend, and loopback-based speaker capture. |
| macOS | Not targeted yet. |
Linux notes:
- Hold hotkeys on Wayland need access to input devices. Defaults are dedicated
F8/F9; do not use modifier-only chords such asAlt+Left, which desktop navigation and key-repeat can interfere with. - Add your user to the
inputgroup, log out, then log back in. - Paste uses clipboard plus a key-injection backend such as
ydotoolorwtype. - Speaker mode follows the current default PipeWire/Pulse speaker monitor.
Windows notes:
- Microphone dictation works through the default recording input.
- Speaker transcription needs a loopback input such as Stereo Mix, VB-CABLE, Virtual Audio Cable, or Voicemeeter.
- gdictate reports missing Windows audio setup instead of installing drivers.
Linux:
git clone https://github.com/bigidulka/gdictate.git
cd gdictate
./install.sh
.venv/bin/python gdictate.py --daemon --no-uiDesktop app in development:
npm run bootstrap:linux
npm run tauri:devWindows:
.\scripts\bootstrap-windows.ps1
npm run tauri:devOptional system dependencies:
./install.sh --install-systemOptional local file transcription dependencies:
./install.sh --with-batchDefault controls:
| Action | Default |
|---|---|
| Hold microphone channel | F8 |
| Hold speaker channel | F9 |
| Stop recording | release the held key |
| Open app settings | tray menu or main window |
| Disable popup | Settings -> Live -> Live popup |
Useful daemon commands:
.venv/bin/python gdictate.py --status
.venv/bin/python gdictate.py --start mic
.venv/bin/python gdictate.py --start speakers
.venv/bin/python gdictate.py --stop
.venv/bin/python gdictate.py --shutdownDiagnostics:
.venv/bin/python gdictate.py --capabilities
.venv/bin/python gdictate.py --diagnostics
.venv/bin/python gdictate.py --preflight
.venv/bin/python gdictate.py --live-report
.venv/bin/python gdictate.py --shortcut-report
.venv/bin/python gdictate.py --settings-snapshotThe app stores settings in the user config directory:
| OS | Settings path |
|---|---|
| Linux | ~/.config/gdictate/settings.json |
| Windows | %APPDATA%\gdictate\settings.json |
Main settings groups:
- General: language, engine, default channel.
- Chrome: browser channel, hidden window mode, profile path, setup flow.
- Transcriber: ChatGPT bridge or generic OpenAI-compatible endpoint, model, timeout, optional API-key environment variable.
- Audio: microphone source, speaker source, Linux routing, Windows loopback input.
- Binds: hold or toggle mode, mic bind, speakers bind, Linux hotkey backend.
- Paste: paste backend, terminal paste combo, live paste, clipboard-only mode.
- Live: popup enabled, click-through, interim text, position.
- Files: local transcription model, output formats, diarization options.
Live popup settings apply immediately. Daemon-bound settings are saved and applied through daemon restart when the runtime is idle.
.venv/bin/python gdictate.py --lang en-US
.venv/bin/python gdictate.py --source mic
.venv/bin/python gdictate.py --source speakers
.venv/bin/python gdictate.py --source both
.venv/bin/python gdictate.py --bind-mode dual-hold
.venv/bin/python gdictate.py --bind-mode toggle --key CTRL+ALT
.venv/bin/python gdictate.py --live-paste
.venv/bin/python gdictate.py --no-live-paste
.venv/bin/python gdictate.py --paste auto
.venv/bin/python gdictate.py --paste type --save-settings
.venv/bin/python gdictate.py --paste copy --save-settings
.venv/bin/python gdictate.py --linux-paste-key ctrl-v
.venv/bin/python gdictate.py --chrome-channel chromium
.venv/bin/python gdictate.py --chrome-profile-dir ./tmp/chrome-profile
.venv/bin/python gdictate.py --no-chrome-hidden
.venv/bin/python gdictate.py --setup
# ChatGPT bridge default: http://127.0.0.1:37182/v1/audio/transcriptions
.venv/bin/python gdictate.py --engine chatgpt --test --no-ui
# Any compatible endpoint; optional bearer token is read from named environment variable.
GDICTATE_TRANSCRIBER_API_KEY=... .venv/bin/python gdictate.py --engine openai --transcriber-endpoint http://127.0.0.1:9000/v1/audio/transcriptions --test --no-ui
.venv/bin/python gdictate.py --save-settings
.venv/bin/python gdictate.py --reset-settingsDaemon IPC:
.venv/bin/python gdictate.py --daemon --no-ui
.venv/bin/python gdictate.py --daemon-hotkeys
.venv/bin/python gdictate.py --toggle mic
.venv/bin/python gdictate.py --toggle speakersFile transcription:
.venv/bin/python gdictate.py --file-report ./call.mp4
.venv/bin/python gdictate.py --transcribe-file ./call.mp4 --model-size small --export-format all
.venv/bin/python gdictate.py --transcribe-file ./call.mp4 --diarize --diarization-backend auto
.venv/bin/python gdictate.py --file-start ./call.mp4
.venv/bin/python gdictate.py --file-jobs
.venv/bin/python gdictate.py --file-job <job-id>
.venv/bin/python gdictate.py --file-cancel <job-id>| Path | Role |
|---|---|
gdictate.py |
Compatibility entrypoint. |
gdictate_core/app.py |
Dictation state machine and transcript routing. |
gdictate_core/chrome.py |
Browser Web Speech bridge, local HTTPS server, WebSocket messages. |
gdictate_core/openai_compatible.py |
Capture-to-OpenAI-compatible STT engine; ChatGPT loopback bridge specialization. |
gdictate_core/audio.py |
Linux PipeWire/Pulse routing and Windows speaker-input guidance. |
gdictate_core/paste.py |
Clipboard, direct typing, live paste, and platform paste backends. |
gdictate_core/hotkeys.py |
Linux evdev hold/toggle listeners. |
gdictate_core/ipc.py |
Local daemon HTTP/WebSocket control server. |
gdictate_core/settings.py |
Shared dataclass settings and schema. |
gdictate_core/platforms.py |
Capability reports, diagnostics, safe system actions. |
gdictate_core/file_jobs.py |
Batch ASR/diarization jobs and exports. |
src/ |
React UI. |
src-tauri/ |
Desktop shell, tray, native windows, popup window, daemon supervisor. |
speech-proxy.html |
Browser Web Speech page. |
The daemon listens on 127.0.0.1:9877.
| Endpoint | Role |
|---|---|
GET /status |
Current daemon state and active channel. |
POST /start |
Start mic, speakers, or both. |
POST /stop |
Stop recording and finalize text. |
POST /toggle |
Toggle recording. |
GET /events |
WebSocket stream of engine, recording, transcript, audio level, and file-job events. |
GET /file-jobs |
List file transcription jobs. |
POST /file-jobs |
Start a file transcription job. |
GET /file-jobs/{id} |
Read job state/result. |
POST /file-jobs/{id}/cancel |
Cancel queued/running job. |
POST /shutdown |
Stop daemon. |
npm run test:python
npm run perf
npm run build
cargo check --manifest-path src-tauri/Cargo.tomlLinux packages:
npm run tauri:build:linuxWindows packages:
npm run tauri:build:windowsThe release workflow builds package assets for tags matching v* and uploads them to the matching GitHub Release.
Release checklist:
# Update package manifests first.
git commit -m "Release <tag>"
git tag <tag>
git push origin main <tag>Built assets:
- Linux AppImage
- Linux deb
- Linux rpm
- Arch pacman package
- Windows setup exe
- Windows MSI
MIT