Speak into any app on Linux. Your words appear as text.
One shortcut to dictate anywhere • Any model from a growing library • An open protocol any app can build on • Built in Rust
Super.STT.Quick.Demo.mp4
Super STT makes speech-to-text with automatic text input trivial on Linux, for everyone. Bind a shortcut (Super+Space by convention), speak, and your words are typed straight into whatever app is focused, e.g. your editor, browser, chat, terminal. No copy-paste, no fiddling.
Under the hood it's two things:
- A model-agnostic engine. A background daemon installs speech models from a library of backends (local or cloud), loads one, and keeps it warm for instant transcription. You pick the model that fits your hardware and swap it whenever you like.
- An open protocol. The daemon speaks a documented HTTP protocol over a local socket, so any app, in any language, can request transcriptions, stream live audio visualizations, or drive recording, with per-app consent. Super STT's own desktop app, CLI, and COSMIC applet are just the first clients. See Developers.
Quick install - detects your system and downloads pre-built binaries:
curl -sSL https://raw.githubusercontent.com/jorge-menjivar/super-stt/main/install.sh | bashAppend -s -- --beta for the latest beta.
Build from source - for the very latest changes (needs the build prerequisites):
git clone https://github.com/jorge-menjivar/super-stt.git
cd super-stt
just installEither way you get the daemon, the stt CLI, a consent helper, the desktop app, and (on COSMIC) the panel applet, wired up as a systemctl --user service. Everything installs system-wide (root-owned, so the installer asks for sudo), while the daemon itself runs unprivileged in your user session. On COSMIC the installer offers to bind Super+Space → stt record --write for you. GPU acceleration comes from the model you run (see Models).
stt record --write # Record and transcribes. Types the result after silence is detected or you run the command again.Bind stt record --write to a key combo. Super+Space is the convention (the COSMIC installer does this for you; on other desktops add a custom shortcut for stt record --write). When you trigger it:
sttasks the running daemon to start transcribing from your mic.- You speak; the daemon transcribes your audio.
- When you stop (on silence or a second trigger), it types the transcription into the focused app.
Manage the daemon with the usual systemd controls:
systemctl --user start super-stt # or: enable / status / restart
journalctl --user -u super-stt -f # follow logsTwo settings shape a session. Both live in the desktop app under Settings (or the daemon config); stop mode can also be set per-recording on the CLI.
Stop mode - how a recording ends:
| Mode | Behavior |
|---|---|
| Silence + Manual (default) | Stops on silence detection or a second shortcut press |
| Silence Only | Stops only when silence is detected |
| Manual Only | Stops only on a second shortcut press |
stt record --write --stop-mode manual_onlyWrite method - how text is injected: Auto (default) tries the XDG Desktop Portal, then ydotool, then direct Wayland input. Force a specific one in Settings if auto-detection picks wrong.
Models come from a library of backends you install on demand. Open the app, go to Library → Browse, install a backend, and it appears in the model selector. Some run locally (your audio never leaves your machine); others are online providers you reach with your own API, which is stored securely in your system keyring (GNOME Keyring, KWallet, …).
Local - everything stays on your device.
| Model | Best for |
|---|---|
| Voxtral (mini / small) | High accuracy; needs an NVIDIA GPU (CUDA). |
| Qwen3-ASR (0.6b / 1.7b) | Fast, multilingual; runs on CPU or an NVIDIA GPU (CUDA). |
| Whisper (tiny → large) | Versatile and battle-tested; tiny/base are great CPU defaults. |
Online - bring your own API key.
| Provider | Notable models |
|---|---|
| Mistral | voxtral-mini-latest, plus a realtime Voxtral model |
| OpenAI | gpt-4o-transcribe, gpt-4o-mini-transcribe, whisper-1 |
| Deepgram | nova-3 |
GPU acceleration is a property of the model, not a separate build of the app. Install a GPU-capable backend (like Voxtral or Qwen3-ASR) and the daemon downloads the build matched to your NVIDIA GPU automatically. You just need an up-to-date driver.
This is a snapshot. The app always shows the current catalog, published live at jorge-menjivar.github.io/super-stt/index.json. Want a model that isn't there? Anyone can publish one.
Models & backends — load a backend, choose a model, and pick where it runs.
Clean start![]() |
Load a backend![]() |
Pick a model![]() |
Run on CPU or GPU![]() |
Select language![]() |
Library — browse the catalog and manage what's installed.
Installed backends![]() |
Browse the catalog![]() |
Settings — tune audio feedback, recording behavior, and text input.
Customization![]() |
Recording![]() |
Input simulation![]() |
Panel visualizer — the COSMIC applet shows your mic input live in the panel, in three styles.
stt: command not found: binaries are installed system-wide and are onPATHby default — restart your terminal so it rehashes, and check the installer finished without errors.- Daemon won't start / misbehaves: check
journalctl --user -u super-stt -n 50. - Transcriptions seem less accurate than they should be: your microphone input volume may be set too high or too low. Adjust the mic volume in your system sound settings and try again.
- Typing doesn't work in some apps:
ydotooltypes reliably across virtually all apps. Install it via your package manager, then try it out first by runningsudo ydotoold --socket-path="$HOME/.ydotool_socket" --socket-own="$(id -u):$(id -g)"in a terminal and setting the write method to ydotool in Settings. If that fixes typing, make it permanent by enabling the service (e.g.systemctl --user enable --now ydotool).
Super STT is built to be built on. The details live in developer-facing docs:
- Build a client — get transcriptions, event streams, or recording control into your own app, in any language, over the documented HTTP protocol → docs/protocol/
- Add your own model — package a speech model as a backend the daemon can install and run, then publish it to the catalog → docs/protocol/ and registry/README.md
- Contribute — build from source, workspace layout, and the PR workflow → CONTRIBUTING.md
Architecture and the security model live in docs/.
Jorge Menjivar • jorge@menjivar.ai












