Skip to content

Set Up Voice (STT and TTS)

Voice lets you speak to an assistant and hear its replies. Two pieces work together: speech-to-text for what you say, and text-to-speech for what the assistant says back. Each uses a voice provider you connect first.

Voice dictation in the chat composer works without any voice-specific setup. Speech-to-text reuses a provider you have already connected (for example, an OpenAI key drives Whisper transcription), and a bundled local Whisper model is the always-available offline floor, so the mic works even with no cloud key at all. On macOS, the signed build asks for microphone access the first time you tap the mic, not at launch.

Local voice models (Whisper for speech-to-text, Kokoro for text-to-speech) download on demand with live progress, so you can run voice fully offline.

Voice runs on the voice providers in the catalog. Speech-to-text uses providers like Flux Voice, Deepgram, and AssemblyAI; text-to-speech uses providers like ElevenLabs. Flux Voice arrives as a speech-to-text provider with clear, actionable messages, so you can talk to Wayland. Connect the ones you want in Settings > Models by pasting their API keys. See Add a Cloud Model Provider.

Go to Settings > Voice. The panel has a speech-to-text section and a text-to-speech section.

  1. In the speech-to-text section, choose the provider that transcribes your speech.
  2. In the text-to-speech section, choose the provider and voice for spoken replies.

If a section needs a provider you have not connected yet, the panel points you to it.

Use the microphone check in the panel to confirm Wayland can hear your input device before you rely on it in a conversation.

With voice set up, a speech input control appears in the chat box. Press it to speak instead of type. When text-to-speech is on, replies are read aloud.