Skip to content

VoltipHold a key, speak, and let go.

The text appears at your cursor, in any app. Recognition runs on your computer or with a cloud service you choose, and AI polish can tidy it first.

Runs on
Windows, macOS (Apple silicon and Intel) and Linux. An Android app is in development.
Your audio
With a local model it never leaves your computer. With a cloud service it goes only to the service you chose.
The Voltip home page with the shortcut, the microphone, the speech model and the latest results.The Voltip home page with the shortcut, the microphone, the speech model and the latest results.
The overlay while you speak, showing the input level and the Esc key that cancels.The overlay while you speak, showing the input level and the Esc key that cancels.
The overlay during AI polish, showing the time the step has taken.The overlay during AI polish, showing the time the step has taken.
The overlay after the text has been inserted, showing the number of characters and the app.The overlay after the text has been inserted, showing the number of characters and the app.

What Voltip does

Every feature at a glance. Features marked "In development" are being built and are not in a release yet.

Dictation wherever you type

  • Hold Ctrl+Alt+Space, speak and release. You can also press once to start and again to stop, or lock a long take with a short press.

  • Right Ctrl, Right Alt, the Fn key on a Mac or a mouse side button can start dictation on their own.

  • The overlay shows the words as they are recognised. The text can be inserted all at once, sentence by sentence, or as you speak.

  • Voice editAvailable

    Select text, hold Ctrl+Alt+E and say what to change, such as "make it more formal" or "translate into English". The selection is replaced in place.

Recognition on your terms

  • Qwen3-ASR, SenseVoice and Paraformer run on your computer. Once a model is downloaded, no audio leaves it.

  • Qwen3-ASR uses the graphics card through Vulkan on Windows and Linux and Metal on macOS, and the processor otherwise.

  • Release packages include a default service. OpenAI, Groq, SiliconFlow or any OpenAI-compatible endpoint can be used with your own key.

Text that reads right

  • AI polishAvailable

    Fixes punctuation, typos and filler words without changing what you said. Works with the default service, your own provider or a local Ollama.

  • Teach Voltip your names and terms, and rewrite phrases with literal or regular-expression rules. Chinese can be written in Simplified or Traditional characters.

  • Scenes by appAvailable

    The app you speak into can choose the polish style, the output mode, the language and extra instructions for the AI.

  • Proofread, prompt, intent, chat, translation and notes presets, switched from the home page, the title bar or the tray, plus ready-made scenes for coding, office writing, chat and professional fields.

Phone, history and long recordings

One dictation, four steps

  1. Hold the shortcut

    CtrlAltSpace

    The microphone opens only now. The overlay at the bottom of the screen shows the input level.

  2. Speak

    With the live transcription model installed, the words appear in the overlay as you say them.

  3. Release the keys

    The speech is recognised, corrected with your dictionary, polished by AI if you want, and rewritten by your rules.

  4. The text is at your cursor

    Voltip pastes the text into the app you were using and keeps a copy in the history. Esc cancels at any point.

Recognition on your computer, or in the cloud ​

Download a model once and dictation works offline. Qwen3-ASR recognises 30 languages and writes the punctuation itself; SenseVoice and Paraformer are small and fast on any processor. When the computer has a graphics card, Qwen3-ASR runs on it through Vulkan or Metal, and on the processor otherwise.

Release packages also include a default cloud service, so dictation works before you set anything up. You can switch to OpenAI, Groq, SiliconFlow or any OpenAI-compatible endpoint with your own key at any time.

Local recognition · Cloud services

ModelSizeLanguagesRuns on
Qwen3-ASR 0.6BRecommended690 MB30 languages, detected automaticallyGPU or CPU
Qwen3-ASR 1.7BMost accurate1.7 GB30 languages, detected automaticallyGPU or CPU
SenseVoice SmallSmallest download240 MBChinese, English, Japanese, Korean, CantoneseCPU
ParaformerNo punctuation227 MBChinese, including dialects, mixed with EnglishCPU
Live transcriptionPreview while you speak169 MBChinese and EnglishCPU

Downloads resume after an interruption and are checked against a SHA-256 hash before use. On an NVIDIA L40S, Qwen3-ASR 0.6B recognised 16 seconds of Chinese speech in 0.16 s on the GPU and in 2.87 s on 8 processor threads.

Text that reads the way you meant it ​

Recognition gets the words; the next steps make them read right. The dictionary corrects the names and terms a recogniser tends to mishear. AI polish removes filler words and fixes the punctuation. Replacement rules rewrite phrases exactly as you specify, and scenes adjust all of this to the app you are writing in.

Presets are in development In development: choose between proofreading, prompt writing, intent, chat, translation and notes from the home page, the title bar or the tray.

AI polish · Dictionary and rules · Scenes

Recognised

um so I think we should uh ship the fix on friday and then update the docs next week

Inserted at your cursor

I think we should ship the fix on Friday and then update the docs next week.

An example of AI polish. Filler words are removed and the punctuation is fixed; the wording stays yours.

Your phone as a microphone and keyboard ​

In development

Pair an Android phone with the computer and hold to talk on the phone. The computer recognises the speech with its own settings and inserts the text at its cursor. The phone can also send typed text or its clipboard.

The Android app is built and tested but not released yet. It will be on the releases page once it is. How the phone works

Pair once
Scan the QR code on the computer, enter its 6-digit code, or tap the computer in the phone's list of computers nearby. Both screens then show the same safety code for you to confirm.
Hold to talk on the phone
The audio streams to the computer, compressed with Opus and encrypted end to end. The computer recognises it and inserts the text at its own cursor.
Send text or the clipboard
Type on the phone or send its clipboard. The computer inserts it like a dictation result.
On any network
On the same network the devices connect directly. Elsewhere, an optional relay forwards the encrypted data without being able to read it.

Platforms

Voltip is the same app on all three desktop systems. The platform notes list what differs between them.

PlatformPackagesShortcut and single keyLocal recognitionGPU
Windows 10 and 11AvailablePer-user installer or portable zip, x64BothYesVulkan
macOS 11 or laterAvailableA dmg for Apple silicon and one for IntelBoth, with the Accessibility permissionYesMetal
LinuxAvailable.deb or AppImage, x64; X11 and WaylandBoth on X11; on Wayland, a system shortcut runs the toggle commandYesVulkan
AndroidIn developmentNot released yetHold to talk in the appDone by the paired computer—

iOS has not been started. Packages for Windows on ARM and Linux on ARM are not planned for now.

What leaves your computer

It depends on the services you choose. The home page and the Speech models page always show which one is in use.

Local recognition

SendsNothing

Recognition runs inside Voltip on a model file on your disk. The only network use is downloading the model, which is verified against a SHA-256 hash.

Cloud recognition

SendsAudio, to the service you chose

The recording goes to the built-in service or to the provider you set up, with your key and, if you keep a dictionary, its terms as a hint. Keys are kept in the system keychain.

AI polish and voice edit

SendsText, to the AI service you chose

The recognised text, or for voice edit the selection and your instruction. The app's name is included unless you turn it off; the window title only if you turn it on. Audio is never sent.

Install ​

One command picks the right package for the computer, checks it against the release's SHA256SUMS and installs it.

powershell
irm https://raw.githubusercontent.com/sunerpy/voltip/main/scripts/install.ps1 | iex
bash
curl -fsSL https://raw.githubusercontent.com/sunerpy/voltip/main/scripts/install.sh | sh

Packages for every platform, checksums and build attestations are on the releases page. The install guide covers each platform, including the first start on a Mac.

What comes next

The next releases, in the expected order. No dates are promised; each feature is described on this site once it ships.

  1. AI presets and built-in scenes

    In development

    The default Proofread preset fixes typos and punctuation, removes filler words, and keeps only the corrected wording when you correct yourself. Further presets cover prompts, intent, chat, translation between Chinese and English, and notes; you can also write your own.

  2. 20,000 history entries and statistics

    In development

    The history limit rises from 500 to 20,000 entries, and the home page shows words dictated, words corrected and time saved by day, week and month. Time saved is speaking time × 1.9, based on Ruan et al. 2016.

  3. Long recordings and computer audio

    In development

    Recordings of up to two hours from the microphone, the computer's own audio or both. Long recordings are recognised in segments and can be exported as SRT subtitles.

  4. The Android app

    In development

    Using the phone as a microphone and keyboard is built and tested. The app is being prepared for release.

  5. Streaming recognition with cloud services

    Planned

    Text while you speak currently uses the local live transcription model. Streaming results from cloud services are planned.

  6. More local models

    Planned

    Whisper among others, and measurements on AMD and Intel graphics cards.

Deliberately left out

  • Listening all the time, or stopping when you fall silent. The microphone opens only while you record.
  • Speaker separation, meeting detection and batch transcription of files.
  • A browser-only version. The global shortcut, typing into other apps and local models need a native app.
  • Syncing the history, dictionary or rules to the phone. The phone is a microphone and keyboard; the computer does the work.

Voltip is released under the Apache License 2.0. Part of FirLab. Content from voltip@8197772.