After using the built-in dictation feature on my Mac for some time, I gradually began to feel that it was not quite enough for my needs. I decided to compare several third-party dictation apps and eventually settled on VoiceInk.
In this article, I’ll explain how I arrived at VoiceInk, what I like about it, and the settings worth knowing when you first start using the app.
Note: This is an English version of an article originally written in Japanese. I primarily use VoiceInk to dictate Japanese, so my comments about recognition accuracy and recommended AI models are based on my experience in a Japanese-language environment.
The built-in dictation feature in macOS has some clear advantages. It feels responsive because your words appear on the screen as you speak, and using a feature built into the operating system provides a certain sense of reliability.
However, I was often frustrated by its difficulty in recognizing technical terms correctly. I frequently dictate IT and digital-related terminology, as well as product names, and manually correcting these terms every time was becoming tedious. That led me to explore third-party dictation apps in search of something that better suited my needs.
I tried four apps: MacWhisper, superwhisper, Aqua Voice, and VoiceInk. Of these, VoiceInk was the one I ultimately decided to keep using.
In my experience, its speech recognition is highly accurate, and it handles a wide range of vocabulary, including technical terminology. If VoiceInk fails to transcribe a particular word correctly, you can add it to the dictionary so that it will be entered correctly in the future.
VoiceInk also felt faster than the other apps I tried. Another deciding factor was its one-time purchase price of $29 at the time of writing, rather than a recurring subscription.
VoiceInk lets you choose which AI model to use for speech recognition. At first, I wasn’t sure which one to select, so I tried several. For Japanese dictation, the model I found most accurate was “Large v3 Turbo (Quantized).”
When using this model, very short Japanese phrases are sometimes translated into English for no apparent reason. I found that I could usually avoid this problem by speaking for a little longer and having VoiceInk process a complete sentence or passage at once.
Another option is “Cohere Transcribe,” which, in my experience, can handle short Japanese phrases more reliably. Processing speed and transcription results vary depending on the model, but I found these two to be broadly comparable overall. I recommend trying both and choosing the one that works best for your language, speaking style, and workflow.
I have now been using VoiceInk regularly for more than two weeks, and it has already become an essential part of my workflow. I especially recommend it to freelancers and others who work from home and are able to use voice input as part of their daily routine.
The following steps are intended for readers who have already downloaded VoiceInk. I’ll walk through the initial setup and then introduce a few more advanced settings and ways to use the app. I hope these tips will help anyone who is just getting started.
First, choose an AI model. Select [AI Models] in the sidebar and download the model you want to use. For Japanese dictation, my recommendation is “Large v3 Turbo (Quantized).” Quantization is a technique used to reduce the size of an AI model, so this version requires less storage than the standard “Large v3 Turbo” model. If you plan to dictate in another language, the model that works best may be different, so it is worth testing several options.After the download is complete, assign the model to a profile. Select [Models] from the sidebar. You should see a profile called “Default.”Click “Default” to open the side panel. Under [Transcription], open the [Model] menu and select the “Large v3 Turbo (Quantized)” model you downloaded earlier.The [Language] setting below it can be left on [Auto-detect], but I find that specifying the language reduces recognition errors. For Japanese dictation, select [Japanese]. Be careful not to choose [Javanese], which appears immediately below it in the list. Javanese is a different language spoken primarily on the Indonesian island of Java. After choosing the language, click [Save Change] to close the panel.Next, open [Setting] and configure the keyboard shortcut used to activate VoiceInk. You can assign a shortcut under [Primary Shortcut]. VoiceInk can distinguish between the left and right [control] and [option] keys, giving you more flexibility when choosing a shortcut that does not interfere with other apps.The drop-down menu next to the shortcut controls how the key operates. [Toggle] starts recording when you press the shortcut once and stops it when you press the shortcut again. [Push to Talk] records only while you hold down the shortcut and stops when you release it. [Hybrid] supports both types of operation. If you use VoiceInk frequently or want to dictate longer passages at once, I think [Toggle] is the most convenient option. These are the only essential settings you need to configure before you begin.Close the settings window and try using VoiceInk. In any app where you can enter text, press the keyboard shortcut you configured earlier. An indicator will appear near the bottom center of the screen. Speak into your Mac’s microphone or your preferred external microphone.After a short wait, the transcribed text will be inserted at the cursor position. When using “Large v3 Turbo (Quantized)” for Japanese, very short phrases may not be recognized as Japanese even when [Language] is set to [Japanese]. I have found that dictating a longer, complete sentence at once generally produces better results. This behavior may differ when dictating in other languages or using a different AI model.
Advanced Settings and Usage Tips
One of the first additional features worth learning is the dictionary. Open [Dictionary] in the sidebar and register any words or phrases that you want VoiceInk to transcribe correctly. At the top of the [Word Replacements] tab, enter the pronunciation or commonly misrecognized version in the field on the left. In the field on the right, enter the exact text you want VoiceInk to produce, then press the [return] key to save the entry. You can also register words under the [Vocabulary] tab. For Japanese, however, I have found [Word Replacements] to be more reliable. Japanese can represent words using kanji, hiragana, katakana, and the Latin alphabet, so specifying both the input and the desired replacement gives me more predictable results. For other languages, the [Vocabulary] feature may work perfectly well, so it is worth trying both options.The next feature to check is History. VoiceInk saves your previous transcriptions, including their audio recordings. If these files are left untouched, they can gradually consume a significant amount of storage on your Mac, so I recommend configuring VoiceInk to delete old audio files automatically.Click the gear button to open the side panel, then turn on [Auto-delete Audio Files]. I have set [Keep Audio For], which controls how long recordings are retained, to [1 day].VoiceInk can also mute or pause audio that is playing when you begin dictating. I often listen to music in Apple’s Music app while I work, so I find this feature particularly useful. Open the audio settings using the [Audio] button, then turn on both [Mute Audio While Recording] and [Pause Media While Recording].I also recommend setting up more than one AI model. VoiceInk lets you assign a different keyboard shortcut to each model. I normally use “Large v3 Turbo (Quantized).” If it does not transcribe something correctly, I use a separate shortcut to activate “Cohere Transcribe” instead. This makes it easy to switch models without repeatedly opening the settings.I have also changed the clipboard settings. VoiceInk appears to insert transcribed text by temporarily placing it on the clipboard and then pasting it into the active app. As far as I can tell, the process works roughly as follows: (1) VoiceInk temporarily stores whatever is already on the clipboard; (2) it records your speech; (3) it places the transcription on the clipboard; (4) it automatically pastes the transcription at the cursor position; and (5) it restores the clipboard’s previous contents. By default, there is a two-second delay between steps 4 and 5. I have shortened this delay to [500ms], which restores the previous clipboard contents more quickly.