@iamcheyan: https://github.com/iamcheyan/omarchy-voxtype-enhance… omarchy自带语音输入法,按下F9就能使用。但是中文支持很差(毕竟原作者也不使用中文)。 做了一个linux原生的语音输入法,…
摘要
Voxtype Enhance is a plugin for the Omarchy platform that adds native voice input with improved Chinese support, using offline AI models like SenseVoice and Paraformer for local CPU inference.
查看缓存全文
缓存时间: 2026/08/23 03:34
https://github.com/iamcheyan/omarchy-voxtype-enhance… omarchy自带语音输入法,按下F9就能使用。但是中文支持很差(毕竟原作者也不使用中文)。 做了一个linux原生的语音输入法,安装后选择一款模型,插件会自动下载,按住F9键直接说话,对于中文和英文包括像linux windows这种词也能识别,并自动粘贴。 纯本地CPU推理,不消耗token。
iamcheyan/omarchy-voxtype-enhance
Source: https://github.com/iamcheyan/omarchy-voxtype-enhance
Voxtype Enhance
Voxtype Enhance is an Omarchy experience-enhancement plugin for Voxtype voice input. It adds a native top-bar microphone control and a compact settings panel for speech models, language, and output behavior.
Screenshot

English: The settings panel is opened from the microphone icon in the top bar. It shows the available speech models, language choices, and output modes.
中文: 点击顶栏麦克风图标即可打开输入法设置面板。面板提供语音模型、识别语言和输出方式设置。
日本語: トップバーのマイクアイコンから音声入力設定パネルを開けます。音声モデル、認識言語、出力方式を設定できます。
Features
- Native Omarchy top-bar microphone indicator.
- Three Sasayaki-compatible offline speech models, with engine selection handled automatically.
- Automatic model download, installation, file-size checking, and SHA-256 verification.
- Chinese, English, Japanese, Korean, and automatic language choices.
- Omarchy universal paste:
Shift+Insertin terminal/TUI windows andCtrl+Vin graphical applications. - Voxtype’s native Type output as an alternative.
- Download progress shown directly inside the selected model card.
- Clear plugin data action for resetting model downloads and plugin-managed settings.
- No replacement recorder, speech recognizer, clipboard daemon, or extra Quickshell service.
Installation
Install Voxtype first through Omarchy’s AI / Dictation setup, then install the plugin:
omarchy plugin add https://github.com/iamcheyan/omarchy-voxtype-enhance.git --enable
The plugin provides a bar-widget entry point. If the microphone icon is not automatically placed in the top bar, add Voxtype Enhance to the right side of the bar layout.
The plugin does not register or replace Voxtype’s recording shortcut. Use the shortcut already configured by Voxtype, for example holding HOME to record and releasing it to transcribe.
Models
The panel shows model names rather than engine names. Selecting a model downloads missing files, verifies them, installs them into Voxtype’s model directory, and then switches the engine:
| Panel option | Internal engine | Voxtype model | Size |
|---|---|---|---|
| SenseVoice Small · int8 | sensevoice | small-int8 | 229 MB |
| SenseVoice Small · full precision | sensevoice | small | 894 MB |
| Paraformer Large · int8 | paraformer | paraformer-zh | 232 MB |
Only the model confirmed as active in Voxtype is highlighted. A downloaded model
that is not active is labelled downloaded, not active; a failed engine switch
shows the Voxtype/systemd error and restores the effective selection. Downloads
use the pinned Sasayaki Hugging Face sources and SHA-256 manifests. Already
verified files are skipped.
SenseVoice and Paraformer require an ONNX-capable Voxtype binary. The default
Voxtype installation uses the standard Whisper binary. When an ONNX model is
selected without ONNX support, the plugin stops before downloading and shows an
explicit Enable ONNX support (administrator approval) action. Only clicking
that action runs the fixed command pkexec voxtype setup onnx --enable; after
approval, select the model again to download and activate it. Cancelling the
authentication leaves the existing Whisper setup unchanged.
If Voxtype itself is missing (for example when only the plugin was installed),
any panel action that changes settings first offers to install Omarchy’s
voxtype-bin package through the fixed command
pkexec pacman -S --noconfirm --needed voxtype-bin. Cancelling the
authentication prompt leaves the system unchanged; read-only status queries
never trigger an installation.
Models are stored under:
~/.local/share/voxtype/models/
Output modes
The plugin intentionally exposes two output modes:
Omarchy universal paste
Voxtype writes the complete transcription to the Wayland clipboard. The plugin then detects the focused Hyprland window using the same terminal-tag policy as Omarchy’s clipboard.lua:
- Omarchy terminal/TUI windows receive
Shift+Insert; - graphical applications receive
Ctrl+V.
This avoids terminal applications interpreting Ctrl+V as a special action, such as an image-paste command. The final shortcut is sent through Hyprland’s hl.dsp.send_key_state dispatcher with explicit key press and release events.
Before output, the plugin records the previous clipboard hash. It sends the shortcut only when Voxtype has produced new non-empty text, so pressing the recording key without speaking does not paste stale clipboard content.
Type
Voxtype’s native simulated keyboard-input mode. This can be useful in applications where direct typing is preferred, but mixed-language text may interact with Fcitx5/Rime character by character.
Settings and reset
The panel updates ~/.config/voxtype/config.toml only after the user selects a setting. Changes restart the existing user-level voxtype.service so Voxtype loads the new engine or output configuration.
The Clear plugin data link removes only the three model directories managed by this plugin and restores SenseVoice int8, Chinese, and Omarchy universal paste. It does not remove Whisper models, recordings, hotkeys, or unrelated Voxtype data. After clearing, choose a model to download it again.
The plugin does not edit Hyprland bindings, change the recording hotkey, or
install a system service. When Voxtype is already installed, model selection
does not require privileges. Exactly two operations elevate through pkexec;
both are visible, user-initiated, and run fixed commands:
pkexec voxtype setup onnx --enable— switch to the ONNX-capable binary;pkexec pacman -S --noconfirm --needed voxtype-bin— offered only when the Voxtype binary is missing entirely.
Dependencies
The following are provided by Omarchy, Voxtype, or the normal desktop setup:
- Omarchy Quickshell;
- Voxtype and its user service;
- Python 3 standard library;
- Hyprland and
hyprctlfor universal paste; wl-copyfor clipboard output;- Voxtype’s normal output dependencies, such as
wtype.
Model downloads require network access only when a selected model is not already installed.
Uninstallation
Disable or remove the plugin using Omarchy’s plugin manager. Removing the plugin removes the top-bar control and plugin files. It does not remove Voxtype, user recordings, or downloaded models; use Clear plugin data before removal if those model files should also be deleted.
Validation
From the plugin repository root:
omarchy plugin validate .
/usr/lib/qt6/bin/qmllint -I /usr/share/omarchy/shell \
bar/widget.qml VoxtypePanel.qml
python3 -m py_compile scripts/*.py
python3 -m unittest discover -s tests -v
License
MIT. See LICENSE.
中文说明
Voxtype Enhance 是一个 Omarchy 体验增强插件,用于增强 Voxtype 语音输入。它在顶栏提供麦克风图标和设置面板,让用户可以管理语音模型、识别语言和输出方式。
主要功能:
- 顶栏麦克风状态图标;
- 三个 Sasayaki 兼容的离线语音模型;
- 选择模型后自动下载、校验、安装并切换 engine;
- 中文、英文、日文、韩文和自动识别;
- 终端使用
Shift+Insert、普通应用使用Ctrl+V的万能粘贴; - Voxtype 原生 Type 输入模式;
- 模型下载进度显示;
- 清除模型并恢复插件默认设置。
安装 Voxtype 后执行:
omarchy plugin add https://github.com/iamcheyan/omarchy-voxtype-enhance.git --enable
点击顶栏麦克风图标打开设置面板。用户不需要了解模型对应的 engine,选择模型后插件会自动处理。插件不会修改录音快捷键,仍然使用 Voxtype 原来的快捷键。
万能粘贴会根据当前窗口判断输出方式:Omarchy 标记为终端的窗口使用 Shift+Insert,其它图形应用使用 Ctrl+V,以避免 CodeX 等终端程序把 Ctrl+V 解释成图片粘贴操作。
点击面板底部的 Clear plugin data 会删除插件管理的三个模型,并恢复 SenseVoice int8、中文和万能粘贴。不会删除 Whisper 模型、录音、快捷键或其它 Voxtype 数据。
日本語
Voxtype Enhance は、Voxtype の音声入力体験を Omarchy 向けに強化するプラグインです。トップバーのマイクアイコンから、音声モデル、認識言語、出力方式を設定できます。
主な機能:
- トップバーのマイク状態表示;
- Sasayaki 互換の 3 つのオフライン音声モデル;
- モデル選択時の自動ダウンロード、検証、インストール、engine 切り替え;
- 中国語、英語、日本語、韓国語、自動認識;
- ターミナルでは
Shift+Insert、GUI アプリではCtrl+Vを使う Omarchy universal paste; - Voxtype 標準の Type 入力;
- モデルのダウンロード進捗表示;
- ダウンロード済みモデルと設定を初期化するリセット機能。
Voxtype を先に Omarchy からインストールし、次のコマンドでプラグインを追加します。
omarchy plugin add https://github.com/iamcheyan/omarchy-voxtype-enhance.git --enable
トップバーのマイクアイコンをクリックして設定パネルを開きます。モデルの engine を意識する必要はありません。モデルを選ぶと、必要なファイルのダウンロードと設定変更が自動的に行われます。録音ショートカットは変更せず、Voxtype の既存設定を使用します。
Universal paste はフォーカス中のウィンドウを判定し、Omarchy のターミナルでは Shift+Insert、通常の GUI アプリでは Ctrl+V を送信します。これにより、CodeX などの TUI が Ctrl+V を画像貼り付けとして扱う問題を避けられます。
パネル下部の Clear plugin data をクリックすると、プラグインが管理する 3 つのモデルを削除し、SenseVoice int8、中国語、Universal paste の初期設定に戻します。Whisper モデル、録音データ、ショートカット、その他の Voxtype 設定は削除しません。
相似文章
@GitTrend0x: 卧槽兄弟们 本地跑语音克隆+电影级视频配音,直接支持646种语言,完全离线、无API密钥、无需联网,ElevenLabs直接被干翻 https://github.com/debpalash/OmniVoice-Studio… 这波开源神器…
OmniVoice Studio is an open-source desktop app that enables local voice cloning and cinematic video dubbing across 646 languages, fully offline with no API keys, positioning itself as a privacy-focused alternative to ElevenLabs.
@FakeMaidenMaker: 炸裂!这个开源项目免费文字转无 AI 味人声,还能克隆任何人的嗓音,并且用文字调整音色! GitHub 狂揽 30K star,出自面壁智能 OpenBMB,VoxCPM 之前拿过 GitHub 和 HuggingFace 双榜第一。 做…
VoxCPM2是OpenBMB开源的语音合成模型,采用无分词器的扩散自回归架构,支持30种语言、语音设计和可控语音克隆,仅需一句话即可克隆音色,或用文字创建全新声音,输出48kHz高质量音频,可商用。
@denziideng: 不安装、不注册,开源离线免费用,Windows离线语音输入值得拥有! 打字慢手酸?语音输入又怕隐私泄露,还得联网? 分享一个 CapsWriter-Offline,GitHub 6.2k的Windows离线语音输入神器,按住CapsLoc…
CapsWriter-Offline 是一款Windows平台的完全离线语音输入工具,按住CapsLock键说话即可自动识别上屏,支持多种模型(SenseVoice、Paraformer等),无需联网和安装,高度自定义且低延迟。
k2-fsa/OmniVoice
OmniVoice 是一款大规模多语言零样本文本转语音模型,支持超过 600 种语言,基于扩散语言模型架构构建,具备快速推理和语音克隆能力。
@Gorden_Sun: ZONOS2:开源MoE TTS模型 8B总参数,0.9B激活参数。支持多语言,支持语音克隆,支持中文,中文效果不错。 模型:
Zyphra released ZONOS2, an open-source MoE text-to-speech model trained on over 6 million hours of multilingual speech, supporting voice cloning and high-quality synthesis across many languages.