dsh-voice-scribe
PensiveFei
DSH voice input plugin: tap Alt to talk, get text in composer. Web Speech default (zero config), optional OpenAI-compatible ASR, polish via DSH LLM.
GITHUB TOPIC
21projects include this topic
The "asr" topic on GitHub groups 21 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by dsh-voice-scribe with 32 GitHub stars. dsh-voice-scribe — DSH voice input plugin: tap Alt to talk, get text in composer. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.
{count} projects
Exact GitHub Topic match
PensiveFei
DSH voice input plugin: tap Alt to talk, get text in composer. Web Speech default (zero config), optional OpenAI-compatible ASR, polish via DSH LLM.
WizisCool
Voice input plugin for DeepSeek Harness (DSH) with text polishing. Supports local GPU Whisper and STT(ASR) API | 一款支持润色整理的 DeepSeek Harness 语音输入插件,为 DSH 提供语音输入能力,支持本地 GPU Whisper 和主流语音识别转写 API
haoku123
Full-duplex voice plugin for DeepSeek Harness: mic → SenseVoice ASR (sherpa-onnx) → LLM → Edge TTS with true barge-in. Full-duplex voice plugin: microphone voice input, SenseVoice local recognition (Simplified Chinese + punctuation + ITN), streaming TTS playback, supports voice interruption. Zero API key.
Nothree-code
DeepSeek Harness (dsh web) voice input plugin: integrates VocoType local offline recognition, results auto-inserted into the chat input box (auto-deploy/anti-duplication/continuous input)
PiyotaHu
Local-first full-duplex voice for DeepSeek Harness, orchestrated by Muxiva
qishuilalala
DSH full-duplex voice conversation mode: streaming zipformer2 recognition into an editable draft, optional wake word, Edge TTS sentence-by-sentence reading + real-time subtitles, barge-in on speech, no API Key needed. Full-duplex voice mode for DeepSeek Harness, no API key.
ShuiHan268
Qwen-MM-Plugins integration plugin for DeepSeek Harness: 12 multimodal MCP tools (vision/OCR/grounding/ASR/audio-video), Web settings page (paste Qwen API Key and go), built-in skills and one-click installer
Zachary7456
DeepSeek Harness (dsh) voice input plugin: record with mic button/hotkey, real-time transcription backfilled into the input box. Three recognition engines: browser Web Speech, local SenseVoice/Paraformer offline backend (one-click deploy), OpenAI-compatible cloud ASR API.
agent-mobile
Speech plugin for the DeepSeek Harness (dsh) web host: ASR, TTS and realtime transcription over pluggable providers
ai-yucheng
Audio Copilot for DeepSeek Harness — transcribe audio (ASR) , with an in-composer voice-input mic button. 🎤
huangdejie
Speech plugin for DeepSeek Harness: per-message voice playback, auto-announce, and dictation via mic — cloud TTS/ASR on Alibaba DashScope or Volcengine (freely combinable), with Web Speech API fallback.
moluyao
DeepSeek Harness global plugin for MiniMax speech recognition (asr-1.0): transcribe_audio tool + settings card + voice input in the input box
zfu691531-hash
Lightweight realtime ASR → DeepSeek Harness → TTS community plugin.
bitterSmilezzz
Speak and it's written · Speak-to-prompt for DeepSeek Harness: cloud ASR speech recognition + Prompt optimization + fill into draft/auto send, cross-platform macOS / Windows.
eddiehuang227-source
Animate any photo into a responsive virtual girl. She talks, turns, smiles, and moves naturally in sync with your conversation. Low-lag, high-detail.
Gammonmush803
Hold-to-talk voice input for DeepSeek Harness Web, runs locally with SenseVoice via sherpa-onnx, offline, no API key.
liyixuan201211
让 AI Agent 听懂音频:转写、鸟种识别、语音情感、环境声与音乐分析。全部本地运行,不用大模型。DeepSeek Harness 技能。
rudyz666
Parse Bilibili video links and extract the full script/subtitles: subtitle track first, local whisper transcription when there is none, export SRT/TXT/JSON. DeepSeek Harness plugin, cross-platform Windows/macOS/Linux (pure Node).
Vorpal-poem
DeepSeek Harness voice input plugin: Volcengine streaming ASR (Doubao Seed ASR), streams back into the input box. Voice input for DSH via Volcengine streaming ASR.
wuwangmao
DSH bundle: Qwen multimodal bridge — vision (qwen3-vl), speech-to-text (qwen3-asr), text-to-image (qwen-image), for DeepSeek Harness
xichow0663-gif
An agent skill that batch-researches a YouTube channel's latest N videos — captions first, local ASR fallback — and writes one cross-video report instead of N summaries. Works in DSH, Codex, Claude Code, Hermes and WorkBuddy.