dsh-vision-toolkit
Anionex
让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.
GITHUB TOPIC
43个项目包含此标签
43 个项目
GitHub Topic 精确匹配
Anionex
让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.
linenxi-ctrl
为 DeepSeek Harness 增加外挂识图模型:圆形鲸鱼按钮、发送图片识图自动回传、模型自主截图+识图工具、多协议自动适配、小白一键安装(未装 Node.js 自动下载)
Flyvhidbwo
DeepSeek Harness 插件:DeepSeek 大脑 + 自动识图。GUI 附加图片自动经 OpenAI 兼容 VLM 转译成文字后交给 DeepSeek 作答;支持百炼/智谱/OpenRouter 等任意 OpenAI 兼容端点(默认 qwen3.7-flash),无 key 自动探测本地 Ollama(图片不出本机);安装时有一问式确认
Sqhao-O
Fully local document intelligence for DeepSeek Harness. Parse PDF, Office files, images, and scanned documents with offline OCR. | DeepSeek Harness 全本地文档智能插件,支持 PDF、Office、图片与离线 OCR
maxwell-feng
该仓库暂未提供项目说明。
Aidenwu0209
PaddleOCR skills for DeepSeek Harness with native tools and GUI configuration
Favio8
DeepEye vision plugin for DeepSeek Harness (DSH): image description, OCR, VQA, UI layout, and clipboard analysis.
GOU-GEE
该仓库暂未提供项目说明。
jing-hy
DSH plugin: pixel-to-text image reading for text-only models. image_scan/image_ocr/image_sample tools + image-reading skill (34-image trained methodology). Pure local, optional PaddleOCR.
Sorwcyra
Paste images into DeepSeek Harness with a four-model vision race, OCR, and an automatic text bridge.
Argonaut790
Image understanding, OCR, and persistent visual evidence for text-only DeepSeek Harness models
ferstar
本地 OCR 插件:让纯文本生成 LLM 也能读懂图片 | Local OCR plugin: give text-only generative LLMs the ability to read images
linkingoscar
Codex-style attachment formats for the DeepSeek Harness Web GUI: PDF text-layer extraction, Office text extraction, scanned-PDF OCR, long-document spill + index cards, image-to-PNG.
niyongsheng
Local‑only vision skill for macOS 本地化识图技能
zouyuanqing
Native interactive visual-reasoning plugin for DeepSeek Harness: precise pixel grounding (SOM grid / zoom / annotate / measure / diff / color / OCR) + MiMo V2.5 multimodal backend, zero external MCP servers.
honghudavy-star
DSH 自建插件集合:微信桥接器 + GUI 微信入口补丁,一键安装
maxwell-feng
该仓库暂未提供项目说明。
tdf1995
Vision for text-only LLMs in DeepSeek Harness (DSH): describe images / OCR / VQA via free Gemini & GLM vision APIs
uknowmyface
Local OCR for DeepSeek Harness — read text from screenshots on your Mac with Apple's Vision framework. No API key, no upload.
Xieweikang123
Give a text-only dsh model eyes: pasted images recognized into text via an OpenAI-compatible vision endpoint.
Aidenwu0209
Unlimited-OCR for DeepSeek Harness with a native tool and GUI configuration
ByronLeeeee
Matter-aware legal workspace dashboard and document agent tools for DeepSeek Harness
DDDFXYqiming
Vision skill plugin for DeepSeek Harness (image analysis and OCR)
go-farther-and-farther
DeepSeek Harness (DSH) 本地视觉眼睛插件:screen 工具(截图/图片交给本地视觉模型描述)+ ocr 工具(Windows 内置 OCR 逐字提取文字)。零云端、OCR 零 GPU、图片不出本机。
hawkongz
让纯文本模型通过桌面豆包看见聊天图片的 DeepSeek Harness 宿主插件(CDP 桥接,全预设生效,识别可取消)
jmjmj009gt
Zero-dependency vision OCR/Q&A toolkit (CLI + local web GUI) for OpenAI-compatible VLMs: Zhipu GLM, Qwen, OpenAI, OpenRouter, SiliconFlow
Kevoyuan
On-device macOS OCR and Apple Vision for DeepSeek Harness — one native plugin with a bundled Skill.
Koreyer
A DeepSeek Harness tool plugin that lets text-only agents "see" local images — auto-detects the real format and returns a detailed text description via any OpenAI-compatible vision model.
Leeminjing
Give text-only DeepSeek models on-demand vision: upload images, DeepSeek answers by calling a view_image tool backed by any OpenAI-compatible vision endpoint (Qwen/DashScope by default).
princefrogdida-ux
Windows-first vision suite with image understanding, OCR, screenshot diffing, and multi-provider routing for DeepSeek Harness.
zjcdkj
DeepSeek Harness (DSH) plugins. qwen-image gives a text-only coding model eyes: an image goes to a Qwen-VL route through ctx.llm and comes back as text, so DeepSeek keeps coding while Qwen looks. Pure ESM, no build permission at install. | DSH 插件集:qwen-image 让纯文本模型借千问 VL 读图,返回文本;纯 ESM,安装无需构建授权。
genusamblyrhynchusbrunooftoul602
Extend DeepSeek Harness composer to accept PDFs and more attachment formats Codex-style, with zero core changes and native pipeline reuse.
henryxiao709
DSH-PDF插件,让 AI 助手读取任意大小的 PDF 文件: 通过 pdfjs-dist 提取完整 Unicode 文本层(中文、英文及其它文字系统),并对扫描件/图片页自动 OCR, 手写笔记也能变成可读文本。DSH-PDF plugin — read any-size PDFs in DeepSeek Harness: full Unicode text (Chinese/English) via pdfjs-dist + automatic OCR (Windows WinRT / tesseract.js) for scanned pages. MIT.
Isanti2016
该仓库暂未提供项目说明。
junhongchashui
零修改、零切换的 DeepSeek Harness 视觉能力插件:纯文本模型粘贴即读图片,云端 + 本地 Ollama 双后端自动切换,ModLens v2 风格结构化证据输出。
kid-tea
Local text-only OCR plugin for DeepSeek Harness: ocr_image tool extracts text from images locally — no vision model, no API key, no external upload.
L-mimimi
一款 Windows 截图工具:截图 · 离线 OCR 文字识别 · 桌面置顶钉图,单文件绿色版,双击即用,无需安装、无需联网。
leozou320-ai
Offline macOS Vision OCR for DeepSeek Harness — accurate, local, API-key free. | DeepSeek Harness 本地离线 OCR 插件
qizhen2021
该仓库暂未提供项目说明。
SKL-666666
图片结构化分析技能:双引擎OCR+形状/表格/图标/布局识别,让纯文本模型看懂图片
wenliang9527
该仓库暂未提供项目说明。
wuwangmao
DSH bundle: Qwen multimodal bridge — vision (qwen3-vl), speech-to-text (qwen3-asr), text-to-image (qwen-image), for DeepSeek Harness
zhuiyueya
Give text-only DeepSeek models eyes — a DeepSeek Harness plugin that transparently converts chat images into OCR text + vision-model descriptions before they reach the LLM. Configure vision backends (GLM-4V, Qwen-VL, Gemini, Ollama…) right in the Models settings page; multi-backend fallback chain, double-layer caching, no config files.