Back to catalog

GITHUB TOPIC

image-to-text

8projects include this topic

The "image-to-text" topic on GitHub groups 8 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by modlens with 4.1k GitHub stars. modlens — The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.

{count} projects

Exact GitHub Topic match

Files & dataSkill

modlens

liustack

The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | The strongest vision add-on plugin for DeepSeek Harness, adding vision capability to text-only models like DeepSeek and GLM, paste an image and get structured JSON evidence (OCR, layout, semantics).

Structure check pending
agent-skillsclaude-codeclaude-skillscodex
Files & dataPlugin

Paste images into DeepSeek Harness with a four-model vision race, OCR, and an automatic text bridge.

Structure check pending
deepseekdeepseek-harnessdsh-pluginimage-to-text
Files & dataPlugin

dsh-auto-vision

NormanFxxkingRockwell

DeepSeek Harness vision bridge: automatically discovers your configured multimodal models and equips text-only main models with a vision tool, returning results as plain text. Zero config, one-command install.

Structure check pending
cordicdeepseek-harnessdshdsh-plugin
Files & dataSkill

DeepSeek Harness (DSH) plugins. qwen-image gives a text-only coding model eyes: an image goes to a Qwen-VL route through ctx.llm and comes back as text, so DeepSeek keeps coding while Qwen looks. Pure ESM, no build permission at install. | DSH plugin set: qwen-image lets text-only models read images via Qwen VL and return text; pure ESM, no build authorization needed at install.

Structure check pending
agent-skillsclaude-codecodexcoding-agent
Files & dataPlugin

pi-pseudo-vision

DDDFXYqiming

Local OCR + color-statistics + pixel-scan + metadata bridge for text-only Pi Coding Agent models. Pi port of dsh-pseudo-vision, no external vision API.

Structure check pending
deepseek-harnessimage-to-textlocal-firstmit-license
Files & dataPlugin

DeepSeek Harness vision plugin: analyze_image (structured OCR evidence) + capture_image (USB camera visual loop). Camera visual loop + structured evidence, supports Ollama / DeepSeek / Xiaomi three backends.

Structure check pending
cameradeepseek-harnessdsh-pluginimage-to-text
Files & dataPlugin

deepseek-visual-plugin

zhangzhimou78-code

dsh-plugin

Structure check pending
cordiscordis-plugindeepseekdeepseek-harness
Files & dataPlugin

DSH plugin: lets text-only (non-multimodal) models send and read images — bypasses the image admission gate, rewrites image blocks to local paths, and OCRs/describes them via configurable vision API providers (image_to_text tool).

Structure check pending
deepseek-harnessdshdsh-pluginimage-to-text