The "image-to-text" topic on GitHub groups 8 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by modlens with 4.1k GitHub stars. modlens — The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.
The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | The strongest vision add-on plugin for DeepSeek Harness, adding vision capability to text-only models like DeepSeek and GLM, paste an image and get structured JSON evidence (OCR, layout, semantics).
DeepSeek Harness vision bridge: automatically discovers your configured multimodal models and equips text-only main models with a vision tool, returning results as plain text. Zero config, one-command install.
DeepSeek Harness (DSH) plugins. qwen-image gives a text-only coding model eyes: an image goes to a Qwen-VL route through ctx.llm and comes back as text, so DeepSeek keeps coding while Qwen looks. Pure ESM, no build permission at install. | DSH plugin set: qwen-image lets text-only models read images via Qwen VL and return text; pure ESM, no build authorization needed at install.
Local OCR + color-statistics + pixel-scan + metadata bridge for text-only Pi Coding Agent models. Pi port of dsh-pseudo-vision, no external vision API.
DSH plugin: lets text-only (non-multimodal) models send and read images — bypasses the image admission gate, rewrites image blocks to local paths, and OCRs/describes them via configurable vision API providers (image_to_text tool).