The "vision-language-model" topic on GitHub groups 11 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by agent-vision-toolkit with 1.2k GitHub stars. agent-vision-toolkit — A better vision toolbox and skill for pure-text models to "see": multi-image understanding, image Q&A, frontend UI restoration, GUI automation and more, with optional seamless integration into multiple mainstream agents, recognizing pasted images directly | A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.
A better vision toolbox and skill for pure-text models to "see": multi-image understanding, image Q&A, frontend UI restoration, GUI automation and more, with optional seamless integration into multiple mainstream agents, recognizing pasted images directly | A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
[dsh] A more powerful visual toolkit for text-only models: one-line install and use, paste images for direct recognition, multi-image Q&A, screenshot-to-frontend UI restoration, and more|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.
Auditable vision and cross-platform Computer Use runtime for DeepSeek Harness — strict evidence, health-checked failover, original pixels, and Token accounting.
On-demand vision for text-only DeepSeek Harness (DSH) sessions: images become markers, and a vision_describe tool sends only image + question to an OpenAI-compatible vision model
DeepSeek Harness all-in-one: no model switching — regular DeepSeek auto-routes to vision & image gen. Multi-backend: Gemini + any OpenAI-compatible (GPT-4o, Qwen-VL, GLM-4V, gpt-image, DALL-E, Flux, OpenRouter). gemini_vision/gemini_generate_image/gemini_optimize_image with vision self-check. Better than modlens.