Back to catalog

GITHUB TOPIC

vision-language-model

11projects include this topic

The "vision-language-model" topic on GitHub groups 11 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by agent-vision-toolkit with 1.2k GitHub stars. agent-vision-toolkit — A better vision toolbox and skill for pure-text models to "see": multi-image understanding, image Q&A, frontend UI restoration, GUI automation and more, with optional seamless integration into multiple mainstream agents, recognizing pasted images directly | A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.

{count} projects

Exact GitHub Topic match

DSH PluginsSkill

A better vision toolbox and skill for pure-text models to "see": multi-image understanding, image Q&A, frontend UI restoration, GUI automation and more, with optional seamless integration into multiple mainstream agents, recognizing pasted images directly | A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

Structure check pending
agentagent-skillsclaude-codecodex
Files & dataSkill

[dsh] A more powerful visual toolkit for text-only models: one-line install and use, paste images for direct recognition, multi-image Q&A, screenshot-to-frontend UI restoration, and more|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.

Structure check pending
agent-skillsagent-vision-toolkitcomputer-visiondeepseek
Files & dataPlugin

Auditable vision and cross-platform Computer Use runtime for DeepSeek Harness — strict evidence, health-checked failover, original pixels, and Token accounting.

Structure check pending
ai-agentai-evaluationauditable-aibrowser-automation
Files & dataPlugin

On-demand vision for text-only DeepSeek Harness (DSH) sessions: images become markers, and a vision_describe tool sends only image + question to an OpenAI-compatible vision model

Structure check pending
agentdeepseekdeepseek-harnessdsh
Files & dataPlugin

Hosted free vision sidecar for DeepSeek Harness with durable session evidence

Structure check pending
deepseek-harnessdsh-pluginglmvision-language-model
Files & dataPlugin

DeepSeek Harness plugin that bridges session images to pluggable vision APIs while keeping DeepSeek as the primary model.

Structure check pending
deepseekdeepseek-aideepseek-harnessdeepseek-harness-desktop
Files & dataPlugin

Image understanding, OCR, and persistent visual evidence for text-only DeepSeek Harness models

Structure check pending
ai-agentscomputer-visiondeepseekdeepseek-harness
Files & dataPlugin

DeepSeek Harness all-in-one: no model switching — regular DeepSeek auto-routes to vision & image gen. Multi-backend: Gemini + any OpenAI-compatible (GPT-4o, Qwen-VL, GLM-4V, gpt-image, DALL-E, Flux, OpenRouter). gemini_vision/gemini_generate_image/gemini_optimize_image with vision self-check. Better than modlens.

Structure check pending
aicordisdeepseekdeepseek-harness
Files & dataPlugin

dsh-vision-bridge

TwistedRiCen

DSH-native Vision Evidence bridge for text-only reasoning models with native image attachments and strict multi-image validation.

Structure check pending
deepseek-harnessdsh-pluginllmmultimodal
Files & dataPlugin

DeepSeek Harness multimodal vision bridge (dsh-vision-bridge): pasted images auto-converted to VL text descriptions via llm/stream (solves UNSUPPORTED_CONTENT) + view_image/ocr_image active vision tools + native multimodal routing auto-skip (rc.7); zero dependencies. Vision bridge for text-only DeepSeek models.

Structure check pending
dashscopedeepseek-harnessdshdsh-plugin
Files & dataPlugin

duhai-vision

hamliy-feng

Visual model adapter for Codex and DeepSeek Harness, powered by PaddleOCR-VL and Qwen.

Structure check pending
ai-agentscodexdeepseek-harnessdsh-plugin