The "multimodal" topic on GitHub groups 84 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by modlens with 4.1k GitHub stars. modlens — The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.
The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | The strongest vision add-on plugin for DeepSeek Harness, adding vision capability to text-only models like DeepSeek and GLM, paste an image and get structured JSON evidence (OCR, layout, semantics).
A better vision toolbox and skill for pure-text models to "see": multi-image understanding, image Q&A, frontend UI restoration, GUI automation and more, with optional seamless integration into multiple mainstream agents, recognizing pasted images directly | A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
DeepSeek Harness (DSH) plugin: dispatch work to DSH agents from Claude Code / Codex — native subagent progress, in-host worker sessions with per-tier presets, and a multimodal bridge that lends the text-only harness vision and image generation.
Self-contained DeepSeek Harness (DSH) plugin for Provider/Auth login, model switching, image fallback, token/cost analytics, and same-port Web restart. Useful? A star helps.
DeepSeek Harness plugin: DeepSeek Pro brain + automatic image recognition. Images attached in the GUI are processed by default with the official deepseek-v4-flash-vision-exp native vision model, converted to text, and passed to DeepSeek for answers (even text-only V4-Pro can see images); supports any OpenAI-compatible VLM such as Bailian/Zhipu/OpenRouter; auto-detects local Ollama without a key; one-question confirmation during install
Dsh-visual-plugin.Give your text-only model eyes: forward user images to any OpenAI-compatible vision model and see the results in a Web UI right panel
One tool = all MiniMax multimodal capabilities: DSH text-only models see images/draw images/generate video/speak/sing/cover/search/check quota | One mmx_bridge tool = all MiniMax multimodal (VLM/image/video/speech/music/cover/search/quota) for DeepSeek Harness (DSH)
On-demand vision for text-only DeepSeek Harness (DSH) sessions: images become markers, and a vision_describe tool sends only image + question to an OpenAI-compatible vision model
[Discontinued] DeepSeek Harness vision bridge plugin: the new Harness natively supports image recognition, please use the native capability, this repository is for historical reference only.
Give DeepSeek a pair of eyes and a paintbrush: paste screenshots/images straight into the conversation, the GLM vision model first transcribes the image content precisely (error messages, code, UI preserved verbatim), then DeepSeek continues with your question —— all in the same turn, seamless throughout; when an illustration is needed, DeepSeek automatically calls the text-to-image backend and shows the image in the conversation.
DeepSeek Harness third-party API and custom model settings plugin: supports request headers, User-Agent, model list, image input and reasoning levels | WebUI plugin for third-party APIs and custom models with request headers, image input, and reasoning levels
A lightweight DeepSeek Harness vision delegation tool for text-only routes, with native OpenAI Responses, Chat Completions, and Anthropic Messages adapters.
DSH plugin: text-only models (e.g. DeepSeek-V4) automatically see images via a vision model. Official surface-replace, cache-friendly, human transcript untouched. Vision bridge for text-only models
DeepSeek Harness plugin. dsh plugin supporting manual selection of model capabilities when creating a custom model, such as whether the model supports image input. The DSH plugin allows users to manually configure model capabilities when creating a custom model, such as image input support.
Silent vision bridge for DeepSeek Harness: route chat images to a fixed vision model, preserve UI originals, and reuse observations across compaction and restarts.
Plug-in vision for text-only models on DSH, with native interaction for image understanding and generation, GUI automation, through layered evidence memory and cache.
Give text-only DeepSeek models on-demand vision: upload images, DeepSeek answers by calling a view_image tool backed by any OpenAI-compatible vision endpoint (Qwen/DashScope by default).
Give DeepSeek Harness text-only models vision: Codex-style drag-and-drop of images into the chat box automatically routes them to the user-configured vision model for text conversion, v4 reads images without switching models. Vision for text-only DSH models: routes chat-box images to a user-configured vision model and returns text descriptions, deepseek-v4-pro reads images without switching.
DeepSeek Harness vision bridge: automatically discovers your configured multimodal models and equips text-only main models with a vision tool, returning results as plain text. Zero config, one-command install.
Scene-aware vision routing layer for DeepSeek Harness (dsh): decides which engine/backend an image should go to (chat/UI/table/code) before other vision plugins route it. Switch-gated, front-loaded, never touches other plugins' tools. Scene-level vision routing layer.
Qwen-MM-Plugins integration plugin for DeepSeek Harness: 12 multimodal MCP tools (vision/OCR/grounding/ASR/audio-video), Web settings page (paste Qwen API Key and go), built-in skills and one-click installer
Design-fidelity QA for DeepSeek Harness: lend any text-only model an eye, then judge whether the implementation matches the mock. Ships the benchmark behind that judgement — four fixtures, 23 injected defects, and every raw model transcript. Retires itself when DeepSeek ships vision.
Eyes for text-only DeepSeek: view_image tool (local Ollama or any OpenAI-compatible VLM) + chat image-attachment bridge — paste/drop images in the chat and the model can see them.
DeepSeek Harness all-in-one: no model switching — regular DeepSeek auto-routes to vision & image gen. Multi-backend: Gemini + any OpenAI-compatible (GPT-4o, Qwen-VL, GLM-4V, gpt-image, DALL-E, Flux, OpenRouter). gemini_vision/gemini_generate_image/gemini_optimize_image with vision self-check. Better than modlens.
DeepSeek Harness (DSH) plugins. qwen-image gives a text-only coding model eyes: an image goes to a Qwen-VL route through ctx.llm and comes back as text, so DeepSeek keeps coding while Qwen looks. Pure ESM, no build permission at install. | DSH plugin set: qwen-image lets text-only models read images via Qwen VL and return text; pure ESM, no build authorization needed at install.
Unified access to the four Image2 and Nano Banana models, covering text-to-image, multi-reference image editing, 2K/4K output, sequential batch tasks, default model persistence and masked Key configuration.
Local OCR + color-statistics + pixel-scan + metadata bridge for text-only Pi Coding Agent models. Pi port of dsh-pseudo-vision, no external vision API.
DeepSeek Harness vision enhancement plugin: hands images to an external vision model for analysis and outputs plain-text evidence with coordinate-based visual primitives, so non-multimodal text models can also understand images, screenshots and documents in conversation.
Vision for DeepSeek Harness agents — paste images in the Web composer, delegate reads to Kimi/MiniMax vision routes on isolated contexts; zero image bytes in the main session
A see_image vision tool plugin for DeepSeek Harness — describe images through any OpenAI-compatible vision model (GitHub Copilot, OpenAI, Ollama, vLLM, LM Studio).
Give text-only LLMs a pair of eyes. A DeepSeek Harness (DSH) native skill + zero-dependency Python CLI, adding image understanding and document parsing (OCR, tables, formulas, PDF → Markdown) to DeepSeek and other text-only models, using third-party multimodal APIs with free tiers first, direct connection on domestic networks, no proxy needed.
Agnes omni-modal plugin for DeepSeek Harness: agnes_vision (image understanding) + agnes_image (text-to-image / image-to-image) + a vision bridge that lets you send images in chat. API key via DSH credentials, never in code.
DSH plugin: keep text-only DeepSeek models (V4-Flash / V4-Pro) and auto-route image-bearing requests to the official vision model (deepseek-v4-flash-vision-exp) - no manual model switching.
Give text-only DeepSeek models eyes — a DeepSeek Harness plugin that transparently converts chat images into OCR text + vision-model descriptions before they reach the LLM. Configure vision backends (GLM-4V, Qwen-VL, Gemini, Ollama…) right in the Models settings page; multi-backend fallback chain, double-layer caching, no config files.