The "vision" topic on GitHub groups 159 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by modlens with 4.1k GitHub stars. modlens — The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.
The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | The strongest vision add-on plugin for DeepSeek Harness, adding vision capability to text-only models like DeepSeek and GLM, paste an image and get structured JSON evidence (OCR, layout, semantics).
A better vision toolbox and skill for pure-text models to "see": multi-image understanding, image Q&A, frontend UI restoration, GUI automation and more, with optional seamless integration into multiple mainstream agents, recognizing pasted images directly | A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
DeepSeek Harness plugin: DeepSeek Pro brain + automatic image recognition. Images attached in the GUI are processed by default with the official deepseek-v4-flash-vision-exp native vision model, converted to text, and passed to DeepSeek for answers (even text-only V4-Pro can see images); supports any OpenAI-compatible VLM such as Bailian/Zhipu/OpenRouter; auto-detects local Ollama without a key; one-question confirmation during install
Dsh-visual-plugin.Give your text-only model eyes: forward user images to any OpenAI-compatible vision model and see the results in a Web UI right panel
Adds external image recognition models to DeepSeek Harness: round whale button, send image for recognition with auto-return, model autonomous screenshot + image recognition tools, automatic multi-protocol adaptation, one-click install for beginners (auto-download if Node.js is not installed)
DSH plugin: images and files straight to text-only models — images keep the native attachment experience, PDF/Office/archives/video/audio show as square chips in the attachment rail, auto-converted to workspace paths on send; pairs with dsh-vision-toolkit for paste-to-view. A DSH plugin that delivers images AND files to text-only models as workspace paths: images keep the native attachment UI, other files show as square chips in the rail, paths append on send — pairs with dsh-vision-toolkit.
One tool = all MiniMax multimodal capabilities: DSH text-only models see images/draw images/generate video/speak/sing/cover/search/check quota | One mmx_bridge tool = all MiniMax multimodal (VLM/image/video/speech/music/cover/search/quota) for DeepSeek Harness (DSH)
Auxiliary models for DeepSeek Harness: vision understanding and context compression through dedicated model routes. DeepSeek Harness auxiliary model plugin: provides independent model routes, tools and system prompts for vision understanding, context compression, approval review, sub-agents, session titles and image generation, never touching the main conversation model.
A vision plugin built for DSH (DeepSeek Harness), now supports Agent-invoked image display / Vision plugin for DSH(DeepSeek Harness),support Proactive Image Display.
On-demand vision for text-only DeepSeek Harness (DSH) sessions: images become markers, and a vision_describe tool sends only image + question to an OpenAI-compatible vision model
[Discontinued] DeepSeek Harness vision bridge plugin: the new Harness natively supports image recognition, please use the native capability, this repository is for historical reference only.
DeepSeek Harness plugin: lets models that can't see images, such as deepseek-v4-flash, handle chat images too, with a built-in image recognition tool. Install: dsh plugin --profile web add dsh-image-pathify
Give DeepSeek a pair of eyes and a paintbrush: paste screenshots/images straight into the conversation, the GLM vision model first transcribes the image content precisely (error messages, code, UI preserved verbatim), then DeepSeek continues with your question —— all in the same turn, seamless throughout; when an illustration is needed, DeepSeek automatically calls the text-to-image backend and shows the image in the conversation.
A lightweight DeepSeek Harness vision delegation tool for text-only routes, with native OpenAI Responses, Chat Completions, and Anthropic Messages adapters.
Codex-style attachment formats for the DeepSeek Harness Web GUI: PDF text-layer extraction, Office text extraction, scanned-PDF OCR, long-document spill + index cards, image-to-PNG.
Standalone screen capture for DeepSeek Harness (dsh): browser hotkeys plus an agent-facing capture+read tool. Forked out of @liustack/modlens#48 (upstream declined the feature).
DSH plugin: text-only models (e.g. DeepSeek-V4) automatically see images via a vision model. Official surface-replace, cache-friendly, human transcript untouched. Vision bridge for text-only models
DeepSeek Harness vision plugin: provides image recognition for models without native vision (Alibaba Cloud Bailian qwen3.5-omni-plus, auto-switches to Zhipu glm-4.6v-flash on failure). Ported and adapted from claude-vision-skill. | Vision tool for DeepSeek Harness
Silent vision bridge for DeepSeek Harness: route chat images to a fixed vision model, preserve UI originals, and reuse observations across compaction and restarts.
Local-first vision for DeepSeek Harness: structured JSON evidence (OCR/layout/semantics) from local VLMs (LM Studio/Ollama), zero API cost, images never leave your machine.
DeepSeek Harness (DSH) local vision eyes plugin: screen tool (screenshots/images handed to a local vision model for description) + ocr tool (Windows built-in OCR extracts text character by character). Zero cloud, zero GPU for OCR, images never leave the machine.
DeepSeek Harness (DSH) vision plugin: image recognition via Edge + Doubao Web, zero cost, no API Key. General recognition + math-modeling diagram specialty (geometry/flowcharts/charts/tables/formulas) + uncertainty clarification loop. Vision plugin for DeepSeek Harness: image understanding via Edge + Doubao Web, zero cost, no API key. General recognition + math-modeling diagrams (geometry/flowcharts/charts/tables/formulas) + clarify loop for uncertainties.
Plug-in vision for text-only models on DSH, with native interaction for image understanding and generation, GUI automation, through layered evidence memory and cache.
A DeepSeek Harness tool plugin that lets text-only agents "see" local images — auto-detects the real format and returns a detailed text description via any OpenAI-compatible vision model.
Give text-only DeepSeek models on-demand vision: upload images, DeepSeek answers by calling a view_image tool backed by any OpenAI-compatible vision endpoint (Qwen/DashScope by default).
Vision-augmented DeepSeek adapter plugin for DeepSeek Harness: a vision-capable model describes image input, then a text-only DeepSeek model reasons over the description
Give DeepSeek Harness text-only models vision: Codex-style drag-and-drop of images into the chat box automatically routes them to the user-configured vision model for text conversion, v4 reads images without switching models. Vision for text-only DSH models: routes chat-box images to a user-configured vision model and returns text descriptions, deepseek-v4-pro reads images without switching.
DeepSeek Harness vision bridge: automatically discovers your configured multimodal models and equips text-only main models with a vision tool, returning results as plain text. Zero config, one-command install.
DeepSeek VisionPlus — official-grade vision extension for DeepSeek Harness. Routes image understanding to a free vision-model pool (Zhipu GLM, SiliconFlow Qwen) with automatic fallback, rate limiting, one-click platform tests and live status lines; text stays on DeepSeek. One-command install. MIT.
Scene-aware vision routing layer for DeepSeek Harness (dsh): decides which engine/backend an image should go to (chat/UI/table/code) before other vision plugins route it. Switch-gated, front-loaded, never touches other plugins' tools. Scene-level vision routing layer.
Design-fidelity QA for DeepSeek Harness: lend any text-only model an eye, then judge whether the implementation matches the mock. Ships the benchmark behind that judgement — four fixtures, 23 injected defects, and every raw model transcript. Retires itself when DeepSeek ships vision.
Browser tool for DeepSeek Harness that can SEE the page: browser-use over CDP driven by deepseek-v4-flash-vision-exp. Reads canvas text, text inside images and rendered charts, returns schema-validated JSON, and reports per-run cost.
Eyes for text-only DeepSeek: view_image tool (local Ollama or any OpenAI-compatible VLM) + chat image-attachment bridge — paste/drop images in the chat and the model can see them.
DeepSeek Harness all-in-one: no model switching — regular DeepSeek auto-routes to vision & image gen. Multi-backend: Gemini + any OpenAI-compatible (GPT-4o, Qwen-VL, GLM-4V, gpt-image, DALL-E, Flux, OpenRouter). gemini_vision/gemini_generate_image/gemini_optimize_image with vision self-check. Better than modlens.
DeepSeek Harness (DSH) plugins. qwen-image gives a text-only coding model eyes: an image goes to a Qwen-VL route through ctx.llm and comes back as text, so DeepSeek keeps coding while Qwen looks. Pure ESM, no build permission at install. | DSH plugin set: qwen-image lets text-only models read images via Qwen VL and return text; pure ESM, no build authorization needed at install.
DeepSeek Harness (DSH) vision plugin — gives the agent screen/window vision + computer-use ability (see/ocr/list_windows + mouse and keyboard control), cross-language calls to the bundled Python cvision, available on Windows / macOS.
DeepSeek Harness (DSH) Web plugin: drag-and-drop file preview / Markdown rendering / file box / smart desktop screenshot (fullscreen·region·window) / external image recognition and OCR text ingestion. Full development history with 21 snapshot tags.
Give text-only DeepSeek Harness (dsh) agents vision — pasted images auto-convert to text descriptions with persistent caching, each image converted only once.
Local OCR + color-statistics + pixel-scan + metadata bridge for text-only Pi Coding Agent models. Pi port of dsh-pseudo-vision, no external vision API.
DeepSeek Harness plugin: lifts the Web GUI image input limit, image attachments are textualized and handed to a vision skill | dsh plugin that lifts the image-input gate and hands attachments to a vision skill
Model-facing read_tiff tool for DeepSeek Harness: decodes TIFF/TIF images (multi-page, LZW/Deflate/PackBits/CCITT/JPEG compression, bilevel, 8/16-bit and float) into viewable PNGs with full header metadata, plus optional vision-model description.
Give text-only models eyes: analyze_image tool for DeepSeek Harness, backed by free Chinese vision APIs (GLM-4V-Flash / Qwen-VL) or any OpenAI-compatible endpoint. dsh plugin that gives text-only models eyes.
DeepSeek Harness host plugin that lets text-only models see chat images through desktop Doubao (CDP bridge, works with all presets, recognition cancellable)
DeepSeek Harness vision enhancement plugin: hands images to an external vision model for analysis and outputs plain-text evidence with coordinate-based visual primitives, so non-multimodal text models can also understand images, screenshots and documents in conversation.
dsh-autovision: paste an image into a text-only model composer and a configured multimodal model transcribes it to text automatically. Twin-provider auto-routing + agent-callable read-image tool. No built-in keys, no relay.
A native DeepSeek Harness (DSH) Cordis plugin that analyzes images through the reverse-engineered chat.deepseek.com vision mode (model_type=vision) — free, no third-party vision API key required. Native DeepSeek Harness (DSH) Cordis plugin: analyzes images via the reverse-engineered chat.deepseek.com vision mode (model_type=vision) — free, no third-party vision API key required.
Automatic model routing for DeepSeek Harness: three difficulty tiers (hard/normal/easy) plus vision routing, picking models you already configured under Settings ? Models. ???? + ?????????
Vision for DeepSeek Harness agents — paste images in the Web composer, delegate reads to Kimi/MiniMax vision routes on isolated contexts; zero image bytes in the main session
Vision tool plugin for DeepSeek Harness (DSH): give text-only models like deepseek-v4-flash image recognition via Alibaba Bailian / any OpenAI-compatible vision API. Add an image recognition tool to DeepSeek Harness models without vision.
A see_image vision tool plugin for DeepSeek Harness — describe images through any OpenAI-compatible vision model (GitHub Copilot, OpenAI, Ollama, vLLM, LM Studio).
Agnes omni-modal plugin for DeepSeek Harness: agnes_vision (image understanding) + agnes_image (text-to-image / image-to-image) + a vision bridge that lets you send images in chat. API key via DSH credentials, never in code.
Image recognition plugin for DeepSeek Harness: automatically detects the current model's vision capability, supports multi-provider vision model management and detection
Give DeepSeek Harness eyes: one-click install the dsh-plugin-deepeye vision plugin with the free Zhipu GLM-4V-Flash model. Paste images into text-only LLMs
Bridge Apple on-device Vision framework (macOS) into DeepSeek Harness: OCR, image classification, face detection, document layout as local dsh tools. No network, no API key.
DSH plugin: keep text-only DeepSeek models (V4-Flash / V4-Pro) and auto-route image-bearing requests to the official vision model (deepseek-v4-flash-vision-exp) - no manual model switching.
Personal Agent System on DeepSeek Harness — everything is a plugin: completion notifications, vision for text models, reply annotations, multi-session tabs, video support, sandbox patch
DeepSeek Harness plugin: upload/paste images; on send transcribe via a vision model (DashScope) or offline Windows OCR. Built with deepseek-harness, Made with the DeepSeek-harness.
DeepSeek Harness plugin: route image-bearing messages to a user-configured OpenAI-compatible vision endpoint when the active text model cannot see images