Back to catalog

GITHUB TOPIC

multimodal

84projects include this topic

The "multimodal" topic on GitHub groups 84 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by modlens with 4.1k GitHub stars. modlens — The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.

{count} projects

Exact GitHub Topic match

Files & dataSkill

modlens

liustack

The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | The strongest vision add-on plugin for DeepSeek Harness, adding vision capability to text-only models like DeepSeek and GLM, paste an image and get structured JSON evidence (OCR, layout, semantics).

Structure check pending
agent-skillsclaude-codeclaude-skillscodex
DSH PluginsSkill

A better vision toolbox and skill for pure-text models to "see": multi-image understanding, image Q&A, frontend UI restoration, GUI automation and more, with optional seamless integration into multiple mainstream agents, recognizing pasted images directly | A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

Structure check pending
agentagent-skillsclaude-codecodex
DSH PluginsPlugin

Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.

Structure check pending
deepseek-harnessdshdsh-pluginmultimodal
Files & dataPlugin

tongflow

tong-io

TongFlow — Multimodal GenAI Studio

Structure check pending
3dagentaiai-tools
DSH PluginsSkill

全息闪卡通用 Agent Skill:多模态模型都能用,不挑宿主——一句话生成可拖转、会流光、带层次景深的 3D 闪卡网页,附可编辑 Blender 工程

Structure check pending
agent-skillagent-skillsblendercollectible-card
Models & MCPPlugin

dsh-crew

ZSeven-W

DeepSeek Harness (DSH) plugin: dispatch work to DSH agents from Claude Code / Codex — native subagent progress, in-host worker sessions with per-tier presets, and a multimodal bridge that lends the text-only harness vision and image generation.

Structure check pending
ai-agentsclaude-codecodexcoding-agent
Files & dataPlugin

dsh-vision

oil-oil

Near-native image understanding for DeepSeek Harness

Structure check pending
deepseek-harnessdsh-pluginimage-understandingmultimodal
Models & MCPPlugin

dsh-AuthInOne

Stormycry-cryp

Self-contained DeepSeek Harness (DSH) plugin for Provider/Auth login, model switching, image fallback, token/cost analytics, and same-port Web restart. Useful? A star helps.

Structure check pending
cost-attributioncost-trackingcustom-apideepseek-harness
Files & dataPlugin

DeepSeek Harness plugin: DeepSeek Pro brain + automatic image recognition. Images attached in the GUI are processed by default with the official deepseek-v4-flash-vision-exp native vision model, converted to text, and passed to DeepSeek for answers (even text-only V4-Pro can see images); supports any OpenAI-compatible VLM such as Bailian/Zhipu/OpenRouter; auto-detects local Ollama without a key; one-question confirmation during install

Structure check pending
dashscopedeepseek-harnessdsh-pluginimage-understanding
Files & dataPlugin

Dsh-visual-plugin.Give your text-only model eyes: forward user images to any OpenAI-compatible vision model and see the results in a Web UI right panel

Structure check pending
deepseekdeepseek-harnessdsh-pluginimage-description
Files & dataPlugin

One tool = all MiniMax multimodal capabilities: DSH text-only models see images/draw images/generate video/speak/sing/cover/search/check quota | One mmx_bridge tool = all MiniMax multimodal (VLM/image/video/speech/music/cover/search/quota) for DeepSeek Harness (DSH)

Structure check pending
agent-toolai-agentcordisdeepseek-harness
Files & dataPlugin

On-demand vision for text-only DeepSeek Harness (DSH) sessions: images become markers, and a vision_describe tool sends only image + question to an OpenAI-compatible vision model

Structure check pending
agentdeepseekdeepseek-harnessdsh
Files & dataPlugin

DeepSeek Harness plugin: describe_image — give a text-only model vision through an OpenAI-compatible VLM endpoint

Structure check pending
deepseekdeepseek-harnessdescribe-imagedsh
Files & dataPlugin

This repository does not yet provide a project description.

Structure check pending
deepseekdeepseek-harnessdshdsh-plugin
Files & dataPlugin

Paste images into DeepSeek Harness with a four-model vision race, OCR, and an automatic text bridge.

Structure check pending
deepseekdeepseek-harnessdsh-pluginimage-to-text
Files & dataPlugin

[Discontinued] DeepSeek Harness vision bridge plugin: the new Harness natively supports image recognition, please use the native capability, this repository is for historical reference only.

Structure check pending
agentaiattachmentchat
Files & dataPlugin

DeepEye vision plugin for DeepSeek Harness (DSH): image description, OCR, VQA, UI layout, and clipboard analysis.

Structure check pending
cordisdeepseek-harnessdshdsh-plugin
Files & dataPlugin

Give DeepSeek a pair of eyes and a paintbrush: paste screenshots/images straight into the conversation, the GLM vision model first transcribes the image content precisely (error messages, code, UI preserved verbatim), then DeepSeek continues with your question —— all in the same turn, seamless throughout; when an illustration is needed, DeepSeek automatically calls the text-to-image backend and shows the image in the conversation.

Structure check pending
deepseek-harnessdshdsh-plugindsh-plugins
Models & MCPPlugin

DeepSeek Harness third-party API and custom model settings plugin: supports request headers, User-Agent, model list, image input and reasoning levels | WebUI plugin for third-party APIs and custom models with request headers, image input, and reasoning levels

Structure check pending
deepseek-harnessdshdsh-pluginmodel-provider
Models & MCPDirectory

Model catalog, portraits, Agent selection, and multimodal runtimes for DeepSeek Harness

Outside plugin validation
aideepseek-harnessdshdsh-plugin
Files & dataPlugin

dsh-xiapan-media

dongsheng123132

Native vision, gpt-image-2 and Seedance plugins for DeepSeek Harness via Xiapan Cloud

Structure check pending
deepseek-harnessdsh-pluginimage-generationmultimodal
Files & dataPlugin

Vision routing and image generation for DeepSeek Harness through a fixed Mix model.

Structure check pending
deepseek-harnessdsh-plugingpt-image-2image-generation
Models & MCPPlugin

dsh-open-eyes

hyper-dsh-plugins

A lightweight DeepSeek Harness vision delegation tool for text-only routes, with native OpenAI Responses, Chat Completions, and Anthropic Messages adapters.

Structure check pending
anthropicdeepseek-harnessdsh-pluginmultimodal
Files & dataPlugin

Config-only DeepSeek Harness bundle for OpenAI-compatible vision models.

Structure check pending
deepseek-harnessdshdsh-pluginmultimodal
Files & dataPlugin

dsh-vision

reimu-create

DSH plugin: text-only models (e.g. DeepSeek-V4) automatically see images via a vision model. Official surface-replace, cache-friendly, human transcript untouched. Vision bridge for text-only models

Structure check pending
deepseek-harnessdsh-pluginmultimodalvision
Files & dataSkill

visual-review

wang-bool

dsh plugin supporting image upload and image recognition. Turns the ds experience multimodal

Structure check pending
deepseek-harnessdshdsh-plugindsh-plugins
InterfacePlugin

DeepSeek Harness plugin. dsh plugin supporting manual selection of model capabilities when creating a custom model, such as whether the model supports image input. The DSH plugin allows users to manually configure model capabilities when creating a custom model, such as image input support.

Structure check pending
deepseek-harnessdsh-pluginmodel-capabilitiesmodel-configuration
LifestylePlugin

dsh-voice

zhuiyueya

Voice for DeepSeek Harness(dsh) — speech-to-text input + read-aloud TTS for text-only DeepSeek, zero API key.

Structure check pending
ai-agentsdeepseekdeepseek-harnessdeepseek-harness-plugin
Files & dataPlugin

Silent vision bridge for DeepSeek Harness: route chat images to a fixed vision model, preserve UI originals, and reuse observations across compaction and restarts.

Structure check pending
deepseek-harnessdshdsh-pluginimage
Files & dataPlugin

dsh-eye-vision

AlloyPlane

This repository does not provide a project description yet.

Structure check pending
aideepseek-harnessdshdsh-plugin
Development toolsPlugin

deepsee

chang416

Vision + smart model routing for DeepSeek Harness. Gemini sees. DeepSeek codes.

Structure check pending
ai-agentsai-coding-agentclaude-codecodex
Models & MCPPlugin

dsh-vision-bridge

GooDAnDReaDY

Universal vision bridge for DeepSeek Harness: attachments with native models, 40+ tools, PDF/OCR/diagrams.

Unrecognized
deepseek-harnessdshdsh-pluginmultimodal
Files & dataPlugin

dsh-mindseye

kanchengw

Plug-in vision for text-only models on DSH, with native interaction for image understanding and generation, GUI automation, through layered evidence memory and cache.

Structure check pending
agentdeepseek-harnessdsh-plugingui-automation
Files & dataPlugin

dsh-eyes

Leeminjing

Give text-only DeepSeek models on-demand vision: upload images, DeepSeek answers by calling a view_image tool backed by any OpenAI-compatible vision endpoint (Qwen/DashScope by default).

Structure check pending
dashscopedeepseek-harnessdsh-pluginmultimodal
Files & dataPlugin

Transparent image preprocessing route for DeepSeek Harness

Structure check pending
ai-agentscordisdeepseekdeepseek-harness
Files & dataPlugin

Give DeepSeek Harness text-only models vision: Codex-style drag-and-drop of images into the chat box automatically routes them to the user-configured vision model for text conversion, v4 reads images without switching models. Vision for text-only DSH models: routes chat-box images to a user-configured vision model and returns text descriptions, deepseek-v4-pro reads images without switching.

Structure check pending
codex-styledeepseek-harnessdsh-pluginimage
Files & dataPlugin

dsh-auto-vision

NormanFxxkingRockwell

DeepSeek Harness vision bridge: automatically discovers your configured multimodal models and equips text-only main models with a vision tool, returning results as plain text. Zero config, one-command install.

Structure check pending
cordicdeepseek-harnessdshdsh-plugin
Models & MCPPlugin

dsh-qwen-mm

RRRosmontis

Qwen-MM-Plugins integration bundle for DeepSeek Harness (dsh) — multimodal MCP tools (vision, OCR, ASR, search, video, Blender, FreeCAD) + image attachment bridge. Enables native multimodal support in DeepSeek Harness.

Structure check pending
agentaideepseekdeepseek-harness
Files & dataPlugin

Multimodal plugin for DeepSeek harness, close to a native experience.

Structure check pending
deepseek-harnessdoubaodshdsh-plugin
Models & MCPPlugin

text-llm-vision

shaoqiuyuavailable

Scene-aware vision routing layer for DeepSeek Harness (dsh): decides which engine/backend an image should go to (chat/UI/table/code) before other vision plugins route it. Switch-gated, front-loaded, never touches other plugins' tools. Scene-level vision routing layer.

Structure check pending
deepseek-harnessdshdsh-pluginlocal-ai
Models & MCPPlugin

Qwen-MM-Plugins integration plugin for DeepSeek Harness: 12 multimodal MCP tools (vision/OCR/grounding/ASR/audio-video), Web settings page (paste Qwen API Key and go), built-in skills and one-click installer

Structure check pending
asrdeepseek-harnessdsh-pluginmcp
Files & dataPlugin

dsh-design-qa

sunxin-ai

Design-fidelity QA for DeepSeek Harness: lend any text-only model an eye, then judge whether the implementation matches the mock. Ships the benchmark behind that judgement — four fixtures, 23 injected defects, and every raw model transcript. Retires itself when DeepSeek ships vision.

Structure check pending
benchmarkdeepseek-harnessdesign-qadesign-review
Files & dataPlugin

mimo-vision

wulusai2333

DeepSeek Harness (DSH) native plugin — describe_image tool: a vision bridge (image → mimo-v2.5 → text description) over the ctx.fs / ctx.credentials seams

Structure check pending
agentcordisdeepseek-harnessdsh
Models & MCPPlugin

Eyes for text-only DeepSeek: view_image tool (local Ollama or any OpenAI-compatible VLM) + chat image-attachment bridge — paste/drop images in the chat and the model can see them.

Structure check pending
deepseek-harnessdshdsh-pluginmultimodal
Files & dataPlugin

DeepSeek Harness all-in-one: no model switching — regular DeepSeek auto-routes to vision & image gen. Multi-backend: Gemini + any OpenAI-compatible (GPT-4o, Qwen-VL, GLM-4V, gpt-image, DALL-E, Flux, OpenRouter). gemini_vision/gemini_generate_image/gemini_optimize_image with vision self-check. Better than modlens.

Structure check pending
aicordisdeepseekdeepseek-harness
Files & dataSkill

DeepSeek Harness (DSH) plugins. qwen-image gives a text-only coding model eyes: an image goes to a Qwen-VL route through ctx.llm and comes back as text, so DeepSeek keeps coding while Qwen looks. Pure ESM, no build permission at install. | DSH plugin set: qwen-image lets text-only models read images via Qwen VL and return text; pure ESM, no build authorization needed at install.

Structure check pending
agent-skillsclaude-codecodexcoding-agent
Files & dataPlugin

dsh-llm-vision

1710782766

Reliable vision + OCR for text-only models on DeepSeek Harness: describe_image (normal/critical) + extract_text tools, auto-preprocessing, retries, and a persistent answer cache.

Structure check pending
deepseek-harnessdshdsh-pluginimage-understanding
Files & dataPlugin

Unified access to the four Image2 and Nano Banana models, covering text-to-image, multi-reference image editing, 2K/4K output, sequential batch tasks, default model persistence and masked Key configuration.

Structure check pending
88apideepseek-harnessdshdsh-plugin
Files & dataPlugin

Add image recognition to text-only DeepSeek Harness models: analyze_image forwards images to any OpenAI-compatible vision endpoint | Vision bridge for text-only DSH models

Structure check pending
deepseek-harnessdsh-pluginmultimodalopenai-compatible
Files & dataPlugin

pi-pseudo-vision

DDDFXYqiming

Local OCR + color-statistics + pixel-scan + metadata bridge for text-only Pi Coding Agent models. Pi port of dsh-pseudo-vision, no external vision API.

Structure check pending
deepseek-harnessimage-to-textlocal-firstmit-license
Files & dataPlugin

dsh-glm-vision

fightingFirefox

Connect Zhipu GLM vision models in dsh, letting text models like DeepSeek see images through the glm_vision tool.

Structure check pending
deepseek-harnessdsh-pluginglmmultimodal
Files & dataPlugin

Vision tools for DeepSeek Harness: OCR, chart extraction, UI review, comparison & image-to-code via OpenAI- or Anthropic-compatible endpoints, with a built-in FREE anonymous vision source and automatic rate-limit failover. Supports OpenAI/Anthropic compatible vision endpoints.

Structure check pending
deepseekdeepseek-harnessdsh-pluginfree
Files & dataPlugin

DeepSeek Harness browser-extension edition: side panel with page-context awareness, vision-bridge multimodal image reading, and voice input

Structure check pending
ai-agentbrowser-extensionchrome-extensiondeepseek
Files & dataPlugin

DeepSeek Harness vision enhancement plugin: hands images to an external vision model for analysis and outputs plain-text evidence with coordinate-based visual primitives, so non-multimodal text models can also understand images, screenshots and documents in conversation.

Structure check pending
deepseek-harnessdeepseek-harness-plugindshdsh-plugin
Files & dataPlugin

dsh-mingmu

Lab-sku

Mingmu VisionBridge - in-house vision bridge: when a blind model receives an image, it automatically calls a vision model to recognize it

Structure check pending
ai-plugindeepseekdeepseek-harnessdsh-plugin
Files & dataPlugin

DeepSeek Harness vision plugin: analyze_image (structured OCR evidence) + capture_image (USB camera visual loop). Camera visual loop + structured evidence, supports Ollama / DeepSeek / Xiaomi three backends.

Structure check pending
cameradeepseek-harnessdsh-pluginimage-to-text
Files & dataPlugin

Vision for DeepSeek Harness agents — paste images in the Web composer, delegate reads to Kimi/MiniMax vision routes on isolated contexts; zero image bytes in the main session

Structure check pending
ai-agentcordisdeepseekdeepseek-harness
Files & dataPlugin

Native-vision Windows computer-use for DeepSeek Harness: screenshots, UIA, OCR, and approval-gated input

Structure check pending
automationcomputer-usedeepseek-harnessdeepseek-v4
Files & dataPlugin

Bailian (DashScope) Kimi LLM adapter plugin for DeepSeek Harness — kimi-k3 with image input, thinking and tool calling. No build step.

Structure check pending
bailiandashscopedeepseek-harnessdsh-plugin
Files & dataPlugin

dsh-vision-link

sprainJinyu

Route-preserving image understanding for text-only models in DeepSeek Harness (DSH).

Structure check pending
deepseek-harnessdsh-pluginimage-understandingjavascript
Models & MCPPlugin

A see_image vision tool plugin for DeepSeek Harness — describe images through any OpenAI-compatible vision model (GitHub Copilot, OpenAI, Ollama, vLLM, LM Studio).

Structure check pending
deepseek-harnessdshdsh-plugingithub-copilot
Files & dataPlugin

dsh-vision-bridge

TwistedRiCen

DSH-native Vision Evidence bridge for text-only reasoning models with native image attachments and strict multi-image validation.

Structure check pending
deepseek-harnessdsh-pluginllmmultimodal
DSH PluginsSkill

Give text-only LLMs a pair of eyes. A DeepSeek Harness (DSH) native skill + zero-dependency Python CLI, adding image understanding and document parsing (OCR, tables, formulas, PDF → Markdown) to DeepSeek and other text-only models, using third-party multimodal APIs with free tiers first, direct connection on domestic networks, no proxy needed.

Structure check pending
dsh-plugindsh-skillfree-firstmultimodal
Files & dataPlugin

dsh-agnes-omni

wumu1111111

Agnes omni-modal plugin for DeepSeek Harness: agnes_vision (image understanding) + agnes_image (text-to-image / image-to-image) + a vision bridge that lets you send images in chat. API key via DSH credentials, never in code.

Structure check pending
agnesdeepseek-harnessdshdsh-plugin
Development toolsPlugin

Bring your Grok subscription into DSH as an ACP subagent, extending native images with audio and video tools.

Structure check pending
acpaudiodeepseek-harnessdsh-plugin
Files & dataPlugin

Vision toolkit for DeepSeek Harness -- give text-only agents eyes

Structure check pending
deepseek-harnessdeepseek-vldsh-pluginmultimodal
Files & dataPlugin

deepseek-visual-plugin

zhangzhimou78-code

dsh-plugin

Structure check pending
cordiscordis-plugindeepseekdeepseek-harness
Files & dataPlugin

dsh-vision

zoahdev

Give DeepSeek Harness eyes: analyze images with an OpenAI-compatible vision model via a vision_analyze tool.

Structure check pending
agentdeepseek-harnessdsh-pluginimage
Files & dataPlugin

Dynamic multimodal-to-text projection and cross-model vision bridge for DeepSeek Harness (dsh)

Structure check pending
claudecordisdeepseekdeepseek-harness
Files & dataPlugin

dsh-plugin-glm-vision

baldovinmarques391-design

GLM Vision plugin for DSH: image translation for non-multimodal models via GLM-4V-Flash

Structure check pending
deepseek-harnessdsh-pluginglmmultimodal
Files & dataPlugin

DSH plugin: auto-detect and configure model capabilities (reasoningEfforts + input modalities) for llm-pi-ai. Successor to dsh-reasoning-efforts.

Structure check pending
deepseek-harnessdshdsh-plugininput-modalities
Files & dataPlugin

Unified vision and image-generation plugin suite for DeepSeek Harness

Structure check pending
deepseek-harnessdsh-pluginimage-generationmultimodal
Files & dataPlugin

dsh-vision-bridge

cyh12345678910

Multi-backend vision plugin for DeepSeek Harness — API (OpenAI Vision) + CDP (Doubao bridge), cross-platform, cached, configurable

Structure check pending
cdpcordisdeepseek-harnessdsh-plugin
Files & dataPlugin

Image input fallback for DeepSeek Harness with native multimodal model detection and Volcengine Ark vision.

Structure check pending
arkdeepseekdeepseek-harnessdsh-plugin
Files & dataPlugin

DeepSeek Harness multimodal vision bridge (dsh-vision-bridge): pasted images auto-converted to VL text descriptions via llm/stream (solves UNSUPPORTED_CONTENT) + view_image/ocr_image active vision tools + native multimodal routing auto-skip (rc.7); zero dependencies. Vision bridge for text-only DeepSeek models.

Structure check pending
dashscopedeepseek-harnessdshdsh-plugin
Models & MCPPlugin

A zero-config, multi-provider vision tool for DeepSeek Harness with automatic local model discovery and privacy-aware remote fallback.

Structure check pending
deepseek-harnessdsh-pluginmultimodalollama
Files & dataPlugin

dsh-vision-relay

junhongchashui

Zero-modification, zero-switching vision plugin for DeepSeek Harness: text-only models read images on paste, cloud + local Ollama dual backends auto-switch, ModLens v2-style structured evidence output.

Structure check pending
cordisdeepseek-harnessdshdsh-plugin
Files & dataPlugin

DSH plugin: keep text-only DeepSeek models (V4-Flash / V4-Pro) and auto-route image-bearing requests to the official vision model (deepseek-v4-flash-vision-exp) - no manual model switching.

Structure check pending
deepseekdeepseek-harnessdshdsh-plugin
Files & dataPlugin

Zero-core-change vision capability for DeepSeek Harness: the describe_image tool + profile bundle, installable via 'dsh plugin add'

Structure check pending
ai-agentsdeepseek-harnessdshdsh-plugin
Files & dataPlugin

Deterministic aggregate media budgets and safe request projections for DeepSeek Harness (dsh)

Structure check pending
cordis-plugindeepseek-harnessdshdsh-plugin
Files & dataPlugin

DSH bundle: Qwen multimodal bridge — vision (qwen3-vl), speech-to-text (qwen3-asr), text-to-image (qwen-image), for DeepSeek Harness

Structure check pending
asrdashscopedeepseek-harnessdsh-plugin
DSH PluginsPlugin

DSH plugin: provides image/video generation tools in chat, based on OpenAI-compatible API

Structure check pending
deepseek-harnessdshdsh-pluginimage-generation
Files & dataPlugin

dsh-visionary

zhuiyueya

Give text-only DeepSeek models eyes — a DeepSeek Harness plugin that transparently converts chat images into OCR text + vision-model descriptions before they reach the LLM. Configure vision backends (GLM-4V, Qwen-VL, Gemini, Ollama…) right in the Models settings page; multi-backend fallback chain, double-layer caching, no config files.

Structure check pending
deepseekdeepseek-harnessdsh-pluginimage-understanding
Files & dataPlugin

Let text-only models read images in DeepSeek Harness

Structure check pending
deepseek-harnessdshdsh-pluginimage-captioning