Back to catalog

GITHUB TOPIC

multimodal-ai

4projects include this topic

The "multimodal-ai" topic on GitHub groups 4 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by dsh-multimodal with 4 GitHub stars. dsh-multimodal — Give DeepSeek a pair of eyes and a paintbrush: paste screenshots/images straight into the conversation, the GLM vision model first transcribes the image content precisely (error messages, code, UI preserved verbatim), then DeepSeek continues with your question —— all in the same turn, seamless throughout; when an illustration is needed, DeepSeek automatically calls the text-to-image backend and shows the image in the conversation. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.

{count} projects

Exact GitHub Topic match

Files & dataPlugin

Give DeepSeek a pair of eyes and a paintbrush: paste screenshots/images straight into the conversation, the GLM vision model first transcribes the image content precisely (error messages, code, UI preserved verbatim), then DeepSeek continues with your question —— all in the same turn, seamless throughout; when an illustration is needed, DeepSeek automatically calls the text-to-image backend and shows the image in the conversation.

Structure check pending
deepseek-harnessdshdsh-plugindsh-plugins
InterfacePlugin

dsh全模态工作站插件,让模型支持视频、图片、语音的输入与输出,支持comfyui图像生成工具调用。Any-to-Any.

Structure check pending
computer-visiondeepseek-harnessdshdsh-plugin
Files & dataPlugin

Image understanding, OCR, and persistent visual evidence for text-only DeepSeek Harness models

Structure check pending
ai-agentscomputer-visiondeepseekdeepseek-harness
Files & dataPlugin

AI-powered PDF reader for DeepSeek Harness with annotations, multi-PDF workflows, mixed image-text evidence, and on-demand OCR.

Structure check pending
ai-pdf-readerchat-with-pdfdeepseek-harnessdsh-plugin