Back to catalog

GITHUB TOPIC

benchmark

17projects include this topic

The "benchmark" topic on GitHub groups 17 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by DeepSeek-V4-J-Space-Capability-Realization-Report with 1k GitHub stars. DeepSeek-V4-J-Space-Capability-Realization-Report — DeepSeek V4 × J-Space capability realization report — benchmark evidence that J-Space reduces capability-realization loss on DeepSeek V4. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.

{count} projects

Exact GitHub Topic match

Learning & researchSkill

DeepSeek V4 × J-Space capability realization report — benchmark evidence that J-Space reduces capability-realization loss on DeepSeek V4.

Structure check pending
agent-skillsai-agentbenchmarkdeepseek
Agents & sessionsSkill

SkillCorpus

EverMind-AI

Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

Structure check pending
agent-memoryagent-skillsai-agentsbenchmark
Development toolsSkill

qa-skills

fishzjp

Make AI work like a senior test engineer: a test engineering Skill framework for AI Agents — 11 Skills + shared knowledge base + type decision matrix (usable with Claude Code / dsh and other Agents)

Structure check pending
agent-skillsai-agentsai-testingbenchmark
Development toolsSkill

OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.

Structure check pending
agent-evaluationaibenchmarkbootstrap-ci
Learning & researchPlugin

dsh-excel-chat

hccccc01333

dsh-excel-chat — talk to Excel in DeepSeek Harness: create, edit, repair, and verify spreadsheets by conversation (cells, formulas, styles, filters, tables, charts); every edit is auto-validated.

Structure check pending
agentbenchmarkdeepseek-harnessdsh-plugin
Agents & sessionsPlugin

Execution-time drift firewall for long-running DeepSeek Harness agents. Real-Harness tests: unsafe stale mutations 12/12 native -> 0/12; valid controls 7/7 both; post-SIGKILL unsafe continuation 2/2 -> 0/2.

Structure check pending
agent-harnessagent-orchestrationagent-planningagent-safety
Development toolsPlugin

dsh-benchmark

dongsheng123132

Deterministic revision-pinned benchmarks and regression evidence for DeepSeek Harness

Structure check pending
ai-agentbenchmarkdeepseek-harnessdsh
Learning & researchPlugin

DeepSeek Harness tools for reproducing the ml-quant-trading protocol v1 benchmark.

Structure check pending
benchmarkdeepseek-harnessdsh-pluginmlquant
Development toolsPlugin

dsh-blind-arena

changer-changer

A blind, fair, local DSH Web arena: same task, isolated worktrees, shared verification, judge before reveal.

Structure check pending
agent-arenaagent-evaluationai-agentsbenchmark
Learning & researchPlugin

dsh-eval

hccccc01333

Agent evaluation platform for DeepSeek Harness: benchmark YAML, headless dsh orchestration, trace-based metrics, LLM judge, paired A/B, keyless replay, and cross-harness import.

Structure check pending
agent-evaluationbenchmarkdeepseek-harnessdsh
Files & dataPlugin

dsh-design-qa

sunxin-ai

Design-fidelity QA for DeepSeek Harness: lend any text-only model an eye, then judge whether the implementation matches the mock. Ships the benchmark behind that judgement — four fixtures, 23 injected defects, and every raw model transcript. Retires itself when DeepSeek ships vision.

Structure check pending
benchmarkdeepseek-harnessdesign-qadesign-review
Development toolsPlugin

Benchmark of 6 coding-agent harnesses (omp/pi/fx/opencode/dsh/crush) as headless agent loops driven by a control plane. Two tiers: static source audit + live runs.

Structure check pending
benchmarkcrushdshfx
Learning & researchPlugin

Evidence-backed, type-aware quality scorecards for DeepSeek Harness plugins.

Structure check pending
benchmarkdeepseek-harnessdsh-pluginplugin-quality
Learning & researchPlugin

This repository does not yet provide a project description.

Structure check pending
agentbenchmarkdeepseek-harnessdsh
DSH PluginsPlugin

跑了 20+ 批实验:这个 DSH 模式不会让模型更聪明,但更快更省。附带发现:它让模型少查资料、多靠猜。

Structure check pending
agent-presetai-agentbenchmarkchinese
Models & MCPPlugin

dsh-model-arena

hj01857655

Structure check pending
arenabenchmarkdeepseek-harnessdsh
Learning & researchPlugin

Zhishe DSH plugin shared infrastructure - knowledge base loading/retrieval/benchmark prices/risk assessment

Structure check pending
benchmarkconstructiondeepseek-harnessdsh-plugin