DeepSeek-V4-J-Space-Capability-Realization-Report
Tiger3807861189
DeepSeek V4 × J-Space capability realization report — benchmark evidence that J-Space reduces capability-realization loss on DeepSeek V4.
GITHUB TOPIC
17projects include this topic
The "benchmark" topic on GitHub groups 17 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by DeepSeek-V4-J-Space-Capability-Realization-Report with 1k GitHub stars. DeepSeek-V4-J-Space-Capability-Realization-Report — DeepSeek V4 × J-Space capability realization report — benchmark evidence that J-Space reduces capability-realization loss on DeepSeek V4. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.
{count} projects
Exact GitHub Topic match
Tiger3807861189
DeepSeek V4 × J-Space capability realization report — benchmark evidence that J-Space reduces capability-realization loss on DeepSeek V4.
EverMind-AI
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
fishzjp
Make AI work like a senior test engineer: a test engineering Skill framework for AI Agents — 11 Skills + shared knowledge base + type decision matrix (usable with Claude Code / dsh and other Agents)
lizhiyao
OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.
hccccc01333
dsh-excel-chat — talk to Excel in DeepSeek Harness: create, edit, repair, and verify spreadsheets by conversation (cells, formulas, styles, filters, tables, charts); every edit is auto-validated.
1052326311
Execution-time drift firewall for long-running DeepSeek Harness agents. Real-Harness tests: unsafe stale mutations 12/12 native -> 0/12; valid controls 7/7 both; post-SIGKILL unsafe continuation 2/2 -> 0/2.
dongsheng123132
Deterministic revision-pinned benchmarks and regression evidence for DeepSeek Harness
initial-d
DeepSeek Harness tools for reproducing the ml-quant-trading protocol v1 benchmark.
changer-changer
A blind, fair, local DSH Web arena: same task, isolated worktrees, shared verification, judge before reveal.
hccccc01333
Agent evaluation platform for DeepSeek Harness: benchmark YAML, headless dsh orchestration, trace-based metrics, LLM judge, paired A/B, keyless replay, and cross-harness import.
sunxin-ai
Design-fidelity QA for DeepSeek Harness: lend any text-only model an eye, then judge whether the implementation matches the mock. Ships the benchmark behind that judgement — four fixtures, 23 injected defects, and every raw model transcript. Retires itself when DeepSeek ships vision.
aaronjmars
Benchmark of 6 coding-agent harnesses (omp/pi/fx/opencode/dsh/crush) as headless agent loops driven by a control plane. Two tiers: static source audit + live runs.
B1lli
Evidence-backed, type-aware quality scorecards for DeepSeek Harness plugins.
bpc-oss
This repository does not yet provide a project description.
walksls
跑了 20+ 批实验:这个 DSH 模式不会让模型更聪明,但更快更省。附带发现:它让模型少查资料、多靠猜。
hj01857655
siweimofang
Zhishe DSH plugin shared infrastructure - knowledge base loading/retrieval/benchmark prices/risk assessment