oh-my-knowledge
lizhiyao
OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.
GITHUB TOPIC
12projects include this topic
The "agent-evaluation" topic on GitHub groups 12 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by oh-my-knowledge with 18 GitHub stars. oh-my-knowledge — OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.
{count} projects
Exact GitHub Topic match
lizhiyao
OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.
timwhitez
Evidence-first, crash-resumable self-evolution engine for DeepSeek Harness and Harbor.
changer-changer
A blind, fair, local DSH Web arena: same task, isolated worktrees, shared verification, judge before reveal.
daydreamer0213
An auditable learning control plane for long-running agents, built on DSH.
hccccc01333
Agent evaluation platform for DeepSeek Harness: benchmark YAML, headless dsh orchestration, trace-based metrics, LLM judge, paired A/B, keyless replay, and cross-harness import.
ShawnSiao
Planned repeatable agent and plugin regression evaluation for DeepSeek Harness
TT-Wang
Paired A/B evaluation infrastructure for DeepSeek Harness (dsh) components: interleaved repeated runs through the real runtime, verifier self-checks, one-variable arms, regression gating, cache- and calendar-aware cost, CLI + web UI
ArmyWas
Post-run contract verifier and persistent SDK canary for DeepSeek Harness subagents.
Harzva
DSH-native multi-runtime baseline, ablation, and reproducible evaluation control plane
Morriaty-The-Murderer
Reproducible paired A/B experiments for DeepSeek Harness plugins, with isolated runners, auditable evidence, and deterministic promotion gates.
Muredsa
Install with npm i dsh-benchup. Reproducible, profile-aware benchmarks for DeepSeek Harness — compare models, plugins, prompts, and agent strategies.
yu-xin-c
Offline evaluation and guarded self-evolution loop for DeepSeek Harness