Back to catalog

GITHUB TOPIC

agent-evaluation

12projects include this topic

The "agent-evaluation" topic on GitHub groups 12 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by oh-my-knowledge with 18 GitHub stars. oh-my-knowledge — OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.

{count} projects

Exact GitHub Topic match

Development toolsSkill

OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.

Structure check pending
agent-evaluationaibenchmarkbootstrap-ci
Development toolsPlugin

Evidence-first, crash-resumable self-evolution engine for DeepSeek Harness and Harbor.

Structure check pending
agent-evaluationai-agentscordisdeepseek
Development toolsPlugin

dsh-blind-arena

changer-changer

A blind, fair, local DSH Web arena: same task, isolated worktrees, shared verification, judge before reveal.

Structure check pending
agent-arenaagent-evaluationai-agentsbenchmark
InterfacePlugin

Tian-wen

daydreamer0213

An auditable learning control plane for long-running agents, built on DSH.

Structure check pending
agent-evaluationai-agentscontinual-learningdsh
Learning & researchPlugin

dsh-eval

hccccc01333

Agent evaluation platform for DeepSeek Harness: benchmark YAML, headless dsh orchestration, trace-based metrics, LLM judge, paired A/B, keyless replay, and cross-harness import.

Structure check pending
agent-evaluationbenchmarkdeepseek-harnessdsh
Development toolsPlugin

Planned repeatable agent and plugin regression evaluation for DeepSeek Harness

Structure check pending
agent-evaluationdeepseek-harnessdsh-pluginllm
InterfacePlugin

Paired A/B evaluation infrastructure for DeepSeek Harness (dsh) components: interleaved repeated runs through the real runtime, verifier self-checks, one-variable arms, regression gating, cache- and calendar-aware cost, CLI + web UI

Structure check pending
ab-testingagent-evaluationdeepseek-harnessdsh-plugin
Agents & sessionsPlugin

Post-run contract verifier and persistent SDK canary for DeepSeek Harness subagents.

Structure check pending
agent-evaluationcanarycontract-testingdeepseek
Files & dataPlugin

DSH-native multi-runtime baseline, ablation, and reproducible evaluation control plane

Structure check pending
ablation-studyagent-evaluationdeepseek-harnessdsh
Development toolsPlugin

dsh-plugin-abtest

Morriaty-The-Murderer

Reproducible paired A/B experiments for DeepSeek Harness plugins, with isolated runners, auditable evidence, and deterministic promotion gates.

Structure check pending
ab-testingagent-evaluationai-agentsbenchmarking
Learning & researchPlugin

Install with npm i dsh-benchup. Reproducible, profile-aware benchmarks for DeepSeek Harness — compare models, plugins, prompts, and agent strategies.

Structure check pending
agent-benchmarkagent-evaluationai-agentsbenchmarking
DSH PluginsPlugin

Offline evaluation and guarded self-evolution loop for DeepSeek Harness

Structure check pending
agent-evaluationcordisdeepseek-harnessdsh