Back to catalog

GITHUB TOPIC

evaluation

9projects include this topic

The "evaluation" topic on GitHub groups 9 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by fable-method with 2.3k GitHub stars. fable-method — The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.

{count} projects

Exact GitHub Topic match

DSH PluginsSkill

fable-method

Sahir619

The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

Structure check pending
agent-skillsai-agentsclaudeclaude-code
Development toolsPlugin

DSH plugin eval tool: YAML case-driven real agent regression eval + baseline comparison PASS/WARN/FAIL gates|Regression eval harness for DeepSeek Harness plugins

Structure check pending
deepseek-harnessdshdsh-pluginevaluation
Development toolsPlugin

DSH-arena

Apageoflove

Local-first experiment and evaluation workbench plugin for DeepSeek Harness (DSH).

Structure check pending
ai-agentarenabenchmarkingcordis
Development toolsPlugin

No project description provided for this repository yet.

Structure check pending
deepseek-harnessdsh-pluginevaluationregression-testing
DSH PluginsPlugin

Reproducible DSH profile and patch experiment matrices with reports and policy gates

Structure check pending
deepseekdeepseek-harnessdsh-pluginevaluation
DSH PluginsPlugin

duo

CZ-ZL

Composable DSH-native evaluation and optimization with explicit evidence modes, bounded budgets and traceable results

Structure check pending
agent-toolscordisdeepseekevaluation
DSH PluginsPlugin

dsh-verdict

hj01857655

Measure whether a change to your dsh setup actually helped: register repeatable cases, run them, diff before/after. DeepSeek Harness plugin.

Structure check pending
deepseek-harnessdshdsh-plugindsh-plugins
Models & MCPPlugin

Public, reusable DeepSeek Harness plugins and skills: workflow canvas toolkit, blind eval harness, LLM cost lab, incident ledger.

Structure check pending
deepseek-harnessdsh-pluginevaluationmcp
InterfaceSkill

Clean, portable agent skills for DeepSeek Harness, Claude Code and any harness reading the Agent Skills standard. Plain text, nothing to build. Every skill ships a checker and a measured evaluation protocol — and says what was not verified.

Structure check pending
agent-skillsai-agentsclaude-codedeepseek-harness