Back to catalog

GITHUB TOPIC

llm-evaluation

15projects include this topic

The "llm-evaluation" topic on GitHub groups 15 open-source projects in the DeepSeek Harness (DSH) ecosystem, led by SkillCorpus with 202 GitHub stars. SkillCorpus — Open-source infrastructure that turns scattered SKILL. Every project here is indexed by DSH Universe with live GitHub data — stars, activity and install status — so you can compare and install directly.

{count} projects

Exact GitHub Topic match

Agents & sessionsSkill

SkillCorpus

EverMind-AI

Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

Structure check pending
agent-memoryagent-skillsai-agentsbenchmark
DSH PluginsSkill

suede-creator-skills

JasonColapietro

Open-source AI skills for SEO, AI search visibility, conversion copy, marketing strategy, and business operations. Reusable workflows for Claude Code and Codex, plus code review, app delivery, and creator tools.

Structure check pending
agent-orchestrationagent-skillagent-skillsai-agents
Development toolsSkill

OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.

Structure check pending
agent-evaluationaibenchmarkbootstrap-ci
Development toolsPlugin

dsh-blind-arena

changer-changer

A blind, fair, local DSH Web arena: same task, isolated worktrees, shared verification, judge before reveal.

Structure check pending
agent-arenaagent-evaluationai-agentsbenchmark
DSH PluginsPlugin

Evidence-gated completion and trusted code verification for DeepSeek Harness agents

Structure check pending
agent-reliabilityai-agentsdeepseek-harnessdsh-plugin
Learning & researchPlugin

dsh-plugin-evaluation-standards

dsh-plugin-evaluation

Open evaluation datasets, test cases, and metrics for DSH plugins.

Structure check pending
benchmarksdeepseek-harnessdeepseek-harness-plugindsh
OperationsPlugin

Controlled request-surface replay and regression workbench for DeepSeek Harness

Structure check pending
agent-observabilitydeepseekdeepseek-harnessdsh
Agents & sessionsPlugin

Visual workflows and multi-model evaluation for DeepSeek Harness

Structure check pending
agent-workflowdeepseekdeepseek-harnessdsh-plugin
InterfacePlugin

Paired A/B evaluation infrastructure for DeepSeek Harness (dsh) components: interleaved repeated runs through the real runtime, verifier self-checks, one-variable arms, regression gating, cache- and calendar-aware cost, CLI + web UI

Structure check pending
ab-testingagent-evaluationdeepseek-harnessdsh-plugin
Agents & sessionsPlugin

Proof-carrying correction learning for DSH: explicit adoption, scoped recall, and reconstructable delivery.

Structure check pending
agent-learningagent-memoryai-agentcordis
DSH PluginsPlugin

rcos

Foshowithit

RCOS — Recursive Capability Operating System: workflow-first control plane with eval-driven capability promotion, wiring Pi + DSH + Archon

Structure check pending
agentic-workflowsai-agentsautomationdeveloper-tools
DSH PluginsPlugin

基于 DSH 与 Harbor 的代码 Agent 外部验收与失败修复实验。

Structure check pending
agent-harnessai-agentsdeepseek-harnessharbor
Development toolsPlugin

dsh-plugin-abtest

Morriaty-The-Murderer

Reproducible paired A/B experiments for DeepSeek Harness plugins, with isolated runners, auditable evidence, and deterministic promotion gates.

Structure check pending
ab-testingagent-evaluationai-agentsbenchmarking
DSH PluginsSkill

Cross-model subagent preset + dispatch skill for DeepSeek Harness (DSH): a reliable component, not just a task solver. Cross-model DSH coding subagent preset + dispatch skill: parseable output, bounded behavior, honest reporting

Structure check pending
agentagent-presetagent-skillagent-skills
Development toolsPlugin

Controlled A/B comparisons and evidence-backed reports for DeepSeek Harness plugins and presets.

Structure check pending
ab-testingai-agentsbenchmarkingdeepseek-harness