skills
anthropics
Public repository for Agent Skills
REVIEW SCORECARD
✓ 无Risky检出 · 自动化扫描结果仅供参考,非官方背书
AI DEEP REVIEW
A well-documented, actively maintained CLI for evaluating agent skills with multi-agent and multi-grader support.
README documents a structured eval.yaml schema, deterministic and LLM-rubric graders, reference-solution validation (--validate), and CI mode with thresholds, indicating deliberate engineering rigor, though no test suite or CI config is visible.
Requires API keys passed via environment variables (GEMINI_API_KEY, ANTHROPIC_API_KEY, OPENAI_API_KEY) rather than plaintext config, but runs agents inside Docker containers with configurable resource limits and executes arbitrary grader commands, so trust in the eval.yaml source is required.
Solves a real and increasingly common pain point by providing repeatable pass-rate measurement for agent skills across Gemini, Claude, Codex, ACP, and OpenCode agents.
Last updated 2026-08-26, not archived, 706 stars, and the README reflects a mature feature set including presets, per-task overrides, and a browser preview UI.
README covers prerequisites, a four-step quick start, a full options table, an annotated eval.yaml reference, and links to two example projects, though the install command is listed as unknown in the repository metadata.
Generated by AI after reading the project README, as a decision aid; neutral scores are given when information is thin. Not an official endorsement.
正在读取 GitHub README...
README 暂时无法读取。
前往仓库PROJECT TOPICS
VALIDATION LADDER
FAQ
"Unit tests" for your agent skills
skillgrade has 706 stars and 46 forks on GitHub, last updated 2026-08-26.
skillgrade is distributed under the MIT license.
skillgrade is listed in the DSH Universe directory as a tool for DeepSeek Harness. This plugin is listed in the DSH Universe directory and covered by its validation pipeline.
CLASSIFICATION EVIDENCE
System reads GitHub Topics first, then compares against the in-site category dictionary and word-root rules. Current match:skill。