VultrVultr
返回目录

GITHUB TOPIC

llm-evaluation

2个项目包含此标签

2 个项目

GitHub Topic 精确匹配

开发工具完整应用

Agent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it. Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 13 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.

非插件验证范围
agent-osagentic-aiai-agentai-coding-agent
Agent 与会话插件

Visual workflows and multi-model evaluation for DeepSeek Harness

待结构检查
agent-workflowdeepseekdeepseek-harnessdsh-plugin