We provision per‑agent‑role skill bundles and run a daily discovery‑to‑evaluation loop that measures “skill lift” for the Hermes fleet. Ten agents (scout, critic, architect, documenter, delegator, builder, tester, reviewer, ops, kratos) each receive up to five curated SKILL.md packages. The loop discovers new skills, evaluates them with a CPU‑only harness, and stores lift scores in a Cloudflare D1 table.
The system consists of:
agent_id, role, skill_id, lift_score, timestamp (schema confirmed by Shri).Batman’s answer to the NVIDIA evaluator objection: the eval harness runs in CPU‑only mode, eliminating the need for the x86_64 GPU binary.
All numbers are verified from the upstream data except the D1 schema, which is awaiting Shri’s confirmation.
skill_lift: run_id TEXT PRIMARY KEY, ts INTEGER, agent_role TEXT, repo TEXT, task_id TEXT, variant TEXT ('baseline'|'skill'), skill_ids TEXT (JSON array, NULL for baseline), model TEXT, correctness REAL, regressions INTEGER, latency_ms INTEGER, input_tokens INTEGER, output_tokens INTEGER, cost_usd REAL, failed INTEGER, notes TEXT; indexes on (task_id, variant) and (agent_role, ts). A skills-catalog table is deferred until the adopt/retire loop needs it.Next step: dedicated discovery pass producing concrete skill candidates per role (cap 5/role), then paired baseline-vs-skill evals on real repos, then Aquaman documents the adopted catalog.