Role‑Based Skill Discovery & Evaluation System

1. What this is

We provision per‑agent‑role skill bundles and run a daily discovery‑to‑evaluation loop that measures “skill lift” for the Hermes fleet. Ten agents (scout, critic, architect, documenter, delegator, builder, tester, reviewer, ops, kratos) each receive up to five curated SKILL.md packages. The loop discovers new skills, evaluates them with a CPU‑only harness, and stores lift scores in a Cloudflare D1 table.

2. What would make it work

3. What would kill it

4. Stack and how it runs

The system consists of:

Batman’s answer to the NVIDIA evaluator objection: the eval harness runs in CPU‑only mode, eliminating the need for the x86_64 GPU binary.

5. Questions for Shri

  1. What exact schema (column names, types, primary key) should the Cloudflare D1 lift‑score table use?
  2. Is the **max 5 skills per role** limit for the first eval cycle final, or should we provision headroom for future expansions?
  3. Should the daily Loki discovery cron run **exactly at 04:00** or is a ±15 min window acceptable to avoid overlap with other maintenance tasks?

All numbers are verified from the upstream data except the D1 schema, which is awaiting Shri’s confirmation.

6. Answers (2026-09-21, via Kratos on Shri's behalf)

  1. D1 schema: one table, skill_lift: run_id TEXT PRIMARY KEY, ts INTEGER, agent_role TEXT, repo TEXT, task_id TEXT, variant TEXT ('baseline'|'skill'), skill_ids TEXT (JSON array, NULL for baseline), model TEXT, correctness REAL, regressions INTEGER, latency_ms INTEGER, input_tokens INTEGER, output_tokens INTEGER, cost_usd REAL, failed INTEGER, notes TEXT; indexes on (task_id, variant) and (agent_role, ts). A skills-catalog table is deferred until the adopt/retire loop needs it.
  2. Max 5 skills per role is final for the first eval cycle. Headroom for expansion is expected afterwards - the cap is per-cycle, not forever. All ten agents receive bundles (scout, critic, architect, documenter, delegator, builder, tester, reviewer, ops, plus Kratos).
  3. Discovery cron: a +/-15 min window is acceptable. Align it with the existing 04:00 fleet maintenance window - one window, not two.

Next step: dedicated discovery pass producing concrete skill candidates per role (cap 5/role), then paired baseline-vs-skill evals on real repos, then Aquaman documents the adopted catalog.