feature · r2-32-eval-harnesses
M
5 tags · 8 out · 8 in · 3 artifacts · 1 gaps · 1 corpus docs
How should agent quality be measured, gated, and promoted — combining deterministic checks, calibrated LLM-as-judge, and statistical regression CI — without blocking every PR or trusting uncalibrated vibes?
This topic was frozen in the corporate agentic OS research corpus (wave 10) to answer architecture and measurement questions that sit under the locked loop (observe→understand→predict→intervene→resolve→verify→learn). Stage-1 criticality: M. Research/Spec/Platform readiness: R5/S4/P0.
Graph edges for this feature — technologies, category lens, companies, DM.
features/r2-32-eval-harnesses/README.md
feature missing document-management.md
T32 — Eval Harnesses & CI Gates for Agents · ok
features/r2-32-eval-harnesses/README.md
Research path: features/r2-32-eval-harnesses