Technology subcategory

Evaluations

Harnesses and judges that score agent or model behavior before and after change.

Evaluations packs measure whether an agent or model got better or worse. They are the falsifiers for autonomy claims.

Without evals, “we added an agent” cannot be distinguished from noise. Pair with Observability for traces and Guardrails for hard policy stops.

See also

Technologies

Features

Stacks

Document management

GitHub in this term

Primary repositories linked from member packs and devices.